{
 "S1sh1::signage": {
  "fp": "f0142979c86dcc17",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S1sh1": {
  "input_fingerprint": "b8503bc5a18f8e77",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닷속 쓰레기 더미 위로 튀어나온 낡은 로봇 손가락에 입을 맞추듯 주둥이를 댄 물고기의 측면.\n\nLOCATION (lock): Underwater at the seabed, beside a submerged rubbish heap with a robot finger protruding from it. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Protruding robot finger (Old and protruding from submerged rubbish, touching the fish's snout) — Seen from the side, with its contact surface level with the fish's snout; used as Small focal anchor to the right of the contact point; Submerged rubbish pile (Accumulated on the seabed around the protruding finger); used as Layered lower-frame context that establishes where the finger emerges.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Blue underwater ambient light and restrained contrast preserve the quiet, tactile contact without exaggerating it into a luminous effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In blue daytime seawater, an old scrap robot lies buried in seabed rubbish with a finger protruding; its body is entangled in netting, and its chest bears a worn Ubik logo. Unidentified fish swim among the submerged debris.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닷속 쓰레기 더미 위로 튀어나온 낡은 로봇 손가락에 입을 맞추듯 주둥이를 댄 물고기의 측면.\n\nLOCATION (lock): Underwater at the seabed, beside a submerged rubbish heap with a robot finger protruding from it. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Protruding robot finger (Old and protruding from submerged rubbish, touching the fish's snout) — Seen from the side, with its contact surface level with the fish's snout; used as Small focal anchor to the right of the contact point; Submerged rubbish pile (Accumulated on the seabed around the protruding finger); used as Layered lower-frame context that establishes where the finger emerges.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Blue underwater ambient light and restrained contrast preserve the quiet, tactile contact without exaggerating it into a luminous effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In blue daytime seawater, an old scrap robot lies buried in seabed rubbish with a finger protruding; its body is entangled in netting, and its chest bears a worn Ubik logo. Unidentified fish swim among the submerged debris.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닷속 쓰레기 더미 위로 튀어나온 낡은 로봇 손가락에 입을 맞추듯 주둥이를 댄 물고기의 측면.\n\nLOCATION (lock): Underwater at the seabed, beside a submerged rubbish heap with a robot finger protruding from it. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Protruding robot finger (Old and protruding from submerged rubbish, touching the fish's snout) — Seen from the side, with its contact surface level with the fish's snout; used as Small focal anchor to the right of the contact point; Submerged rubbish pile (Accumulated on the seabed around the protruding finger); used as Layered lower-frame context that establishes where the finger emerges.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Blue underwater ambient light and restrained contrast preserve the quiet, tactile contact without exaggerating it into a luminous effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In blue daytime seawater, an old scrap robot lies buried in seabed rubbish with a finger protruding; its body is entangled in netting, and its chest bears a worn Ubik logo. Unidentified fish swim among the submerged debris.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수평으로 헤엄치는 물고기의 주둥이가 화면 우측에서 뻗어 나온 로봇 손가락 끝을 향해 정확히 맞닿아 있습니다.",
    "built_space": "푸른 빛이 도는 수중 해저면에 플라스틱, 그물 등 다양한 질감의 쓰레기 더미가 층을 이루어 쌓여 있습니다.",
    "entities": "측면 구도의 물고기, 부식된 질감의 낡은 로봇 손과 손가락, 바닷속 수중 쓰레기 더미가 모두 명확히 확인됩니다.",
    "hard_violations": [],
    "physics": "물고기는 수중에 자연스럽게 떠 있으며, 로봇 손은 쓰레기 더미 속에 물리적으로 안정되게 박혀 지지받고 있습니다."
   },
   {
    "label": "B",
    "direction": "물고기가 사선으로 뻗어 나온 로봇 손가락 끝을 향해 입을 대고 있습니다.",
    "built_space": "바닷속 해저면에 타이어와 플라스틱 등 여러 형태의 폐기물들이 흩어져 쌓여 있습니다.",
    "entities": "물고기의 측면, 금속 질감의 로봇 손과 위로 솟은 손가락, 해저 쓰레기 더미가 식별됩니다.",
    "hard_violations": [],
    "physics": "물고기는 물의 부력으로 떠 있고, 로봇 손은 바닥의 쓰레기 더미 위에 얹혀 지탱되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 대로 물고기의 주둥이와 로봇 손가락의 접촉면이 수평을 이루는 측면 구도를 매우 정확하고 자연스럽게 구현했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "물고기와 로봇 손가락의 접촉은 잘 표현되었으나, 손가락이 위로 솟구친 사선 형태여서 접촉면이 수평을 이룬다는 지시사항의 정확도가 A에 비해 약간 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수평으로 헤엄치는 물고기의 주둥이가 화면 우측에서 뻗어 나온 로봇 손가락 끝을 향해 정확히 맞닿아 있습니다.",
        "built_space": "푸른 빛이 도는 수중 해저면에 플라스틱, 그물 등 다양한 질감의 쓰레기 더미가 층을 이루어 쌓여 있습니다.",
        "entities": "측면 구도의 물고기, 부식된 질감의 낡은 로봇 손과 손가락, 바닷속 수중 쓰레기 더미가 모두 명확히 확인됩니다.",
        "hard_violations": [],
        "physics": "물고기는 수중에 자연스럽게 떠 있으며, 로봇 손은 쓰레기 더미 속에 물리적으로 안정되게 박혀 지지받고 있습니다."
       },
       {
        "label": "B",
        "direction": "물고기가 사선으로 뻗어 나온 로봇 손가락 끝을 향해 입을 대고 있습니다.",
        "built_space": "바닷속 해저면에 타이어와 플라스틱 등 여러 형태의 폐기물들이 흩어져 쌓여 있습니다.",
        "entities": "물고기의 측면, 금속 질감의 로봇 손과 위로 솟은 손가락, 해저 쓰레기 더미가 식별됩니다.",
        "hard_violations": [],
        "physics": "물고기는 물의 부력으로 떠 있고, 로봇 손은 바닥의 쓰레기 더미 위에 얹혀 지탱되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 대로 물고기의 주둥이와 로봇 손가락의 접촉면이 수평을 이루는 측면 구도를 매우 정확하고 자연스럽게 구현했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "물고기와 로봇 손가락의 접촉은 잘 표현되었으나, 손가락이 위로 솟구친 사선 형태여서 접촉면이 수평을 이룬다는 지시사항의 정확도가 A에 비해 약간 떨어집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수평으로 헤엄치는 물고기의 주둥이가 화면 우측에서 뻗어 나온 로봇 손가락 끝을 향해 정확히 맞닿아 있습니다.",
        "built_space": "푸른 빛이 도는 수중 해저면에 플라스틱, 그물 등 다양한 질감의 쓰레기 더미가 층을 이루어 쌓여 있습니다.",
        "entities": "측면 구도의 물고기, 부식된 질감의 낡은 로봇 손과 손가락, 바닷속 수중 쓰레기 더미가 모두 명확히 확인됩니다.",
        "hard_violations": [],
        "physics": "물고기는 수중에 자연스럽게 떠 있으며, 로봇 손은 쓰레기 더미 속에 물리적으로 안정되게 박혀 지지받고 있습니다."
       },
       {
        "label": "B",
        "direction": "물고기가 사선으로 뻗어 나온 로봇 손가락 끝을 향해 입을 대고 있습니다.",
        "built_space": "바닷속 해저면에 타이어와 플라스틱 등 여러 형태의 폐기물들이 흩어져 쌓여 있습니다.",
        "entities": "물고기의 측면, 금속 질감의 로봇 손과 위로 솟은 손가락, 해저 쓰레기 더미가 식별됩니다.",
        "hard_violations": [],
        "physics": "물고기는 물의 부력으로 떠 있고, 로봇 손은 바닥의 쓰레기 더미 위에 얹혀 지탱되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "물고기 측면을 더 크게 담고 접촉점 오른쪽의 손가락을 비교적 작게 유지하여, 주둥이 접촉을 중심으로 한 클로즈업 지시에 더 충실하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "수평으로 마주 닿은 주둥이와 손끝은 정확하지만, 물고기보다 로봇 손과 넓은 해저 맥락의 비중이 커 작은 손가락을 보조 초점으로 삼으라는 구도에서 다소 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 향한 측면이며, 주둥이가 오른쪽 로봇 손끝에 직접 닿는다. 손가락은 오른쪽 아래에서 왼쪽 위로 뻗어 물고기의 입 높이에서 만난다. 배경 물고기들은 여러 방향으로 놓여 있으나 흐려서 개별 시선이나 이동 방향은 확정하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 하단에 겹겹이 쌓인 폐기물 더미가 있고, 그 오른쪽에서 로봇 손 하나가 드러난다. 손 뒤에는 타이어 하나, 왼쪽 아래에는 밧줄과 그물성 섬유, 주변에는 파손된 용기와 부품들이 보인다. 손의 밑부분은 잔해 속에 묻혀 있어 지정된 해저 장소와 맞는다.",
        "entities": "큰 비늘과 지느러미를 가진 물고기 한 마리가 주 피사체이며 배경에도 여러 물고기가 있다. 접촉 대상은 부식과 관절 구조가 보이는 낡은 금속 로봇 손가락이다. 푸른 낮 수중광, 해저 쓰레기, 얽힌 섬유가 표현되어 있다. 로봇 몸통과 가슴은 가려져 그물의 몸통 결박 상태와 가슴 표식은 확인할 수 없지만, 클로즈업에서 드러낼 필요는 없다. 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 부력과 지느러미로 유영 자세를 유지하므로 공중 부유 문제가 없다. 로봇 손가락은 관절을 통해 손에 연결되고, 손의 밑부분은 해저 잔해에 지지된다. 타이어와 다른 폐기물도 바닥이나 서로 겹친 잔해에 놓여 있다. 주둥이와 손끝의 가벼운 접촉은 물리적으로 자연스럽다."
       },
       {
        "label": "B",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 바라보는 측면이며, 입술이 왼쪽으로 뻗은 로봇 손끝에 닿는다. 손가락은 거의 수평으로 뻗어 접촉면과 주둥이 높이가 일치한다. 배경의 작은 물고기들은 흐려서 정확한 시선과 진행 방향을 판별하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 오른쪽 하단의 쓰레기 더미에서 로봇 손 하나가 드러나고, 왼쪽에는 모래 해저가 넓게 보인다. 손 주변에 그물과 굵은 밧줄, 기울어진 폐용기와 판재가 층을 이룬다. 지정된 장소에는 맞지만 손 전체와 주변 잔해가 차지하는 화면 비중이 크다.",
        "entities": "주 피사체는 어두운 비늘의 물고기 한 마리이며 배경에도 작은 물고기들이 보인다. 상대 물체는 녹슬고 마모된 관절식 로봇 손으로, 펼친 손가락 하나와 접힌 다른 손가락들이 식별된다. 푸른 수중광과 해저 폐기물, 그물이 있다. 묻힌 몸통과 가슴 표식은 보이지 않아 확인 대상 밖이다. 사람, 읽을 수 있는 글자나 그래픽은 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 지느러미를 펼친 정상적인 유영 자세다. 뻗은 손가락은 로봇 손과 연결되어 있고, 손과 손목은 아래쪽 잔해에 받쳐져 있다. 그물과 밧줄은 폐기물 위에 걸쳐 있으며 주변 용기들도 잔해에 기대어 있다. 주둥이 접촉이나 물체의 지지에서 물리적 모순은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "물고기 측면을 더 크게 담고 접촉점 오른쪽의 손가락을 비교적 작게 유지하여, 주둥이 접촉을 중심으로 한 클로즈업 지시에 더 충실하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "수평으로 마주 닿은 주둥이와 손끝은 정확하지만, 물고기보다 로봇 손과 넓은 해저 맥락의 비중이 커 작은 손가락을 보조 초점으로 삼으라는 구도에서 다소 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 향한 측면이며, 주둥이가 오른쪽 로봇 손끝에 직접 닿는다. 손가락은 오른쪽 아래에서 왼쪽 위로 뻗어 물고기의 입 높이에서 만난다. 배경 물고기들은 여러 방향으로 놓여 있으나 흐려서 개별 시선이나 이동 방향은 확정하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 하단에 겹겹이 쌓인 폐기물 더미가 있고, 그 오른쪽에서 로봇 손 하나가 드러난다. 손 뒤에는 타이어 하나, 왼쪽 아래에는 밧줄과 그물성 섬유, 주변에는 파손된 용기와 부품들이 보인다. 손의 밑부분은 잔해 속에 묻혀 있어 지정된 해저 장소와 맞는다.",
        "entities": "큰 비늘과 지느러미를 가진 물고기 한 마리가 주 피사체이며 배경에도 여러 물고기가 있다. 접촉 대상은 부식과 관절 구조가 보이는 낡은 금속 로봇 손가락이다. 푸른 낮 수중광, 해저 쓰레기, 얽힌 섬유가 표현되어 있다. 로봇 몸통과 가슴은 가려져 그물의 몸통 결박 상태와 가슴 표식은 확인할 수 없지만, 클로즈업에서 드러낼 필요는 없다. 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 부력과 지느러미로 유영 자세를 유지하므로 공중 부유 문제가 없다. 로봇 손가락은 관절을 통해 손에 연결되고, 손의 밑부분은 해저 잔해에 지지된다. 타이어와 다른 폐기물도 바닥이나 서로 겹친 잔해에 놓여 있다. 주둥이와 손끝의 가벼운 접촉은 물리적으로 자연스럽다."
       },
       {
        "label": "A",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 바라보는 측면이며, 입술이 왼쪽으로 뻗은 로봇 손끝에 닿는다. 손가락은 거의 수평으로 뻗어 접촉면과 주둥이 높이가 일치한다. 배경의 작은 물고기들은 흐려서 정확한 시선과 진행 방향을 판별하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 오른쪽 하단의 쓰레기 더미에서 로봇 손 하나가 드러나고, 왼쪽에는 모래 해저가 넓게 보인다. 손 주변에 그물과 굵은 밧줄, 기울어진 폐용기와 판재가 층을 이룬다. 지정된 장소에는 맞지만 손 전체와 주변 잔해가 차지하는 화면 비중이 크다.",
        "entities": "주 피사체는 어두운 비늘의 물고기 한 마리이며 배경에도 작은 물고기들이 보인다. 상대 물체는 녹슬고 마모된 관절식 로봇 손으로, 펼친 손가락 하나와 접힌 다른 손가락들이 식별된다. 푸른 수중광과 해저 폐기물, 그물이 있다. 묻힌 몸통과 가슴 표식은 보이지 않아 확인 대상 밖이다. 사람, 읽을 수 있는 글자나 그래픽은 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 지느러미를 펼친 정상적인 유영 자세다. 뻗은 손가락은 로봇 손과 연결되어 있고, 손과 손목은 아래쪽 잔해에 받쳐져 있다. 그물과 밧줄은 폐기물 위에 걸쳐 있으며 주변 용기들도 잔해에 기대어 있다. 주둥이 접촉이나 물체의 지지에서 물리적 모순은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1875
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "프롬프트가 요구한 대로 물고기의 주둥이와 로봇 손가락의 접촉면이 수평을 이루는 측면 구도를 매우 정확하고 자연스럽게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "물고기와 로봇 손가락의 접촉은 잘 표현되었으나, 손가락이 위로 솟구친 사선 형태여서 접촉면이 수평을 이룬다는 지시사항의 정확도가 A에 비해 약간 떨어집니다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-7d01-79c1-ad2e-5c6ec233544d",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S1sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:45:43.737998+00:00",
  "fingerprint": "b0c147d7b0e90dbf70411ddea65fd3242b49e40f54a6210ee03502f955f07e1a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S1sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S1sh1_sel.png",
  "source_sha256": "31585cd1caef41d122e806cf99793db1e6651e0ddf18c54ee9489a9c91764d02",
  "file": "S1sh1_cine.png",
  "staged_sha256": "d12bac9d46b562171b8c7dd50ff590d8e816301f84130d95dcca37c98f8f9d7c",
  "latency_ms": 11212
 },
 "S2sh3::signage": {
  "fp": "f33ea41b6728da5d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S2sh3": {
  "input_fingerprint": "320b3540247f4d8b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is sprawled motionless among the rubbish deposited on the collection ship's deck, with a net wrapped around his body. The source does not establish which side of his body faces upward or the individual positions of his head, arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubbish has accumulated on the collection ship's deck, with the old scrap robot Charlie lying among it, still entangled in netting and bearing a worn Ubik chest logo. Floating waste around the ship includes shattered helicopter wreckage and a ship broken in half.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is sprawled motionless among the rubbish deposited on the collection ship's deck, with a net wrapped around his body. The source does not establish which side of his body faces upward or the individual positions of his head, arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubbish has accumulated on the collection ship's deck, with the old scrap robot Charlie lying among it, still entangled in netting and bearing a worn Ubik chest logo. Floating waste around the ship includes shattered helicopter wreckage and a ship broken in half.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is sprawled motionless among the rubbish deposited on the collection ship's deck, with a net wrapped around his body. The source does not establish which side of his body faces upward or the individual positions of his head, arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubbish has accumulated on the collection ship's deck, with the old scrap robot Charlie lying among it, still entangled in netting and bearing a worn Ubik chest logo. Floating waste around the ship includes shattered helicopter wreckage and a ship broken in half.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S2sh3__bgfirst_bg.png",
     "asset_id": "22001d0d-2139-4c65-801f-8fe7b7978313",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S2sh3.png",
     "asset_id": "b077dc93-9217-44ee-9e06-e7c3ca0dcd21",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_collection_deck_b9e7b9.png",
     "asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "로봇의 얼굴과 몸이 하늘을 향해 있음.",
    "built_space": "갑판 위 쓰레기 더미, 그물, 배경의 난파선 2척 등 위치 레퍼런스의 구성 요소가 정확히 배치됨.",
    "entities": "고철 로봇이 그물에 감겨 있으나, 찰리의 필수 외형(코트, 모자, 마스크 얼굴)이 완전히 누락된 대체 디자인임.",
    "hard_violations": [
     "[gemini-pro] 지정된 캐릭터(찰리)의 정체성 및 복장 완전 누락"
    ],
    "physics": "쓰레기 더미 위에 바닥의 지지를 받으며 중력에 맞게 자연스럽게 누워 있음."
   },
   {
    "label": "B",
    "direction": "로봇의 얼굴과 몸이 위를 향해 있음.",
    "built_space": "쓰레기 수거선 갑판, 배경의 바다와 난파선, 주변의 쓰레기 더미 등 공간적 요소가 잘 구현됨.",
    "entities": "찰리의 상반신(모자, 코트, 얼굴, 가슴 장갑)은 레퍼런스와 일치하나, 하반신은 위치 레퍼런스에 있던 흰색 원통형 로봇 다리가 그대로 섞여서 나타남.",
    "hard_violations": [
     "[gpt-high] 가슴 옷깃에 판독 가능한 'UBIK' 표기가 노출되어, 읽을 수 있는 글자와 로고를 금지한 최종 지시를 위반한다."
    ],
    "physics": "쓰레기 더미 위에 널브러져 있으며, 모든 신체 부위가 바닥의 지지를 받고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 핵심 외형(중절모, 코트, 마스크)과 공간적 배경을 잘 구현했으나, 하반신이 레퍼런스의 다른 로봇 다리와 섞여 묘사된 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "공간적 배경과 포즈는 적절하나, 찰리의 지정된 복장과 얼굴 형태가 전혀 반영되지 않은 다른 로봇이 등장하여 캐릭터 일치도가 매우 낮습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 얼굴과 몸이 하늘을 향해 있음.",
        "built_space": "갑판 위 쓰레기 더미, 그물, 배경의 난파선 2척 등 위치 레퍼런스의 구성 요소가 정확히 배치됨.",
        "entities": "고철 로봇이 그물에 감겨 있으나, 찰리의 필수 외형(코트, 모자, 마스크 얼굴)이 완전히 누락된 대체 디자인임.",
        "hard_violations": [
         "지정된 캐릭터(찰리)의 정체성 및 복장 완전 누락"
        ],
        "physics": "쓰레기 더미 위에 바닥의 지지를 받으며 중력에 맞게 자연스럽게 누워 있음."
       },
       {
        "label": "B",
        "direction": "로봇의 얼굴과 몸이 위를 향해 있음.",
        "built_space": "쓰레기 수거선 갑판, 배경의 바다와 난파선, 주변의 쓰레기 더미 등 공간적 요소가 잘 구현됨.",
        "entities": "찰리의 상반신(모자, 코트, 얼굴, 가슴 장갑)은 레퍼런스와 일치하나, 하반신은 위치 레퍼런스에 있던 흰색 원통형 로봇 다리가 그대로 섞여서 나타남.",
        "hard_violations": [],
        "physics": "쓰레기 더미 위에 널브러져 있으며, 모든 신체 부위가 바닥의 지지를 받고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 핵심 외형(중절모, 코트, 마스크)과 공간적 배경을 잘 구현했으나, 하반신이 레퍼런스의 다른 로봇 다리와 섞여 묘사된 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "공간적 배경과 포즈는 적절하나, 찰리의 지정된 복장과 얼굴 형태가 전혀 반영되지 않은 다른 로봇이 등장하여 캐릭터 일치도가 매우 낮습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 얼굴과 몸이 하늘을 향해 있음.",
        "built_space": "갑판 위 쓰레기 더미, 그물, 배경의 난파선 2척 등 위치 레퍼런스의 구성 요소가 정확히 배치됨.",
        "entities": "고철 로봇이 그물에 감겨 있으나, 찰리의 필수 외형(코트, 모자, 마스크 얼굴)이 완전히 누락된 대체 디자인임.",
        "hard_violations": [
         "지정된 캐릭터(찰리)의 정체성 및 복장 완전 누락"
        ],
        "physics": "쓰레기 더미 위에 바닥의 지지를 받으며 중력에 맞게 자연스럽게 누워 있음."
       },
       {
        "label": "B",
        "direction": "로봇의 얼굴과 몸이 위를 향해 있음.",
        "built_space": "쓰레기 수거선 갑판, 배경의 바다와 난파선, 주변의 쓰레기 더미 등 공간적 요소가 잘 구현됨.",
        "entities": "찰리의 상반신(모자, 코트, 얼굴, 가슴 장갑)은 레퍼런스와 일치하나, 하반신은 위치 레퍼런스에 있던 흰색 원통형 로봇 다리가 그대로 섞여서 나타남.",
        "hard_violations": [],
        "physics": "쓰레기 더미 위에 널브러져 있으며, 모든 신체 부위가 바닥의 지지를 받고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "모자·코트·흰 얼굴은 인물 참조에 가깝지만, 판독 가능한 가슴 글자가 무문자 지시를 위반하고 앞쪽 발이 화면 아래에서 잘려 전신 구도를 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "쓰레기에 받쳐 누운 전신과 몸을 감싼 그물, 지정된 선박 갑판을 충실히 보여주지만, 참조의 모자·코트가 없고 짧은 다리의 고릴라형 체형 재현은 부족하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "머리는 화면 오른쪽 뒤쪽, 발은 왼쪽과 아래쪽 전경으로 향한다. 흰 얼굴은 위쪽과 카메라 쪽으로 기울어 있으며 눈이 향하는 특정 대상은 없다. 요구된 응시 대상이나 무기 조준은 없는 장면이다.",
        "built_space": "녹슨 강철 불워크와 그 사이로 드러난 젖은 갑판이 보인다. 가장자리에 기둥 여섯 개, 뒤쪽 중앙에 케이블 윈치 한 대, 오른쪽에 크레인 한 대와 직사각형 갑판 함체 한 개가 보인다. 왼쪽 전경의 개방형 금속 구획도 장소 참조와 대응한다. 로봇은 난간 안쪽 쓰레기 더미에 놓여 있다. 다만 아래쪽으로 뻗은 발 끝이 프레임에 잘려 전신이 완결되지 않는다.",
        "entities": "로봇 한 개체만 있고 추가 인물은 없다. 갈색 모자, 올리브색 코트와 벨트, 베이지색 어깨·팔 장갑, 흰 기계식 얼굴은 인물 참조와 가깝다. 하체는 상체보다 훨씬 희고 심하게 부식되어 있으며, 길게 뻗은 다리는 짧은 다리라는 설정과 차이가 있다. 가슴 옷깃의 'UBIK' 글자가 판독된다. 몸과 하체에 어망이 걸려 있고 주변에는 플라스틱 상자, 부표, 타이어와 파편이 있다. 바다에는 헬리콥터 잔해와 두 동강 난 선박이 보인다.",
        "hard_violations": [
         "가슴 옷깃에 판독 가능한 'UBIK' 표기가 노출되어, 읽을 수 있는 글자와 로고를 금지한 최종 지시를 위반한다."
        ],
        "physics": "골반과 다리는 쓰레기와 갑판에 얹혀 있고 양팔과 손도 주변 파편에 내려앉아 있다. 상체는 뒤쪽 쓰레기 더미에 기대어 비스듬히 높아져 있으며 머리는 목 연결부 위에 남아 있어 축 늘어진 인상은 약하다. 다만 명백하게 지지 없이 공중에 뜬 부위는 확인되지 않는다. 그물은 몸의 굴곡을 따라 걸리고 나머지는 쓰레기 위에 처져 있다."
       },
       {
        "label": "B",
        "direction": "머리는 화면 오른쪽 뒤쪽, 양발은 왼쪽과 아래쪽 전경으로 뻗어 있다. 얼굴은 하늘 쪽으로 돌아가 있고 특정 대상을 바라보는 행동은 없다. 팔과 손은 몸 양옆 쓰레기 쪽으로 떨어져 있어 무언가를 가리키거나 제시하지 않는다.",
        "built_space": "녹슨 불워크 안쪽의 열린 수거선 갑판이며, 쓰레기 사이와 머리 뒤쪽으로 젖은 바닥이 비스듬히 드러난다. 가장자리 기둥 여섯 개, 뒤쪽 중앙 윈치 한 대, 오른쪽 크레인 한 대와 직사각형 함체 한 개, 왼쪽 전경의 개방형 금속 구획이 장소 참조와 대응한다. 로봇은 이 설비들 앞의 쓰레기 더미에 누워 있다. 양발을 포함한 전신이 프레임 안에 들어오며, 일부 가장자리만 전경 쓰레기에 자연스럽게 가려진다.",
        "entities": "추가 인물 없이 고철 로봇 한 개체가 보인다. 각진 샌드 베이지 장갑과 점·선 형태의 흰 마스크 얼굴은 설명에 부합하지만, 인물 참조의 모자·코트·벨트가 없고 얼굴 및 흉부 설계도 다르다. 다리는 요구된 짧은 비율보다 길다. 표면에 마모는 있으나 장소 참조 속 고철보다 부식이 적다. 굵은 그물이 가슴부터 다리까지 교차해 감싸며 주변에는 상자, 부표, 타이어와 폐기물이 쌓여 있다. 바다의 헬리콥터 잔해와 반으로 부서진 선박도 보인다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반은 쓰레기 더미에, 다리와 발은 전경의 폐기물과 갑판에 받쳐져 있다. 화면 오른쪽 손은 파편 위에 떨어져 있고 반대쪽 팔도 몸 옆 더미에 놓여 있다. 머리는 뒤쪽 더미에 낮게 기대어 있으며, 몸을 능동적으로 들어 올리는 자세는 보이지 않는다. 그물은 장갑 표면에 걸쳐지고 낮은 부분에서는 쓰레기 위로 처진다. 지지 없이 떠 있는 물체나 신체 부위는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "모자·코트·흰 얼굴은 인물 참조에 가깝지만, 판독 가능한 가슴 글자가 무문자 지시를 위반하고 앞쪽 발이 화면 아래에서 잘려 전신 구도를 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "쓰레기에 받쳐 누운 전신과 몸을 감싼 그물, 지정된 선박 갑판을 충실히 보여주지만, 참조의 모자·코트가 없고 짧은 다리의 고릴라형 체형 재현은 부족하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "머리는 화면 오른쪽 뒤쪽, 발은 왼쪽과 아래쪽 전경으로 향한다. 흰 얼굴은 위쪽과 카메라 쪽으로 기울어 있으며 눈이 향하는 특정 대상은 없다. 요구된 응시 대상이나 무기 조준은 없는 장면이다.",
        "built_space": "녹슨 강철 불워크와 그 사이로 드러난 젖은 갑판이 보인다. 가장자리에 기둥 여섯 개, 뒤쪽 중앙에 케이블 윈치 한 대, 오른쪽에 크레인 한 대와 직사각형 갑판 함체 한 개가 보인다. 왼쪽 전경의 개방형 금속 구획도 장소 참조와 대응한다. 로봇은 난간 안쪽 쓰레기 더미에 놓여 있다. 다만 아래쪽으로 뻗은 발 끝이 프레임에 잘려 전신이 완결되지 않는다.",
        "entities": "로봇 한 개체만 있고 추가 인물은 없다. 갈색 모자, 올리브색 코트와 벨트, 베이지색 어깨·팔 장갑, 흰 기계식 얼굴은 인물 참조와 가깝다. 하체는 상체보다 훨씬 희고 심하게 부식되어 있으며, 길게 뻗은 다리는 짧은 다리라는 설정과 차이가 있다. 가슴 옷깃의 'UBIK' 글자가 판독된다. 몸과 하체에 어망이 걸려 있고 주변에는 플라스틱 상자, 부표, 타이어와 파편이 있다. 바다에는 헬리콥터 잔해와 두 동강 난 선박이 보인다.",
        "hard_violations": [
         "가슴 옷깃에 판독 가능한 'UBIK' 표기가 노출되어, 읽을 수 있는 글자와 로고를 금지한 최종 지시를 위반한다."
        ],
        "physics": "골반과 다리는 쓰레기와 갑판에 얹혀 있고 양팔과 손도 주변 파편에 내려앉아 있다. 상체는 뒤쪽 쓰레기 더미에 기대어 비스듬히 높아져 있으며 머리는 목 연결부 위에 남아 있어 축 늘어진 인상은 약하다. 다만 명백하게 지지 없이 공중에 뜬 부위는 확인되지 않는다. 그물은 몸의 굴곡을 따라 걸리고 나머지는 쓰레기 위에 처져 있다."
       },
       {
        "label": "A",
        "direction": "머리는 화면 오른쪽 뒤쪽, 양발은 왼쪽과 아래쪽 전경으로 뻗어 있다. 얼굴은 하늘 쪽으로 돌아가 있고 특정 대상을 바라보는 행동은 없다. 팔과 손은 몸 양옆 쓰레기 쪽으로 떨어져 있어 무언가를 가리키거나 제시하지 않는다.",
        "built_space": "녹슨 불워크 안쪽의 열린 수거선 갑판이며, 쓰레기 사이와 머리 뒤쪽으로 젖은 바닥이 비스듬히 드러난다. 가장자리 기둥 여섯 개, 뒤쪽 중앙 윈치 한 대, 오른쪽 크레인 한 대와 직사각형 함체 한 개, 왼쪽 전경의 개방형 금속 구획이 장소 참조와 대응한다. 로봇은 이 설비들 앞의 쓰레기 더미에 누워 있다. 양발을 포함한 전신이 프레임 안에 들어오며, 일부 가장자리만 전경 쓰레기에 자연스럽게 가려진다.",
        "entities": "추가 인물 없이 고철 로봇 한 개체가 보인다. 각진 샌드 베이지 장갑과 점·선 형태의 흰 마스크 얼굴은 설명에 부합하지만, 인물 참조의 모자·코트·벨트가 없고 얼굴 및 흉부 설계도 다르다. 다리는 요구된 짧은 비율보다 길다. 표면에 마모는 있으나 장소 참조 속 고철보다 부식이 적다. 굵은 그물이 가슴부터 다리까지 교차해 감싸며 주변에는 상자, 부표, 타이어와 폐기물이 쌓여 있다. 바다의 헬리콥터 잔해와 반으로 부서진 선박도 보인다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반은 쓰레기 더미에, 다리와 발은 전경의 폐기물과 갑판에 받쳐져 있다. 화면 오른쪽 손은 파편 위에 떨어져 있고 반대쪽 팔도 몸 옆 더미에 놓여 있다. 머리는 뒤쪽 더미에 낮게 기대어 있으며, 몸을 능동적으로 들어 올리는 자세는 보이지 않는다. 그물은 장갑 표면에 걸쳐지고 낮은 부분에서는 쓰레기 위로 처진다. 지지 없이 떠 있는 물체나 신체 부위는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.125
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 캐릭터(찰리)의 정체성 및 복장 완전 누락"
    ],
    "B": [
     "[gpt-high] 가슴 옷깃에 판독 가능한 'UBIK' 표기가 노출되어, 읽을 수 있는 글자와 로고를 금지한 최종 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "찰리의 핵심 외형(중절모, 코트, 마스크)과 공간적 배경을 잘 구현했으나, 하반신이 레퍼런스의 다른 로봇 다리와 섞여 묘사된 점이 아쉽습니다.  ★위반: [gpt-high] 가슴 옷깃에 판독 가능한 'UBIK' 표기가 노출되어, 읽을 수 있는 글자와 로고를 금지한 최종 지시를 위반한다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "공간적 배경과 포즈는 적절하나, 찰리의 지정된 복장과 얼굴 형태가 전혀 반영되지 않은 다른 로봇이 등장하여 캐릭터 일치도가 매우 낮습니다.  ★위반: [gemini-pro] 지정된 캐릭터(찰리)의 정체성 및 복장 완전 누락"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_collection_deck_b9e7b9.png",
    "asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-7eac-701e-ba95-02e6b1d777f5",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S2sh3__bgfirst_bg.png",
   "bg_asset_id": "22001d0d-2139-4c65-801f-8fe7b7978313",
   "bg_record_key": "S2sh3::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "collection_deck",
   "groupbg_asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S2sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T15:56:31.543705+00:00",
  "fingerprint": "b2aa6d2f37f4c4e69f0e6d1a3aac93eda577c1daddb0b3c0b2dad96365ce6905",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S2sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S2sh3_sel.png",
  "source_sha256": "a86e937414996c8ea61a962f6b58c1260160a5a2bb61078c13b8769e63a247a7",
  "file": "S2sh3_cine.png",
  "staged_sha256": "578811e74f1d62b8e15accfd4b4d59c9f27f76b72efa5d6e468cede9cabc1f73",
  "latency_ms": 10192
 },
 "S3sh2::confined_fp_apt": {
  "applies": true,
  "reason_ko": "덤프트럭 운전석 내부라는 제한된 통제 공간에서 운전대와 운전자의 위치 및 방향 관계가 정확하게 묘사되어야 하므로 레이아웃 가이드가 필요합니다.",
  "input_fingerprint": "5886dc4de21d8287"
 },
 "S3sh2::signage": {
  "fp": "7203ed35b954bc2c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::a1ed7fb485b1": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_a1ed7fb485b1.png",
  "place_text": "Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield.",
  "input_fingerprint": "c8d955a7a74b3395"
 },
 "S3sh2::confined_fp": {
  "reads": {
   "controls": "A steering wheel is located at the left-side seat.",
   "mirrors": "No mirrors are depicted in the diagram.",
   "camera": "Positioned behind and between the two seats, angled diagonally forward and left to point directly at the left seat.",
   "occupants": "A person occupies the left seat (driver's seat)."
  },
  "mismatches": [],
  "scene_description_en": "The camera views the cab interior diagonally, looking forward and left from a point between the seats. The driver sits in the left-hand seat, appearing in the center-left of the frame from a rear-oblique angle. A steering wheel sits directly in front of the driver on the left side of the screen. The windshield stretches across the background space ahead. No mirrors or reflective surfaces are visible from this camera position.",
  "fixed": false,
  "input_fingerprint": "a5d7961103c50693"
 },
 "era_assess::96792db9ad07ec93": {
  "subjects": [],
  "subject_text": "인천 난민촌 외곽 도로와 무인점포 앞\n임시 주거지 외곽에서 도심 방향으로 이어지는 도로. 길 건너편에 무인점포의 전면과 출입구가 보인다.",
  "identity": "canonical",
  "scope_id": "L194",
  "scope_role": "location_exterior",
  "scope_sha": "194604be5aa0a228"
 },
 "S3sh2": {
  "input_fingerprint": "5ca986d0e5e7ceb3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 덤프트럭 운전석에서 고개를 떨군 채 조는 운전사의 모습.\n\nLOCATION (lock): Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dump truck cab interior (Occupied by the dozing driver while the truck travels autonomously) — Viewed diagonally from the passenger side toward the driving position; used as Close spatial enclosure around the driver's lowered head and shoulders.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light within the cab maintains natural facial detail and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The dump-truck driver is seated asleep in the driver's seat with his head drooping forward. The source does not specify the placement of his arms and legs or whether his torso rests against the seat back.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The self-driving dump truck carries a full load of scrap, including the old robot Charlie, still entangled in netting with a worn Ubik logo on its chest.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera views the cab interior diagonally, looking forward and left from a point between the seats. The driver sits in the left-hand seat, appearing in the center-left of the frame from a rear-oblique angle. A steering wheel sits directly in front of the driver on the left side of the screen. The windshield stretches across the background space ahead. No mirrors or reflective surfaces are visible from this camera position.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 덤프트럭 운전석에서 고개를 떨군 채 조는 운전사의 모습.\n\nLOCATION (lock): Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light within the cab maintains natural facial detail and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The dump-truck driver is seated asleep in the driver's seat with his head drooping forward. The source does not specify the placement of his arms and legs or whether his torso rests against the seat back.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The self-driving dump truck carries a full load of scrap, including the old robot Charlie, still entangled in netting with a worn Ubik logo on its chest.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera views the cab interior diagonally, looking forward and left from a point between the seats. The driver sits in the left-hand seat, appearing in the center-left of the frame from a rear-oblique angle. A steering wheel sits directly in front of the driver on the left side of the screen. The windshield stretches across the background space ahead. No mirrors or reflective surfaces are visible from this camera position.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 덤프트럭 운전석에서 고개를 떨군 채 조는 운전사의 모습.\n\nLOCATION (lock): Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light within the cab maintains natural facial detail and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The dump-truck driver is seated asleep in the driver's seat with his head drooping forward. The source does not specify the placement of his arms and legs or whether his torso rests against the seat back.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The self-driving dump truck carries a full load of scrap, including the old robot Charlie, still entangled in netting with a worn Ubik logo on its chest.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S3sh2_confinedfp.png",
     "asset_id": null,
     "role": null
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S3sh2_confinedfp.png",
     "asset_id": null,
     "role": null
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 조수석 방향에서 운전석을 향해 조준함.",
    "built_space": "운전실 내부. 운전대와 좌석이 측면 창문(사이드 미러와 가로 방향 모션 블러 관찰됨)을 향하고 있어 차량 구조가 논리에 맞지 않음.",
    "entities": "조는 상태의 한국인 남성 운전자 1명. 프롬프트 지시대로 고개를 앞으로 떨구고 있음.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 공간 연출 (운전대와 운전석이 차량 측면 창문을 향해 배치됨)"
    ],
    "physics": "운전자는 좌석에 앉아 있으며 양손은 무릎에 놓여 안정적으로 지지됨."
   },
   {
    "label": "B",
    "direction": "카메라가 조수석 방향에서 운전석을 향해 조준함.",
    "built_space": "운전실 내부. 좌핸들 차량의 대시보드, 운전대, 정면 및 측면 창문이 올바른 구조로 배치됨.",
    "entities": "수면 중인 한국인 남성 운전자 1명. 고개를 앞으로 떨구라는 지시와 다르게 뒤로 젖히고 있음.",
    "hard_violations": [],
    "physics": "운전자는 좌석에 안착해 고개를 등받이에 기대고 있으며, 양손은 다리와 팔걸이에 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "운전대와 좌석이 차량의 측면 창문을 향해 배치되는 물리적으로 불가능한 공간 연출이 발생하여 실격입니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "공간 구조와 카메라는 지시대로 정확히 구현되었으나, 지정된 자세(고개를 앞으로 떨군 채)와 달리 운전자가 고개를 뒤로 젖히고 있어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 조수석 방향에서 운전석을 향해 조준함.",
        "built_space": "운전실 내부. 운전대와 좌석이 측면 창문(사이드 미러와 가로 방향 모션 블러 관찰됨)을 향하고 있어 차량 구조가 논리에 맞지 않음.",
        "entities": "조는 상태의 한국인 남성 운전자 1명. 프롬프트 지시대로 고개를 앞으로 떨구고 있음.",
        "hard_violations": [
         "물리적으로 불가능한 공간 연출 (운전대와 운전석이 차량 측면 창문을 향해 배치됨)"
        ],
        "physics": "운전자는 좌석에 앉아 있으며 양손은 무릎에 놓여 안정적으로 지지됨."
       },
       {
        "label": "B",
        "direction": "카메라가 조수석 방향에서 운전석을 향해 조준함.",
        "built_space": "운전실 내부. 좌핸들 차량의 대시보드, 운전대, 정면 및 측면 창문이 올바른 구조로 배치됨.",
        "entities": "수면 중인 한국인 남성 운전자 1명. 고개를 앞으로 떨구라는 지시와 다르게 뒤로 젖히고 있음.",
        "hard_violations": [],
        "physics": "운전자는 좌석에 안착해 고개를 등받이에 기대고 있으며, 양손은 다리와 팔걸이에 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "운전대와 좌석이 차량의 측면 창문을 향해 배치되는 물리적으로 불가능한 공간 연출이 발생하여 실격입니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "공간 구조와 카메라는 지시대로 정확히 구현되었으나, 지정된 자세(고개를 앞으로 떨군 채)와 달리 운전자가 고개를 뒤로 젖히고 있어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 조수석 방향에서 운전석을 향해 조준함.",
        "built_space": "운전실 내부. 운전대와 좌석이 측면 창문(사이드 미러와 가로 방향 모션 블러 관찰됨)을 향하고 있어 차량 구조가 논리에 맞지 않음.",
        "entities": "조는 상태의 한국인 남성 운전자 1명. 프롬프트 지시대로 고개를 앞으로 떨구고 있음.",
        "hard_violations": [
         "물리적으로 불가능한 공간 연출 (운전대와 운전석이 차량 측면 창문을 향해 배치됨)"
        ],
        "physics": "운전자는 좌석에 앉아 있으며 양손은 무릎에 놓여 안정적으로 지지됨."
       },
       {
        "label": "B",
        "direction": "카메라가 조수석 방향에서 운전석을 향해 조준함.",
        "built_space": "운전실 내부. 좌핸들 차량의 대시보드, 운전대, 정면 및 측면 창문이 올바른 구조로 배치됨.",
        "entities": "수면 중인 한국인 남성 운전자 1명. 고개를 앞으로 떨구라는 지시와 다르게 뒤로 젖히고 있음.",
        "hard_violations": [],
        "physics": "운전자는 좌석에 안착해 고개를 등받이에 기대고 있으며, 양손은 다리와 팔걸이에 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "조수석 쪽 사선 구도와 운전석에서 자는 상황은 맞지만, 머리가 앞으로 떨어지지 않고 머리받침 쪽으로 기대어 있어 핵심 자세 지시와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "조수석 쪽에서 본 운전사가 고개를 앞으로 떨구고 손을 무릎에 내려놓은 모습과 낮의 주행 배경이 핵심 장면을 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "운전사는 눈을 감고 있으며 얼굴은 앞쪽 아래보다는 비스듬한 전방을 향한다. 머리는 뒤의 머리받침 쪽에 기대어 있어 요구된 전방 고개 숙임과 다르다. 운전대는 운전사 앞에 있고 카메라는 조수석 쪽에서 운전석을 사선으로 본다.",
        "built_space": "운전석 하나와 오른쪽 전경의 빈 조수석 하나, 운전대 하나, 대시보드 하나, 변속 레버 하나와 주차 브레이크 레버 하나가 보인다. 운전석 옆문과 창, 앞유리, 뒤쪽 창, 앞 기둥의 손잡이 하나가 좁은 운전실을 구성한다. 운전사는 운전대 뒤의 올바른 좌석에 앉아 있으며 평면도의 좌석 관계와 부합한다. 불가능한 반사나 중복 좌석은 보이지 않는다.",
        "entities": "중년의 동아시아계 남성 한 명이 회색 반소매 상의와 어두운 바지를 입고 있다. 한국인 운전사 설정과 시각적 충돌은 없지만 국적 자체는 외모만으로 확정할 수 없다. 낡은 대형 트럭 운전실과 낮의 외부 풍경이 보인다. 적재함의 고철·찰리·그물은 이 실내 구도 밖이므로 확인할 수 없다. 명확히 읽히는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이와 몸통은 좌석 및 등받침에 받쳐지고 머리는 머리받침 쪽으로 기울어 있다. 양손은 허벅지에 놓였으며 가까운 팔은 팔걸이에도 받쳐져 있다. 지지 없이 들린 신체나 떠 있는 물체는 없다. 잠든 자세로는 가능하지만 앞으로 고개를 떨군 자세는 아니다."
       },
       {
        "label": "B",
        "direction": "운전사의 눈은 감겨 있고 턱과 얼굴이 가슴 및 무릎 쪽으로 내려가 있다. 고개를 앞으로 떨구고 조는 방향이 명확하다. 운전대는 운전사를 향해 정상적으로 설치되어 있으며 카메라는 조수석 쪽에서 운전석을 사선으로 바라본다.",
        "built_space": "운전석 하나와 오른쪽 전경의 빈 조수석 하나, 운전대 하나, 대시보드 하나, 변속 레버 하나와 주차 브레이크 레버 하나가 보인다. 앞유리와 운전석 측면 창, 뒤쪽 창, 앞 기둥 손잡이 하나 및 상단 거울 하나가 보인다. 운전사는 운전대 뒤에 정상적으로 앉아 있고 두 좌석의 관계도 평면도에 맞는다. 머리와 어깨를 둘러싼 좁은 운전실이 유지되며 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "중년의 동아시아계 남성 운전사 한 명이 회색 긴소매 상의와 어두운 바지를 입고 있다. 한국인 설정에 어긋나는 외형은 없으며 별도의 인물 외형 참조는 제공되지 않았다. 낡은 트럭 운전실 밖으로 낮의 도로와 흐릿한 차량들이 보인다. 적재물과 찰리는 구도 밖이라 확인 대상이 아니며, 외부 간판과 실내 표식은 판독되지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이는 좌석에, 등은 등받침에 지지되고 머리는 목에서 자연스럽게 앞으로 처져 있다. 양손과 아래팔은 무릎 및 허벅지에 내려놓여 있어 잠든 사람이 팔을 들고 있는 문제가 없다. 운전대와 레버는 차량에 고정되어 있다. 외부 차량과 도로의 흐림은 주행 중이라는 설정을 뒷받침하며, 운전사가 운전대를 잡지 않은 상태도 자율주행 설정과 맞는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "조수석 쪽 사선 구도와 운전석에서 자는 상황은 맞지만, 머리가 앞으로 떨어지지 않고 머리받침 쪽으로 기대어 있어 핵심 자세 지시와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "조수석 쪽에서 본 운전사가 고개를 앞으로 떨구고 손을 무릎에 내려놓은 모습과 낮의 주행 배경이 핵심 장면을 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "운전사는 눈을 감고 있으며 얼굴은 앞쪽 아래보다는 비스듬한 전방을 향한다. 머리는 뒤의 머리받침 쪽에 기대어 있어 요구된 전방 고개 숙임과 다르다. 운전대는 운전사 앞에 있고 카메라는 조수석 쪽에서 운전석을 사선으로 본다.",
        "built_space": "운전석 하나와 오른쪽 전경의 빈 조수석 하나, 운전대 하나, 대시보드 하나, 변속 레버 하나와 주차 브레이크 레버 하나가 보인다. 운전석 옆문과 창, 앞유리, 뒤쪽 창, 앞 기둥의 손잡이 하나가 좁은 운전실을 구성한다. 운전사는 운전대 뒤의 올바른 좌석에 앉아 있으며 평면도의 좌석 관계와 부합한다. 불가능한 반사나 중복 좌석은 보이지 않는다.",
        "entities": "중년의 동아시아계 남성 한 명이 회색 반소매 상의와 어두운 바지를 입고 있다. 한국인 운전사 설정과 시각적 충돌은 없지만 국적 자체는 외모만으로 확정할 수 없다. 낡은 대형 트럭 운전실과 낮의 외부 풍경이 보인다. 적재함의 고철·찰리·그물은 이 실내 구도 밖이므로 확인할 수 없다. 명확히 읽히는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이와 몸통은 좌석 및 등받침에 받쳐지고 머리는 머리받침 쪽으로 기울어 있다. 양손은 허벅지에 놓였으며 가까운 팔은 팔걸이에도 받쳐져 있다. 지지 없이 들린 신체나 떠 있는 물체는 없다. 잠든 자세로는 가능하지만 앞으로 고개를 떨군 자세는 아니다."
       },
       {
        "label": "A",
        "direction": "운전사의 눈은 감겨 있고 턱과 얼굴이 가슴 및 무릎 쪽으로 내려가 있다. 고개를 앞으로 떨구고 조는 방향이 명확하다. 운전대는 운전사를 향해 정상적으로 설치되어 있으며 카메라는 조수석 쪽에서 운전석을 사선으로 바라본다.",
        "built_space": "운전석 하나와 오른쪽 전경의 빈 조수석 하나, 운전대 하나, 대시보드 하나, 변속 레버 하나와 주차 브레이크 레버 하나가 보인다. 앞유리와 운전석 측면 창, 뒤쪽 창, 앞 기둥 손잡이 하나 및 상단 거울 하나가 보인다. 운전사는 운전대 뒤에 정상적으로 앉아 있고 두 좌석의 관계도 평면도에 맞는다. 머리와 어깨를 둘러싼 좁은 운전실이 유지되며 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "중년의 동아시아계 남성 운전사 한 명이 회색 긴소매 상의와 어두운 바지를 입고 있다. 한국인 설정에 어긋나는 외형은 없으며 별도의 인물 외형 참조는 제공되지 않았다. 낡은 트럭 운전실 밖으로 낮의 도로와 흐릿한 차량들이 보인다. 적재물과 찰리는 구도 밖이라 확인 대상이 아니며, 외부 간판과 실내 표식은 판독되지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이는 좌석에, 등은 등받침에 지지되고 머리는 목에서 자연스럽게 앞으로 처져 있다. 양손과 아래팔은 무릎 및 허벅지에 내려놓여 있어 잠든 사람이 팔을 들고 있는 문제가 없다. 운전대와 레버는 차량에 고정되어 있다. 외부 차량과 도로의 흐림은 주행 중이라는 설정을 뒷받침하며, 운전사가 운전대를 잡지 않은 상태도 자율주행 설정과 맞는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.667
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 공간 연출 (운전대와 운전석이 차량 측면 창문을 향해 배치됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1667
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "운전대와 좌석이 차량의 측면 창문을 향해 배치되는 물리적으로 불가능한 공간 연출이 발생하여 실격입니다.  ★위반: [gemini-pro] 물리적으로 불가능한 공간 연출 (운전대와 운전석이 차량 측면 창문을 향해 배치됨)"
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "공간 구조와 카메라는 지시대로 정확히 구현되었으나, 지정된 자세(고개를 앞으로 떨군 채)와 달리 운전자가 고개를 뒤로 젖히고 있어 감점되었습니다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S3sh2_confinedfp.png",
    "asset_id": null,
    "role": null
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-8383-74e3-b293-f5f0611e1fc4",
  "confined_fp": {
   "base_key": "confinedfp::a1ed7fb485b1",
   "apt_reason": "덤프트럭 운전석 내부라는 제한된 통제 공간에서 운전대와 운전자의 위치 및 방향 관계가 정확하게 묘사되어야 하므로 레이아웃 가이드가 필요합니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S3sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T15:56:42.119379+00:00",
  "fingerprint": "47bb95cf4363cf998a7a084a6430365d408e35a40ff47d0ffd6a9bd476bc28c8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S3sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S3sh2_sel.png",
  "source_sha256": "17e17e46e202e712858d0245e147eca51f7d785551c90937f5eb4457f044000f",
  "file": "S3sh2_cine.png",
  "staged_sha256": "aff0326ecf002982a727db4ab208c050ea094bb56ca5ed02696dedc849d3e339",
  "latency_ms": 10075
 },
 "S4sh1::confined_fp_apt": {
  "applies": true,
  "reason_ko": "차량 내부의 앞좌석이라는 통제 장치가 있는 제한된 공간을 배경으로 하며, 페드로가 운전석에 앉고 현우가 조수석에 앉아 상호작용하는 정확한 위치 관계가 이야기 전달에 필수적이므로 평면도 레이아웃 보조가 필요합니다.",
  "input_fingerprint": "70ce91fadbf89495"
 },
 "S4sh1::signage": {
  "fp": "f07dec417db507b8",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "구겨진 도면"
   }
  ],
  "dropped": []
 },
 "confinedfp::510c33afc573": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_510c33afc573.png",
  "place_text": "Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows.",
  "input_fingerprint": "f9b6d15cd2a96a4b"
 },
 "S4sh1::confined_fp": {
  "reads": {
   "controls": "The steering wheel is in front of the left seat. A center console sits between the left and right seats.",
   "mirrors": "A rearview mirror is mounted at the top center, facing the rear of the cabin.",
   "camera": "The camera is located behind the left seat, angled diagonally forward and to the right.",
   "occupants": "페드로 (Pedro) occupies the left seat (driver's side). 현우 (Hyunwoo) occupies the right seat (passenger's side)."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned in the rear left of the car's interior, looking diagonally forward and right from behind the driver's seat. In the left foreground, the back of Pedro's head and shoulder are visible as he occupies the driver's seat. To his right, in the center midground, sits the center console where Pedro's hand reaches toward a crumpled floor plan. In the right midground, Hyunwoo occupies the passenger seat, angled slightly to look down at the console. The steering wheel is located in front of Pedro on the left. High in the upper center background, the rearview mirror faces rearward, positioned to physically reflect the rear of the cabin and the occupants' faces from the camera's perspective.",
  "fixed": true,
  "input_fingerprint": "df64102d5ba7dda3"
 },
 "era_assess::947f776320e3b609": {
  "subjects": [],
  "subject_text": "페드로의 자동차 내부\n운전석과 조수석이 나란한 자동차 앞좌석 공간. 운전대와 대시보드, 노트북과 구겨진 도면이 있으며 창으로 낮빛이 들어온다.",
  "identity": "canonical",
  "scope_id": "L150",
  "scope_role": "location_interior",
  "scope_sha": "5d50a17e313ca4db"
 },
 "S4sh1": {
  "input_fingerprint": "1e67f81116469cc1",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 정차해 있는 차 안, 구겨진 도면을 가리키는 페드로의 손과 이를 미간을 찌푸린 채 내려다보는 현우의 모습.\n\nLOCATION (lock): Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Crumpled drawing (Crumpled and poorly drawn, being examined and indicated) — The marked face tilts upward toward the occupants and remains visible to the camera from above; used as Shared lower-center attention anchor linking the hand and 현우's reaction; Stationary car interior (Occupied by 페드로 in the driving position and 현우 beside him) — Viewed diagonally forward from behind the driving position; used as Maintains the compressed relationship between the two occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light inside the car provides clear facial and paper detail with understated contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car is parked in the upscale residential district, and the house plan is crumpled and badly drawn. 페드로: Pedro occupies the driver's seat and wears an earpiece. 현우: Hyunwoo is seated in the passenger position, holding and examining the crumpled plan.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the rear left of the car's interior, looking diagonally forward and right from behind the driver's seat. In the left foreground, the back of Pedro's head and shoulder are visible as he occupies the driver's seat. To his right, in the center midground, sits the center console where Pedro's hand reaches toward a crumpled floor plan. In the right midground, Hyunwoo occupies the passenger seat, angled slightly to look down at the console. The steering wheel is located in front of Pedro on the left. High in the upper center background, the rearview mirror faces rearward, positioned to physically reflect the rear of the cabin and the occupants' faces from the camera's perspective.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 정차해 있는 차 안, 구겨진 도면을 가리키는 페드로의 손과 이를 미간을 찌푸린 채 내려다보는 현우의 모습.\n\nLOCATION (lock): Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light inside the car provides clear facial and paper detail with understated contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car is parked in the upscale residential district, and the house plan is crumpled and badly drawn. 페드로: Pedro occupies the driver's seat and wears an earpiece. 현우: Hyunwoo is seated in the passenger position, holding and examining the crumpled plan.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the rear left of the car's interior, looking diagonally forward and right from behind the driver's seat. In the left foreground, the back of Pedro's head and shoulder are visible as he occupies the driver's seat. To his right, in the center midground, sits the center console where Pedro's hand reaches toward a crumpled floor plan. In the right midground, Hyunwoo occupies the passenger seat, angled slightly to look down at the console. The steering wheel is located in front of Pedro on the left. High in the upper center background, the rearview mirror faces rearward, positioned to physically reflect the rear of the cabin and the occupants' faces from the camera's perspective.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 정차해 있는 차 안, 구겨진 도면을 가리키는 페드로의 손과 이를 미간을 찌푸린 채 내려다보는 현우의 모습.\n\nLOCATION (lock): Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light inside the car provides clear facial and paper detail with understated contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car is parked in the upscale residential district, and the house plan is crumpled and badly drawn. 페드로: Pedro occupies the driver's seat and wears an earpiece. 현우: Hyunwoo is seated in the passenger position, holding and examining the crumpled plan.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S4sh1_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "페드로",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S4sh1_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "페드로",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "페드로의 오른손 검지가 구겨진 도면을 가리키고 있으며, 현우의 시선은 그 도면을 향해 아래로 내려다봄.",
    "built_space": "자동차 내부 운전석 뒤에서 조수석을 향하는 구도. 전면에 주택가 풍경이 보임. 룸미러에 얼굴이 반사되어 있으나 카메라와 인물의 각도상 불가능한 뷰임.",
    "entities": "페드로(비니, 이어피스, 후드티 착용)와 현우(회색 셔츠, 미간을 찌푸린 표정) 모두 프롬프트에 명시된 외양 및 상황과 일치함.",
    "hard_violations": [
     "[gemini-pro] 페드로의 고개는 우측을 향하고 카메라가 운전석 뒤에 위치함에도 룸미러에는 정면을 응시하는 얼굴이 반사되어 있어 광학적/물리적으로 불가능함"
    ],
    "physics": "현우의 양손이 도면의 양옆을 안정적으로 쥐고 있으며, 페드로의 오른팔은 허공을 가로질러 손가락으로 도면을 지탱 없이 가리킴."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 아래의 도면을 향하지만, 페드로의 얼굴은 가려져 시선을 확인할 수 없으며 가리키는 손동작도 없음.",
    "built_space": "운전석 헤드레스트 뒤쪽에서 조수석을 향하는 구도. 차량 내부 구조와 창밖 주택가 배경은 명시된 공간과 일치함.",
    "entities": "현우는 찌푸린 표정으로 지시와 일치하나, 페드로는 비니를 쓰지 않았으며 뒷모습만 프레임에 잡혀 식별이 어려움.",
    "hard_violations": [],
    "physics": "도면은 페드로의 것으로 추정되는 오른손에 의해서만 들려 있으며, 현우의 손은 화면에 나타나지 않아 도면을 전혀 지지하지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 구도와 인물의 행동(도면을 쥔 현우와 이를 가리키는 페드로)을 매우 정확하게 구현했으나 룸미러 반사 오류가 존재함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우가 도면을 들고 있어야 한다는 핵심 지시와 페드로의 손가락 지시 동작이 누락되었으며, 인물 외양(비니 등)도 일치하지 않음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 오른손 검지가 구겨진 도면을 가리키고 있으며, 현우의 시선은 그 도면을 향해 아래로 내려다봄.",
        "built_space": "자동차 내부 운전석 뒤에서 조수석을 향하는 구도. 전면에 주택가 풍경이 보임. 룸미러에 얼굴이 반사되어 있으나 카메라와 인물의 각도상 불가능한 뷰임.",
        "entities": "페드로(비니, 이어피스, 후드티 착용)와 현우(회색 셔츠, 미간을 찌푸린 표정) 모두 프롬프트에 명시된 외양 및 상황과 일치함.",
        "hard_violations": [
         "페드로의 고개는 우측을 향하고 카메라가 운전석 뒤에 위치함에도 룸미러에는 정면을 응시하는 얼굴이 반사되어 있어 광학적/물리적으로 불가능함"
        ],
        "physics": "현우의 양손이 도면의 양옆을 안정적으로 쥐고 있으며, 페드로의 오른팔은 허공을 가로질러 손가락으로 도면을 지탱 없이 가리킴."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 아래의 도면을 향하지만, 페드로의 얼굴은 가려져 시선을 확인할 수 없으며 가리키는 손동작도 없음.",
        "built_space": "운전석 헤드레스트 뒤쪽에서 조수석을 향하는 구도. 차량 내부 구조와 창밖 주택가 배경은 명시된 공간과 일치함.",
        "entities": "현우는 찌푸린 표정으로 지시와 일치하나, 페드로는 비니를 쓰지 않았으며 뒷모습만 프레임에 잡혀 식별이 어려움.",
        "hard_violations": [],
        "physics": "도면은 페드로의 것으로 추정되는 오른손에 의해서만 들려 있으며, 현우의 손은 화면에 나타나지 않아 도면을 전혀 지지하지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 구도와 인물의 행동(도면을 쥔 현우와 이를 가리키는 페드로)을 매우 정확하게 구현했으나 룸미러 반사 오류가 존재함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우가 도면을 들고 있어야 한다는 핵심 지시와 페드로의 손가락 지시 동작이 누락되었으며, 인물 외양(비니 등)도 일치하지 않음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 오른손 검지가 구겨진 도면을 가리키고 있으며, 현우의 시선은 그 도면을 향해 아래로 내려다봄.",
        "built_space": "자동차 내부 운전석 뒤에서 조수석을 향하는 구도. 전면에 주택가 풍경이 보임. 룸미러에 얼굴이 반사되어 있으나 카메라와 인물의 각도상 불가능한 뷰임.",
        "entities": "페드로(비니, 이어피스, 후드티 착용)와 현우(회색 셔츠, 미간을 찌푸린 표정) 모두 프롬프트에 명시된 외양 및 상황과 일치함.",
        "hard_violations": [
         "페드로의 고개는 우측을 향하고 카메라가 운전석 뒤에 위치함에도 룸미러에는 정면을 응시하는 얼굴이 반사되어 있어 광학적/물리적으로 불가능함"
        ],
        "physics": "현우의 양손이 도면의 양옆을 안정적으로 쥐고 있으며, 페드로의 오른팔은 허공을 가로질러 손가락으로 도면을 지탱 없이 가리킴."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 아래의 도면을 향하지만, 페드로의 얼굴은 가려져 시선을 확인할 수 없으며 가리키는 손동작도 없음.",
        "built_space": "운전석 헤드레스트 뒤쪽에서 조수석을 향하는 구도. 차량 내부 구조와 창밖 주택가 배경은 명시된 공간과 일치함.",
        "entities": "현우는 찌푸린 표정으로 지시와 일치하나, 페드로는 비니를 쓰지 않았으며 뒷모습만 프레임에 잡혀 식별이 어려움.",
        "hard_violations": [],
        "physics": "도면은 페드로의 것으로 추정되는 오른손에 의해서만 들려 있으며, 현우의 손은 화면에 나타나지 않아 도면을 전혀 지지하지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "현우의 찌푸린 표정과 후방 시점은 맞지만, 페드로의 손이 도면을 가리키기보다 붙잡고 있고 현우의 도면 파지 관계도 불명확하며 좌석 등받이가 장면을 과도하게 가립니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "운전석 뒤의 미디엄 구도에서 페드로의 지시 손가락, 현우의 도면 파지와 내려다보는 시선을 명확히 연결하며, 비니와 이어피스도 참조에 부합하지만 미간의 찌푸림은 약합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 상체를 안쪽과 뒤쪽으로 돌려 아래의 종이 쪽을 보며 미간을 뚜렷하게 찌푸린다. 페드로도 오른쪽의 현우와 종이 쪽으로 고개를 돌렸다. 다만 페드로의 보이는 손은 도면 위쪽 가장자리를 붙잡는 모양이며, 특정 도면 선을 가리키는 손가락은 명확하지 않다. 도면의 그려진 면은 위로 펼쳐져 카메라에서도 보인다.",
        "built_space": "앞좌석 등받이와 머리받침이 각각 두 개, 왼쪽 운전대 한 개, 두 좌석 사이 콘솔 한 개가 보인다. 페드로는 운전대 앞 좌석에, 현우는 오른쪽 조수석에 있다. 카메라는 운전석 뒤에서 대각선 전방을 보지만 등받이들이 화면의 큰 부분을 차지해 두 사람과 손의 관계를 가린다. 실내 후사경은 보이지 않으며, 창밖에는 낮의 주택가가 보인다.",
        "entities": "등장인물은 두 명뿐이다. 현우는 앳된 동아시아계 남성의 외형, 헝클어진 검은 머리와 회색 셔츠로 참조에 대체로 부합한다. 페드로는 짙은 머리와 어두운 상의를 입고 귀에 장치를 착용하지만 참조의 검은 비니는 없다. 얼굴이 부분적으로만 보여 참조 인상과의 일치를 충분히 확인하기 어렵다. 종이는 실제로 구겨져 있고 거친 건축 도면 선이 그려져 있다. 읽을 수 있는 문구는 식별되지 않는다. 현우 자신의 손이 도면을 확실히 쥐고 있는지는 가려져 불명확하다.",
        "hard_violations": [],
        "physics": "두 사람은 각자의 앞좌석에 앉아 있으며 현우의 몸통 회전은 앉은 상태에서 가능한 동작이다. 종이는 페드로의 손이 위쪽 가장자리를 잡고 있고 아래쪽은 중앙 콘솔 부근에 걸쳐 있어 무지지 부유로 보이지 않는다. 손과 팔의 연결에 명백한 해부학적 불가능은 없지만, 종이가 좌석 뒤쪽으로 많이 내려와 현우가 직접 들고 검토하는 동작은 약하게 전달된다."
       },
       {
        "label": "B",
        "direction": "페드로의 뻗은 검지는 구겨진 도면 중앙 위쪽의 방 구획선을 직접 가리킨다. 현우는 고개와 눈을 아래로 향해 그 도면을 살펴본다. 페드로의 고개도 종이 쪽으로 돌아가 있다. 종이의 그려진 면은 두 사람 사이에서 위쪽으로 기울어져 있어 현우가 읽는 동시에 뒤쪽 카메라에도 보이는 배치다. 현우의 표정은 집중되어 있으나 미간의 주름은 A보다 약하다.",
        "built_space": "왼쪽 운전석과 오른쪽 조수석의 등받이·머리받침이 각각 하나씩 보이고, 운전대 한 개, 중앙 콘솔 한 개, 실내 후사경 한 개가 배치되어 있다. 페드로는 운전석에, 현우는 그 옆 조수석에 앉아 각자의 몸을 중앙으로 돌렸다. 운전석 뒤에서 대각선 전방을 보는 미디엄 구도가 손과 도면, 현우의 얼굴을 함께 담는다. 후사경 왼쪽에는 운전자의 눈과 이마 일부가 보이며, 이 화면만으로 불가능한 반사라고 단정할 근거는 없다. 창밖의 밝은 단독주택가도 장소 조건과 맞는다.",
        "entities": "두 명의 인물만 보인다. 현우는 젊은 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 회색 셔츠와 올리브색 바지로 참조에 대체로 부합한다. 페드로는 젊은 남성의 부분 얼굴과 짙은 머리, 검은 비니·후드 상의 및 이어피스를 보여 참조와 착용 조건을 잘 따른다. 구겨진 종이에는 서툰 평면도 선이 있고, 현우의 회색 소매에서 이어진 손들이 종이 가장자리를 잡는다. 식별 가능한 글자나 별도의 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 몸은 각자의 좌석에 지지되고 등은 각 좌석 등받이 쪽에 놓여 있다. 현우의 양손이 종이의 오른쪽과 아래쪽 가장자리를 받쳐 종이의 기울기를 유지한다. 페드로의 팔은 어깨에서 자연스럽게 뻗어 검지로 도면을 지시한다. 종이의 주름과 처짐도 손으로 든 상태에 부합하며 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우의 찌푸린 표정과 후방 시점은 맞지만, 페드로의 손이 도면을 가리키기보다 붙잡고 있고 현우의 도면 파지 관계도 불명확하며 좌석 등받이가 장면을 과도하게 가립니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "운전석 뒤의 미디엄 구도에서 페드로의 지시 손가락, 현우의 도면 파지와 내려다보는 시선을 명확히 연결하며, 비니와 이어피스도 참조에 부합하지만 미간의 찌푸림은 약합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 상체를 안쪽과 뒤쪽으로 돌려 아래의 종이 쪽을 보며 미간을 뚜렷하게 찌푸린다. 페드로도 오른쪽의 현우와 종이 쪽으로 고개를 돌렸다. 다만 페드로의 보이는 손은 도면 위쪽 가장자리를 붙잡는 모양이며, 특정 도면 선을 가리키는 손가락은 명확하지 않다. 도면의 그려진 면은 위로 펼쳐져 카메라에서도 보인다.",
        "built_space": "앞좌석 등받이와 머리받침이 각각 두 개, 왼쪽 운전대 한 개, 두 좌석 사이 콘솔 한 개가 보인다. 페드로는 운전대 앞 좌석에, 현우는 오른쪽 조수석에 있다. 카메라는 운전석 뒤에서 대각선 전방을 보지만 등받이들이 화면의 큰 부분을 차지해 두 사람과 손의 관계를 가린다. 실내 후사경은 보이지 않으며, 창밖에는 낮의 주택가가 보인다.",
        "entities": "등장인물은 두 명뿐이다. 현우는 앳된 동아시아계 남성의 외형, 헝클어진 검은 머리와 회색 셔츠로 참조에 대체로 부합한다. 페드로는 짙은 머리와 어두운 상의를 입고 귀에 장치를 착용하지만 참조의 검은 비니는 없다. 얼굴이 부분적으로만 보여 참조 인상과의 일치를 충분히 확인하기 어렵다. 종이는 실제로 구겨져 있고 거친 건축 도면 선이 그려져 있다. 읽을 수 있는 문구는 식별되지 않는다. 현우 자신의 손이 도면을 확실히 쥐고 있는지는 가려져 불명확하다.",
        "hard_violations": [],
        "physics": "두 사람은 각자의 앞좌석에 앉아 있으며 현우의 몸통 회전은 앉은 상태에서 가능한 동작이다. 종이는 페드로의 손이 위쪽 가장자리를 잡고 있고 아래쪽은 중앙 콘솔 부근에 걸쳐 있어 무지지 부유로 보이지 않는다. 손과 팔의 연결에 명백한 해부학적 불가능은 없지만, 종이가 좌석 뒤쪽으로 많이 내려와 현우가 직접 들고 검토하는 동작은 약하게 전달된다."
       },
       {
        "label": "A",
        "direction": "페드로의 뻗은 검지는 구겨진 도면 중앙 위쪽의 방 구획선을 직접 가리킨다. 현우는 고개와 눈을 아래로 향해 그 도면을 살펴본다. 페드로의 고개도 종이 쪽으로 돌아가 있다. 종이의 그려진 면은 두 사람 사이에서 위쪽으로 기울어져 있어 현우가 읽는 동시에 뒤쪽 카메라에도 보이는 배치다. 현우의 표정은 집중되어 있으나 미간의 주름은 A보다 약하다.",
        "built_space": "왼쪽 운전석과 오른쪽 조수석의 등받이·머리받침이 각각 하나씩 보이고, 운전대 한 개, 중앙 콘솔 한 개, 실내 후사경 한 개가 배치되어 있다. 페드로는 운전석에, 현우는 그 옆 조수석에 앉아 각자의 몸을 중앙으로 돌렸다. 운전석 뒤에서 대각선 전방을 보는 미디엄 구도가 손과 도면, 현우의 얼굴을 함께 담는다. 후사경 왼쪽에는 운전자의 눈과 이마 일부가 보이며, 이 화면만으로 불가능한 반사라고 단정할 근거는 없다. 창밖의 밝은 단독주택가도 장소 조건과 맞는다.",
        "entities": "두 명의 인물만 보인다. 현우는 젊은 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 회색 셔츠와 올리브색 바지로 참조에 대체로 부합한다. 페드로는 젊은 남성의 부분 얼굴과 짙은 머리, 검은 비니·후드 상의 및 이어피스를 보여 참조와 착용 조건을 잘 따른다. 구겨진 종이에는 서툰 평면도 선이 있고, 현우의 회색 소매에서 이어진 손들이 종이 가장자리를 잡는다. 식별 가능한 글자나 별도의 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 몸은 각자의 좌석에 지지되고 등은 각 좌석 등받이 쪽에 놓여 있다. 현우의 양손이 종이의 오른쪽과 아래쪽 가장자리를 받쳐 종이의 기울기를 유지한다. 페드로의 팔은 어깨에서 자연스럽게 뻗어 검지로 도면을 지시한다. 종이의 주름과 처짐도 손으로 든 상태에 부합하며 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.167
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "A": [
     "[gemini-pro] 페드로의 고개는 우측을 향하고 카메라가 운전석 뒤에 위치함에도 룸미러에는 정면을 응시하는 얼굴이 반사되어 있어 광학적/물리적으로 불가능함"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지시된 구도와 인물의 행동(도면을 쥔 현우와 이를 가리키는 페드로)을 매우 정확하게 구현했으나 룸미러 반사 오류가 존재함.  ★위반: [gemini-pro] 페드로의 고개는 우측을 향하고 카메라가 운전석 뒤에 위치함에도 룸미러에는 정면을 응시하는 얼굴이 반사되어 있어 광학적/물리적으로 불가능함"
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "현우가 도면을 들고 있어야 한다는 핵심 지시와 페드로의 손가락 지시 동작이 누락되었으며, 인물 외양(비니 등)도 일치하지 않음."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S4sh1_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "페드로",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-86ba-71d3-a45c-975a18b5e8dd",
  "confined_fp": {
   "base_key": "confinedfp::510c33afc573",
   "apt_reason": "차량 내부의 앞좌석이라는 통제 장치가 있는 제한된 공간을 배경으로 하며, 페드로가 운전석에 앉고 현우가 조수석에 앉아 상호작용하는 정확한 위치 관계가 이야기 전달에 필수적이므로 평면도 레이아웃 보조가 필요합니다.",
   "fixed": true,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S4sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T15:58:16.763263+00:00",
  "fingerprint": "94aaa94d0fa4cd66272f6244468ac6191f6ebed22ff409e7f003f4c8a492d3b7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S4sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S4sh1_sel.png",
  "source_sha256": "6dab13d7ee7a1b25e03fb2cb2493db56054a47782e3d4c3e13c8466c85729752",
  "file": "S4sh1_cine.png",
  "staged_sha256": "a3326d9c35aaa4242713439cde865ca58b23c084a58f08fa672c41c31a3ab4de",
  "latency_ms": 10117
 },
 "S5sh1::signage": {
  "fp": "8d5506ac21c64cce",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::683a2594bb638b15": {
  "subjects": [],
  "subject_text": "고급주택 거실\n큰 통유리창으로 외부가 보이는 고급주택 거실. 열리는 창과 현관 쪽 동선, 2층으로 올라가는 계단 입구가 연결된다.",
  "identity": "canonical",
  "scope_id": "L152",
  "scope_role": "location_interior",
  "scope_sha": "468a1a21e5d74ad9"
 },
 "S5sh1::bgfirst_bg": {
  "input_fingerprint": "35c60ca0064c7215",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1__bgfirst_bg.png",
  "asset_id": "33fddc9a-f643-4e6e-9f6c-5eeb879142cd",
  "input_asset_ids": [
   "dde54bdc-8795-4460-ac4c-7a559b6d6b14",
   "1ced485f-4388-40e8-9db1-c1e4432bbd3c"
  ]
 },
 "S5sh1": {
  "input_fingerprint": "98aa4496816dd7d3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The house has a full-height living-room window and an old-fashioned keyed front-door lock. The young female housekeeping robot is intact at this point. 현우: Hyunwoo has crossed the boundary wall and is keeping close to the front entrance.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The house has a full-height living-room window and an old-fashioned keyed front-door lock. The young female housekeeping robot is intact at this point. 현우: Hyunwoo has crossed the boundary wall and is keeping close to the front entrance.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The house has a full-height living-room window and an old-fashioned keyed front-door lock. The young female housekeeping robot is intact at this point. 현우: Hyunwoo has crossed the boundary wall and is keeping close to the front entrance.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1__bgfirst_bg.png",
     "asset_id": "33fddc9a-f643-4e6e-9f6c-5eeb879142cd",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S5sh1.png",
     "asset_id": "dde54bdc-8795-4460-ac4c-7a559b6d6b14",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:896577>",
     "asset_id": "d19d1741-7b2a-408b-a66b-96afe66a9293",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B03.png",
     "asset_id": "1ced485f-4388-40e8-9db1-c1e4432bbd3c",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:896577>",
     "asset_id": "d19d1741-7b2a-408b-a66b-96afe66a9293",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우는 창문 너머의 로봇을 바라보고 있으며, 로봇은 청소 중인 곳으로 시선을 향하고 있습니다.",
    "built_space": "참조된 현관과 유리창 구조가 정확히 묘사되었고, 인물들이 공간에 맞게 배치되었습니다.",
    "entities": "현우와 로봇 모두 참조 이미지의 인상, 복장, 기계 팔다리 등을 정확히 반영했습니다.",
    "hard_violations": [
     "[gpt-high] 현우가 창 옆 불투명한 외벽 뒤에 머물러 있어, 요구된 여자에 대한 직접적인 응시가 공간적으로 성립하지 않습니다."
    ],
    "physics": "두 인물 모두 바닥에 두 발을 딛고 안정적으로 지탱되어 있습니다."
   },
   {
    "label": "B",
    "direction": "현우는 고개를 돌려 창문 안의 로봇을 응시하고, 로봇은 가구를 닦으며 아래를 보고 있습니다.",
    "built_space": "기본적인 배치에는 맞으나, 현우가 위치한 벽면 모서리와 계단 구조가 심하게 왜곡되어 있습니다.",
    "entities": "현우와 로봇의 외형 및 복장은 참조 이미지의 조건을 충족합니다.",
    "hard_violations": [
     "[gemini-pro] 현우의 하체 및 다리 주변 벽면과 계단 구조가 절단된 것처럼 겹치고 깨지는 심각한 공간 콜라주/렌더링 오류",
     "[gpt-high] 현우의 눈이 불투명한 외벽 뒤에 남아 있어 창 너머 여자를 직접 응시할 수 없는 위치로 배치되었습니다."
    ],
    "physics": "바닥과 벽면 구조가 깨져 있어 현우가 물리적으로 불가능한 지형 위에 서 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 공간과 인물의 신원을 정확히 구현했으며, 물리적 오류 없이 지시문의 상황을 안정적으로 연출했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우의 역동적인 자세는 돋보이나, 그가 서 있는 하체 주변의 건축 구조가 기형적으로 깨지는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 창문 너머의 로봇을 바라보고 있으며, 로봇은 청소 중인 곳으로 시선을 향하고 있습니다.",
        "built_space": "참조된 현관과 유리창 구조가 정확히 묘사되었고, 인물들이 공간에 맞게 배치되었습니다.",
        "entities": "현우와 로봇 모두 참조 이미지의 인상, 복장, 기계 팔다리 등을 정확히 반영했습니다.",
        "hard_violations": [],
        "physics": "두 인물 모두 바닥에 두 발을 딛고 안정적으로 지탱되어 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 돌려 창문 안의 로봇을 응시하고, 로봇은 가구를 닦으며 아래를 보고 있습니다.",
        "built_space": "기본적인 배치에는 맞으나, 현우가 위치한 벽면 모서리와 계단 구조가 심하게 왜곡되어 있습니다.",
        "entities": "현우와 로봇의 외형 및 복장은 참조 이미지의 조건을 충족합니다.",
        "hard_violations": [
         "현우의 하체 및 다리 주변 벽면과 계단 구조가 절단된 것처럼 겹치고 깨지는 심각한 공간 콜라주/렌더링 오류"
        ],
        "physics": "바닥과 벽면 구조가 깨져 있어 현우가 물리적으로 불가능한 지형 위에 서 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 공간과 인물의 신원을 정확히 구현했으며, 물리적 오류 없이 지시문의 상황을 안정적으로 연출했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우의 역동적인 자세는 돋보이나, 그가 서 있는 하체 주변의 건축 구조가 기형적으로 깨지는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 창문 너머의 로봇을 바라보고 있으며, 로봇은 청소 중인 곳으로 시선을 향하고 있습니다.",
        "built_space": "참조된 현관과 유리창 구조가 정확히 묘사되었고, 인물들이 공간에 맞게 배치되었습니다.",
        "entities": "현우와 로봇 모두 참조 이미지의 인상, 복장, 기계 팔다리 등을 정확히 반영했습니다.",
        "hard_violations": [],
        "physics": "두 인물 모두 바닥에 두 발을 딛고 안정적으로 지탱되어 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 돌려 창문 안의 로봇을 응시하고, 로봇은 가구를 닦으며 아래를 보고 있습니다.",
        "built_space": "기본적인 배치에는 맞으나, 현우가 위치한 벽면 모서리와 계단 구조가 심하게 왜곡되어 있습니다.",
        "entities": "현우와 로봇의 외형 및 복장은 참조 이미지의 조건을 충족합니다.",
        "hard_violations": [
         "현우의 하체 및 다리 주변 벽면과 계단 구조가 절단된 것처럼 겹치고 깨지는 심각한 공간 콜라주/렌더링 오류"
        ],
        "physics": "바닥과 벽면 구조가 깨져 있어 현우가 물리적으로 불가능한 지형 위에 서 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "문에 밀착한 자세는 맞지만 외벽에 가로막힌 시선, 변형된 현관 공간, 불명확한 먼지 털기 동작 때문에 핵심 장면의 충실도가 낮습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "장소의 구조와 와이드 배치, 여자의 먼지 털기를 더 충실히 구현했지만 현우가 외벽 뒤에 있어 여자를 직접 응시하는 시선 관계는 성립하지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 오른쪽 창가의 여자 방향으로 돌리고 있습니다. 그러나 눈과 코가 창 옆 두꺼운 외벽의 왼쪽 뒤에 머물러 있어, 여자까지의 시선은 불투명한 벽에 막힙니다. 여자는 고개를 숙여 자신의 손과 수납장 쪽을 보며 현우를 돌아보지 않습니다.",
        "built_space": "왼쪽에 닫힌 현관문 하나와 세로형 잠금장치 하나, 그 오른쪽에 벽등 하나가 보입니다. 오른쪽에는 중앙 세로 프레임으로 나뉜 대형 거실창과 양쪽 커튼이 있습니다. 실내의 소파, 낮은 탁자, 텔레비전과 긴 수납장은 참조 장소와 유사합니다. 다만 현관문이 참조보다 전면으로 당겨져 보이고, 현우 아래 바닥과 오른쪽 계단의 연결도 달라져 원래의 깊은 현관 공간이 약해졌습니다. 나무와 맞은편 건물의 유리 반사는 가능한 종류지만 상당히 강합니다.",
        "entities": "인물은 현우와 젊은 여성형 로봇 두 명뿐입니다. 현우의 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 회색 셔츠, 카키색 작업 바지와 부츠는 참조에 부합합니다. 여자는 검은 머리와 사람 같은 얼굴, 남색 옷과 흰 앞치마, 노출된 기계 팔·다리 및 흰 신발을 갖추고 있으며 손상은 보이지 않습니다. 손끝의 밝은 물체는 보이지만 먼지털이의 손잡이와 형태는 분명하지 않습니다. 문에는 구식 열쇠 자물쇠보다는 전자식 도어록에 가까운 장치가 있습니다. 읽을 수 있는 문구는 확인되지 않습니다.",
        "hard_violations": [
         "현우의 눈이 불투명한 외벽 뒤에 남아 있어 창 너머 여자를 직접 응시할 수 없는 위치로 배치되었습니다."
        ],
        "physics": "현우는 두 부츠를 현관 바닥과 매트에 딛고 다리를 벌린 채 몸을 기울이며 문 옆에 붙어 있습니다. 과장된 자세지만 바닥 지지가 있어 부유하지 않습니다. 여자는 두 발로 실내 바닥을 딛고 허리를 숙이고 있으며, 앞으로 내민 팔과 손도 몸에 연결되어 있습니다. 손끝 물체의 정확한 파지 형태는 작아서 확정하기 어렵지만 독립적으로 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 오른쪽 거실창 쪽을 향합니다. 다만 머리가 현관 측벽의 바깥 모서리를 넘어 나오지 않아 여자까지의 직접적인 시선은 두꺼운 외벽을 통과해야 합니다. 여자는 아래쪽 소파와 먼지털이 끝을 바라보고 있으며 관찰자를 의식하지 않는 모습입니다. 먼지털이는 손에서 왼쪽 아래 소파 좌석 방향으로 뻗어 있습니다.",
        "built_space": "왼쪽의 깊게 들어간 현관에 닫힌 문 하나, 문 부착 잠금장치 하나, 벽등 하나, 호출기 하나, 천장 매입등 하나가 보입니다. 현관 바닥과 전면 계단, 문과 창 사이의 두꺼운 회색 벽, 오른쪽 대형 창의 구성이 참조와 가깝습니다. 현우는 문 바로 앞보다는 현관 오른쪽 바깥 모서리에 붙어 있습니다. 닫힌 창 너머 중간 오른쪽 배경에 여자가 있고, 실내에는 소파와 낮은 탁자, 텔레비전 및 긴 수납장이 보입니다. 유리의 나무와 건물 반사는 가능하지만 요구한 절제된 반사보다 강합니다.",
        "entities": "현우와 젊은 여성형 로봇만 등장합니다. 현우의 젊은 동아시아계 얼굴, 검은 머리, 체형과 회색 셔츠·카키색 바지·갈색 부츠가 참조와 대체로 맞습니다. 여자의 검은 머리, 사람 같은 젊은 동아시아계 얼굴, 남색 옷, 흰 앞치마, 기계 팔다리와 흰 신발도 맞으며 온전한 상태입니다. 여자가 쥔 손잡이 달린 흰 먼지털이는 명확합니다. 현관 잠금장치는 요청한 구식 열쇠식보다 전자식에 가깝습니다. 읽을 수 있는 글자나 추가 인물은 보이지 않습니다.",
        "hard_violations": [
         "현우가 창 옆 불투명한 외벽 뒤에 머물러 있어, 요구된 여자에 대한 직접적인 응시가 공간적으로 성립하지 않습니다."
        ],
        "physics": "현우의 두 부츠는 현관 바닥에 닿고 어깨는 벽에 가까이 기울어 있어 지지와 체중 배분이 가능합니다. 여자의 두 신발은 실내 바닥에 놓여 있고, 한 손이 먼지털이 손잡이를 잡아 소파 쪽으로 내립니다. 몸과 도구 모두 지지가 확인되며 먼지를 터는 동작으로 가능한 자세입니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "문에 밀착한 자세는 맞지만 외벽에 가로막힌 시선, 변형된 현관 공간, 불명확한 먼지 털기 동작 때문에 핵심 장면의 충실도가 낮습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "장소의 구조와 와이드 배치, 여자의 먼지 털기를 더 충실히 구현했지만 현우가 외벽 뒤에 있어 여자를 직접 응시하는 시선 관계는 성립하지 않습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 오른쪽 창가의 여자 방향으로 돌리고 있습니다. 그러나 눈과 코가 창 옆 두꺼운 외벽의 왼쪽 뒤에 머물러 있어, 여자까지의 시선은 불투명한 벽에 막힙니다. 여자는 고개를 숙여 자신의 손과 수납장 쪽을 보며 현우를 돌아보지 않습니다.",
        "built_space": "왼쪽에 닫힌 현관문 하나와 세로형 잠금장치 하나, 그 오른쪽에 벽등 하나가 보입니다. 오른쪽에는 중앙 세로 프레임으로 나뉜 대형 거실창과 양쪽 커튼이 있습니다. 실내의 소파, 낮은 탁자, 텔레비전과 긴 수납장은 참조 장소와 유사합니다. 다만 현관문이 참조보다 전면으로 당겨져 보이고, 현우 아래 바닥과 오른쪽 계단의 연결도 달라져 원래의 깊은 현관 공간이 약해졌습니다. 나무와 맞은편 건물의 유리 반사는 가능한 종류지만 상당히 강합니다.",
        "entities": "인물은 현우와 젊은 여성형 로봇 두 명뿐입니다. 현우의 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 회색 셔츠, 카키색 작업 바지와 부츠는 참조에 부합합니다. 여자는 검은 머리와 사람 같은 얼굴, 남색 옷과 흰 앞치마, 노출된 기계 팔·다리 및 흰 신발을 갖추고 있으며 손상은 보이지 않습니다. 손끝의 밝은 물체는 보이지만 먼지털이의 손잡이와 형태는 분명하지 않습니다. 문에는 구식 열쇠 자물쇠보다는 전자식 도어록에 가까운 장치가 있습니다. 읽을 수 있는 문구는 확인되지 않습니다.",
        "hard_violations": [
         "현우의 눈이 불투명한 외벽 뒤에 남아 있어 창 너머 여자를 직접 응시할 수 없는 위치로 배치되었습니다."
        ],
        "physics": "현우는 두 부츠를 현관 바닥과 매트에 딛고 다리를 벌린 채 몸을 기울이며 문 옆에 붙어 있습니다. 과장된 자세지만 바닥 지지가 있어 부유하지 않습니다. 여자는 두 발로 실내 바닥을 딛고 허리를 숙이고 있으며, 앞으로 내민 팔과 손도 몸에 연결되어 있습니다. 손끝 물체의 정확한 파지 형태는 작아서 확정하기 어렵지만 독립적으로 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 오른쪽 거실창 쪽을 향합니다. 다만 머리가 현관 측벽의 바깥 모서리를 넘어 나오지 않아 여자까지의 직접적인 시선은 두꺼운 외벽을 통과해야 합니다. 여자는 아래쪽 소파와 먼지털이 끝을 바라보고 있으며 관찰자를 의식하지 않는 모습입니다. 먼지털이는 손에서 왼쪽 아래 소파 좌석 방향으로 뻗어 있습니다.",
        "built_space": "왼쪽의 깊게 들어간 현관에 닫힌 문 하나, 문 부착 잠금장치 하나, 벽등 하나, 호출기 하나, 천장 매입등 하나가 보입니다. 현관 바닥과 전면 계단, 문과 창 사이의 두꺼운 회색 벽, 오른쪽 대형 창의 구성이 참조와 가깝습니다. 현우는 문 바로 앞보다는 현관 오른쪽 바깥 모서리에 붙어 있습니다. 닫힌 창 너머 중간 오른쪽 배경에 여자가 있고, 실내에는 소파와 낮은 탁자, 텔레비전 및 긴 수납장이 보입니다. 유리의 나무와 건물 반사는 가능하지만 요구한 절제된 반사보다 강합니다.",
        "entities": "현우와 젊은 여성형 로봇만 등장합니다. 현우의 젊은 동아시아계 얼굴, 검은 머리, 체형과 회색 셔츠·카키색 바지·갈색 부츠가 참조와 대체로 맞습니다. 여자의 검은 머리, 사람 같은 젊은 동아시아계 얼굴, 남색 옷, 흰 앞치마, 기계 팔다리와 흰 신발도 맞으며 온전한 상태입니다. 여자가 쥔 손잡이 달린 흰 먼지털이는 명확합니다. 현관 잠금장치는 요청한 구식 열쇠식보다 전자식에 가깝습니다. 읽을 수 있는 글자나 추가 인물은 보이지 않습니다.",
        "hard_violations": [
         "현우가 창 옆 불투명한 외벽 뒤에 머물러 있어, 요구된 여자에 대한 직접적인 응시가 공간적으로 성립하지 않습니다."
        ],
        "physics": "현우의 두 부츠는 현관 바닥에 닿고 어깨는 벽에 가까이 기울어 있어 지지와 체중 배분이 가능합니다. 여자의 두 신발은 실내 바닥에 놓여 있고, 한 손이 먼지털이 손잡이를 잡아 소파 쪽으로 내립니다. 몸과 도구 모두 지지가 확인되며 먼지를 터는 동작으로 가능한 자세입니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.179
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.929
   },
   "violations": {
    "B": [
     "[gemini-pro] 현우의 하체 및 다리 주변 벽면과 계단 구조가 절단된 것처럼 겹치고 깨지는 심각한 공간 콜라주/렌더링 오류",
     "[gpt-high] 현우의 눈이 불투명한 외벽 뒤에 남아 있어 창 너머 여자를 직접 응시할 수 없는 위치로 배치되었습니다."
    ],
    "A": [
     "[gpt-high] 현우가 창 옆 불투명한 외벽 뒤에 머물러 있어, 요구된 여자에 대한 직접적인 응시가 공간적으로 성립하지 않습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 공간과 인물의 신원을 정확히 구현했으며, 물리적 오류 없이 지시문의 상황을 안정적으로 연출했습니다.  ★위반: [gpt-high] 현우가 창 옆 불투명한 외벽 뒤에 머물러 있어, 요구된 여자에 대한 직접적인 응시가 공간적으로 성립하지 않습니다."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "현우의 역동적인 자세는 돋보이나, 그가 서 있는 하체 주변의 건축 구조가 기형적으로 깨지는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 현우의 하체 및 다리 주변 벽면과 계단 구조가 절단된 것처럼 겹치고 깨지는 심각한 공간 콜라주/렌더링 오류 / [gpt-high] 현우의 눈이 불투명한 외벽 뒤에 남아 있어 창 너머 여자를 직접 응시할 수 없는 위치로 배치되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B03.png",
    "asset_id": "1ced485f-4388-40e8-9db1-c1e4432bbd3c",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:896577>",
    "asset_id": "d19d1741-7b2a-408b-a66b-96afe66a9293",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-8a04-7641-855e-f2eec579cd00",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1__bgfirst_bg.png",
   "bg_asset_id": "33fddc9a-f643-4e6e-9f6c-5eeb879142cd",
   "bg_record_key": "S5sh1::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S5sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T15:59:42.422904+00:00",
  "fingerprint": "0e9fae181392f4e28ac1383f4d1ff4c818a178ae3d0c68d15e54da853e0326e8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S5sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S5sh1_sel.png",
  "source_sha256": "2e0fe7ca1470016ea1b5430dba0de2b392d6196bcf8682efb865e1e3d7765be6",
  "file": "S5sh1_cine.png",
  "staged_sha256": "1e001c1bd4fd45392704c187003492bef94b1ef08b6b8792cee0cec7670fb510",
  "latency_ms": 10748
 },
 "S5sh9::signage": {
  "fp": "b3c980b8f47dc17b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S5sh9": {
  "input_fingerprint": "569927066f675b06",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 셰퍼드에게 물어뜯긴 여자의 머리 가죽 아래로 은빛 금속 재질이 드러난 상태.\n\nLOCATION (lock): On the ground immediately outside the opened living-room window of an upscale house, where the fallen household robot is attacked. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Living-room floor (Supporting the fallen robot woman after impact); used as A narrow contextual plane around the head and shoulder; Open living-room window (Opened by the woman before her collapse) — Only an edge of the opening remains visible behind the low foreground action; used as Maintains continuity with the exterior approach and the next movement into the house.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light makes the exposed silver metal distinct from the torn outer covering without adding a supernatural glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The disabled female robot is collapsed on the floor after being dropped by Hyunwoo, with her head at floor level as the shepherd tears its covering and exposes the metal underneath. The source does not specify her torso's orientation or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The living-room window remains open, and the female housekeeping robot lies on the floor with torn head skin exposing metal underneath. The phone used for the call remains at the house, and the German shepherd is biting the robot's head.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리); 셰퍼드 (성견 셰퍼드, 검정·황갈색 털, 곧게 선 삼각형 귀, 긴 주둥이, 두툼한 꼬리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 셰퍼드에게 물어뜯긴 여자의 머리 가죽 아래로 은빛 금속 재질이 드러난 상태.\n\nLOCATION (lock): On the ground immediately outside the opened living-room window of an upscale house, where the fallen household robot is attacked. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Living-room floor (Supporting the fallen robot woman after impact); used as A narrow contextual plane around the head and shoulder; Open living-room window (Opened by the woman before her collapse) — Only an edge of the opening remains visible behind the low foreground action; used as Maintains continuity with the exterior approach and the next movement into the house.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light makes the exposed silver metal distinct from the torn outer covering without adding a supernatural glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The disabled female robot is collapsed on the floor after being dropped by Hyunwoo, with her head at floor level as the shepherd tears its covering and exposes the metal underneath. The source does not specify her torso's orientation or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The living-room window remains open, and the female housekeeping robot lies on the floor with torn head skin exposing metal underneath. The phone used for the call remains at the house, and the German shepherd is biting the robot's head.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리); 셰퍼드 (성견 셰퍼드, 검정·황갈색 털, 곧게 선 삼각형 귀, 긴 주둥이, 두툼한 꼬리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 셰퍼드에게 물어뜯긴 여자의 머리 가죽 아래로 은빛 금속 재질이 드러난 상태.\n\nLOCATION (lock): On the ground immediately outside the opened living-room window of an upscale house, where the fallen household robot is attacked. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Living-room floor (Supporting the fallen robot woman after impact); used as A narrow contextual plane around the head and shoulder; Open living-room window (Opened by the woman before her collapse) — Only an edge of the opening remains visible behind the low foreground action; used as Maintains continuity with the exterior approach and the next movement into the house.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light makes the exposed silver metal distinct from the torn outer covering without adding a supernatural glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The disabled female robot is collapsed on the floor after being dropped by Hyunwoo, with her head at floor level as the shepherd tears its covering and exposes the metal underneath. The source does not specify her torso's orientation or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The living-room window remains open, and the female housekeeping robot lies on the floor with torn head skin exposing metal underneath. The phone used for the call remains at the house, and the German shepherd is biting the robot's head.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리); 셰퍼드 (성견 셰퍼드, 검정·황갈색 털, 곧게 선 삼각형 귀, 긴 주둥이, 두툼한 꼬리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "개의 주둥이가 로봇의 얼굴 가장자리를 향해 물어뜯고 있음.",
    "built_space": "열린 거실 창문 및 내부 공간. 로봇이 창틀 부분에 기대어 눕혀 있고 배경에 거실 가구가 배치됨.",
    "entities": "개는 셰퍼드로 잘 묘사됨. 로봇은 얼굴은 일치하나, 요구된 남색 셔츠와 흰색 앞치마가 전혀 없고 맨 금속 몸체로 묘사되어 외형 일치 조건에서 어긋남.",
    "hard_violations": [
     "[gpt-high] 바닥 높이에 완전히 쓰러진 머리 대신, 외부 지면보다 높은 창턱 위에 머리를 받친 상태로 배치했습니다."
    ],
    "physics": "로봇의 머리와 몸통은 바닥 및 창틀에 닿아 지탱되고 있으며, 개는 다리로 바닥을 디디고 서 있음."
   },
   {
    "label": "B",
    "direction": "개의 주둥이가 로봇의 머리 윗부분을 향해 피부 가죽을 물어당기고 있음.",
    "built_space": "거실 창문 밖에서 내부를 바라본 구도. 로봇이 거실 안쪽 나무 바닥에 엎어져 있음.",
    "entities": "개는 셰퍼드. 로봇은 피 묻은 흰 앞치마와 남색 셔츠, 기계 팔 등 레퍼런스의 지정된 모든 외형 요소를 정확히 재현함.",
    "hard_violations": [
     "[gpt-high] 창문 바로 바깥 지면에 쓰러져 있어야 할 로봇의 머리와 몸통을 거실 안쪽 바닥에 배치했습니다."
    ],
    "physics": "로봇의 몸통과 팔다리는 거실 바닥에 완전히 밀착되어 지탱되며, 개는 네 발로 바닥을 안정적으로 딛고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 클로즈업보다 약간 넓은 구도이며, 로봇의 지정된 의상(남색 셔츠, 앞치마)을 완전히 생략하고 임의의 금속 몸통으로 왜곡하여 레퍼런스 일관성에서 크게 실패했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 클로즈업 샷보다 프레임이 다소 넓게(미디엄 샷) 잡혔으나, 로봇의 표정, 지정된 의상, 기계 팔 등 레퍼런스 외형을 완벽히 재현하여 전반적인 프롬프트 충실도가 훨씬 뛰어납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "개의 주둥이가 로봇의 얼굴 가장자리를 향해 물어뜯고 있음.",
        "built_space": "열린 거실 창문 및 내부 공간. 로봇이 창틀 부분에 기대어 눕혀 있고 배경에 거실 가구가 배치됨.",
        "entities": "개는 셰퍼드로 잘 묘사됨. 로봇은 얼굴은 일치하나, 요구된 남색 셔츠와 흰색 앞치마가 전혀 없고 맨 금속 몸체로 묘사되어 외형 일치 조건에서 어긋남.",
        "hard_violations": [],
        "physics": "로봇의 머리와 몸통은 바닥 및 창틀에 닿아 지탱되고 있으며, 개는 다리로 바닥을 디디고 서 있음."
       },
       {
        "label": "B",
        "direction": "개의 주둥이가 로봇의 머리 윗부분을 향해 피부 가죽을 물어당기고 있음.",
        "built_space": "거실 창문 밖에서 내부를 바라본 구도. 로봇이 거실 안쪽 나무 바닥에 엎어져 있음.",
        "entities": "개는 셰퍼드. 로봇은 피 묻은 흰 앞치마와 남색 셔츠, 기계 팔 등 레퍼런스의 지정된 모든 외형 요소를 정확히 재현함.",
        "hard_violations": [],
        "physics": "로봇의 몸통과 팔다리는 거실 바닥에 완전히 밀착되어 지탱되며, 개는 네 발로 바닥을 안정적으로 딛고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 클로즈업보다 약간 넓은 구도이며, 로봇의 지정된 의상(남색 셔츠, 앞치마)을 완전히 생략하고 임의의 금속 몸통으로 왜곡하여 레퍼런스 일관성에서 크게 실패했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 클로즈업 샷보다 프레임이 다소 넓게(미디엄 샷) 잡혔으나, 로봇의 표정, 지정된 의상, 기계 팔 등 레퍼런스 외형을 완벽히 재현하여 전반적인 프롬프트 충실도가 훨씬 뛰어납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "개의 주둥이가 로봇의 얼굴 가장자리를 향해 물어뜯고 있음.",
        "built_space": "열린 거실 창문 및 내부 공간. 로봇이 창틀 부분에 기대어 눕혀 있고 배경에 거실 가구가 배치됨.",
        "entities": "개는 셰퍼드로 잘 묘사됨. 로봇은 얼굴은 일치하나, 요구된 남색 셔츠와 흰색 앞치마가 전혀 없고 맨 금속 몸체로 묘사되어 외형 일치 조건에서 어긋남.",
        "hard_violations": [],
        "physics": "로봇의 머리와 몸통은 바닥 및 창틀에 닿아 지탱되고 있으며, 개는 다리로 바닥을 디디고 서 있음."
       },
       {
        "label": "B",
        "direction": "개의 주둥이가 로봇의 머리 윗부분을 향해 피부 가죽을 물어당기고 있음.",
        "built_space": "거실 창문 밖에서 내부를 바라본 구도. 로봇이 거실 안쪽 나무 바닥에 엎어져 있음.",
        "entities": "개는 셰퍼드. 로봇은 피 묻은 흰 앞치마와 남색 셔츠, 기계 팔 등 레퍼런스의 지정된 모든 외형 요소를 정확히 재현함.",
        "hard_violations": [],
        "physics": "로봇의 몸통과 팔다리는 거실 바닥에 완전히 밀착되어 지탱되며, 개는 네 발로 바닥을 안정적으로 딛고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "의상과 은빛 두피 노출은 충실하지만, 로봇이 창밖 지면이 아닌 실내에 놓였고 창틀과 상체까지 넓게 담아 지정된 머리 중심 클로즈업에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "머리를 무는 순간과 얼굴 중심 클로즈업은 더 정확하지만, 머리가 지면 대신 높은 창턱에 놓였고 화면에 보이는 어깨와 가슴의 고정 의상이 사라졌습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "셰퍼드는 주둥이를 아래로 내려 로봇의 정수리와 관자놀이 쪽 찢어진 외피에 물고 있습니다. 눈도 머리 쪽을 향합니다. 로봇은 얼굴을 옆으로 돌린 채 눈을 아래쪽으로 느슨하게 뜨고 있으며 뚜렷한 응시 대상은 없습니다.",
        "built_space": "하나의 열린 미닫이창 개구부와 양옆의 검은 창틀, 아래쪽 창 레일과 회색 창턱이 보입니다. 실내에는 밝은 나무 바닥, 소파 일부, 낮은 목재 탁자 하나와 뒤쪽 수납장이 보이며 참고 장소의 재료와 대체로 맞습니다. 그러나 로봇의 머리와 몸통은 개구부 안쪽 나무 바닥에 있고 한쪽 손만 바깥 창턱에 걸쳐 있습니다. 창 가장자리만 남기라는 지시보다 창 구조와 실내가 넓게 드러납니다.",
        "entities": "젊은 동아시아 여성형 얼굴과 검은 머리, 남색 옷과 흰 앞치마, 기계 전완이 보입니다. 참고 로봇의 전반적인 외형과 의상에 부합합니다. 찢어진 머리 외피 아래에는 반사되는 은색 두개부와 배선이 드러납니다. 성견 셰퍼드는 검정·황갈색 털, 긴 주둥이와 선 삼각형 귀를 갖췄습니다. 다른 사람이나 읽을 수 있는 글자는 없으며 전화기는 보이지 않습니다.",
        "hard_violations": [
         "창문 바로 바깥 지면에 쓰러져 있어야 할 로봇의 머리와 몸통을 거실 안쪽 바닥에 배치했습니다."
        ],
        "physics": "로봇의 머리와 몸통은 나무 바닥에, 화면 오른쪽 손은 회색 창턱에 받쳐져 있어 무지지 부유는 보이지 않습니다. 셰퍼드의 한 앞발은 로봇의 어깨 부근을 누르고 다른 앞다리는 바닥으로 내려갑니다. 턱이 찢어진 두피에 접촉해 물어뜯는 동작과 금속 노출의 관계가 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "셰퍼드는 화면 오른쪽에서 왼쪽 아래로 주둥이를 뻗어 로봇의 관자놀이와 두피 외피를 물고 있습니다. 시선도 물고 있는 머리 쪽을 향합니다. 로봇의 눈은 화면 왼쪽 아래로 풀려 있고 특정 대상을 응시하지 않습니다.",
        "built_space": "열린 창의 아래 레일 하나와 오른쪽 수직 창틀 일부, 왼쪽 커튼 한 폭이 보입니다. 뒤에는 나무 바닥, 밝은 소파 하나와 낮은 목재 탁자 하나가 있어 참고 장소의 기본 재료가 이어집니다. 얼굴과 개의 머리를 크게 담아 A보다 클로즈업에 가깝습니다. 다만 로봇의 머리는 창밖 지면이 아니라 높이가 있는 회색 창턱에 놓였고, 몸통은 그 바깥쪽 아래로 이어집니다. 거실 바닥은 머리를 직접 받치지 않습니다.",
        "entities": "검은 머리의 젊은 동아시아 여성형 얼굴과 성견 셰퍼드 한 마리가 보입니다. 개의 털색, 긴 주둥이와 선 귀는 참고에 부합합니다. 두피 아래 은빛 금속은 보이지만 머리보다 목·어깨·가슴의 커다란 기계 구조가 두드러집니다. 화면에 포함된 어깨와 가슴에 참고의 남색 옷과 흰 앞치마가 없어 의상 연속성이 깨집니다. 다른 사람, 읽을 수 있는 글자, 전화기는 보이지 않습니다.",
        "hard_violations": [
         "바닥 높이에 완전히 쓰러진 머리 대신, 외부 지면보다 높은 창턱 위에 머리를 받친 상태로 배치했습니다."
        ],
        "physics": "머리 옆면과 머리카락은 회색 창턱에 받쳐져 있고, 기계 몸통은 화면 오른쪽 아래로 비스듬히 이어집니다. 하체와 몸통의 최종 지지점은 프레임 밖이므로 공중에 떠 있다고 단정할 수는 없습니다. 셰퍼드의 몸과 앞다리도 아래쪽 프레임 밖으로 이어지며 발 접지는 보이지 않습니다. 들린 두피 조각은 개의 이와 남아 있는 외피에 연결되어 있어 지지 없이 떠 있는 물체는 아닙니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "의상과 은빛 두피 노출은 충실하지만, 로봇이 창밖 지면이 아닌 실내에 놓였고 창틀과 상체까지 넓게 담아 지정된 머리 중심 클로즈업에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "머리를 무는 순간과 얼굴 중심 클로즈업은 더 정확하지만, 머리가 지면 대신 높은 창턱에 놓였고 화면에 보이는 어깨와 가슴의 고정 의상이 사라졌습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "셰퍼드는 주둥이를 아래로 내려 로봇의 정수리와 관자놀이 쪽 찢어진 외피에 물고 있습니다. 눈도 머리 쪽을 향합니다. 로봇은 얼굴을 옆으로 돌린 채 눈을 아래쪽으로 느슨하게 뜨고 있으며 뚜렷한 응시 대상은 없습니다.",
        "built_space": "하나의 열린 미닫이창 개구부와 양옆의 검은 창틀, 아래쪽 창 레일과 회색 창턱이 보입니다. 실내에는 밝은 나무 바닥, 소파 일부, 낮은 목재 탁자 하나와 뒤쪽 수납장이 보이며 참고 장소의 재료와 대체로 맞습니다. 그러나 로봇의 머리와 몸통은 개구부 안쪽 나무 바닥에 있고 한쪽 손만 바깥 창턱에 걸쳐 있습니다. 창 가장자리만 남기라는 지시보다 창 구조와 실내가 넓게 드러납니다.",
        "entities": "젊은 동아시아 여성형 얼굴과 검은 머리, 남색 옷과 흰 앞치마, 기계 전완이 보입니다. 참고 로봇의 전반적인 외형과 의상에 부합합니다. 찢어진 머리 외피 아래에는 반사되는 은색 두개부와 배선이 드러납니다. 성견 셰퍼드는 검정·황갈색 털, 긴 주둥이와 선 삼각형 귀를 갖췄습니다. 다른 사람이나 읽을 수 있는 글자는 없으며 전화기는 보이지 않습니다.",
        "hard_violations": [
         "창문 바로 바깥 지면에 쓰러져 있어야 할 로봇의 머리와 몸통을 거실 안쪽 바닥에 배치했습니다."
        ],
        "physics": "로봇의 머리와 몸통은 나무 바닥에, 화면 오른쪽 손은 회색 창턱에 받쳐져 있어 무지지 부유는 보이지 않습니다. 셰퍼드의 한 앞발은 로봇의 어깨 부근을 누르고 다른 앞다리는 바닥으로 내려갑니다. 턱이 찢어진 두피에 접촉해 물어뜯는 동작과 금속 노출의 관계가 자연스럽습니다."
       },
       {
        "label": "A",
        "direction": "셰퍼드는 화면 오른쪽에서 왼쪽 아래로 주둥이를 뻗어 로봇의 관자놀이와 두피 외피를 물고 있습니다. 시선도 물고 있는 머리 쪽을 향합니다. 로봇의 눈은 화면 왼쪽 아래로 풀려 있고 특정 대상을 응시하지 않습니다.",
        "built_space": "열린 창의 아래 레일 하나와 오른쪽 수직 창틀 일부, 왼쪽 커튼 한 폭이 보입니다. 뒤에는 나무 바닥, 밝은 소파 하나와 낮은 목재 탁자 하나가 있어 참고 장소의 기본 재료가 이어집니다. 얼굴과 개의 머리를 크게 담아 A보다 클로즈업에 가깝습니다. 다만 로봇의 머리는 창밖 지면이 아니라 높이가 있는 회색 창턱에 놓였고, 몸통은 그 바깥쪽 아래로 이어집니다. 거실 바닥은 머리를 직접 받치지 않습니다.",
        "entities": "검은 머리의 젊은 동아시아 여성형 얼굴과 성견 셰퍼드 한 마리가 보입니다. 개의 털색, 긴 주둥이와 선 귀는 참고에 부합합니다. 두피 아래 은빛 금속은 보이지만 머리보다 목·어깨·가슴의 커다란 기계 구조가 두드러집니다. 화면에 포함된 어깨와 가슴에 참고의 남색 옷과 흰 앞치마가 없어 의상 연속성이 깨집니다. 다른 사람, 읽을 수 있는 글자, 전화기는 보이지 않습니다.",
        "hard_violations": [
         "바닥 높이에 완전히 쓰러진 머리 대신, 외부 지면보다 높은 창턱 위에 머리를 받친 상태로 배치했습니다."
        ],
        "physics": "머리 옆면과 머리카락은 회색 창턱에 받쳐져 있고, 기계 몸통은 화면 오른쪽 아래로 비스듬히 이어집니다. 하체와 몸통의 최종 지지점은 프레임 밖이므로 공중에 떠 있다고 단정할 수는 없습니다. 셰퍼드의 몸과 앞다리도 아래쪽 프레임 밖으로 이어지며 발 접지는 보이지 않습니다. 들린 두피 조각은 개의 이와 남아 있는 외피에 연결되어 있어 지지 없이 떠 있는 물체는 아닙니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.8
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.55
   },
   "violations": {
    "B": [
     "[gpt-high] 창문 바로 바깥 지면에 쓰러져 있어야 할 로봇의 머리와 몸통을 거실 안쪽 바닥에 배치했습니다."
    ],
    "A": [
     "[gpt-high] 바닥 높이에 완전히 쓰러진 머리 대신, 외부 지면보다 높은 창턱 위에 머리를 받친 상태로 배치했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1550
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "지정된 클로즈업보다 약간 넓은 구도이며, 로봇의 지정된 의상(남색 셔츠, 앞치마)을 완전히 생략하고 임의의 금속 몸통으로 왜곡하여 레퍼런스 일관성에서 크게 실패했습니다.  ★위반: [gpt-high] 바닥 높이에 완전히 쓰러진 머리 대신, 외부 지면보다 높은 창턱 위에 머리를 받친 상태로 배치했습니다."
   },
   {
    "label": "B",
    "score": 1550,
    "verdict_ko": "요구된 클로즈업 샷보다 프레임이 다소 넓게(미디엄 샷) 잡혔으나, 로봇의 표정, 지정된 의상, 기계 팔 등 레퍼런스 외형을 완벽히 재현하여 전반적인 프롬프트 충실도가 훨씬 뛰어납니다.  ★위반: [gpt-high] 창문 바로 바깥 지면에 쓰러져 있어야 할 로봇의 머리와 몸통을 거실 안쪽 바닥에 배치했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1_sel.png",
    "asset_id": "195e2985-7b31-486a-aabe-d8758d9b63f0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1301425>",
    "asset_id": "2444bfb9-b526-4c89-8276-89bc94110e2e",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 셰퍼드: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1175105>",
    "asset_id": "9bb4af18-89d8-41de-87e2-7ccfcebbd5eb",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-8d57-7204-8f8b-47ee7146513d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S5sh1"
  },
  "staged_characters_added": [
   "C45"
  ]
 },
 "S5sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T10:55:23.510851+00:00",
  "fingerprint": "7a3cf5a1f9a848b5e2fb45c0142859c1c142760a880a54126e12fb8e2d317b9f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S5sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S5sh9_sel.png",
  "source_sha256": "a15a7d8290428f184ab9893dd66a0410ffb4b5afce07b0fb7c8e1225b88eb810",
  "file": "S5sh9_cine.png",
  "staged_sha256": "617e55290eab768dae735553b450cce3b28f0259205f1deffc5f83482adbe1fc",
  "latency_ms": 9579
 },
 "S5sh16::signage": {
  "fp": "3ff843c6248d527a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S5sh16::bgfirst_bg": {
  "input_fingerprint": "415cd6aa16584117",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh16__bgfirst_bg.png",
  "asset_id": "87ad0c21-6893-428f-8870-cce81e611fed",
  "input_asset_ids": [
   "32035306-c502-4408-8bfc-f7eedf49376d",
   "e2cccbcb-d0f6-4b30-ac00-d35ae2ddd10b"
  ]
 },
 "S5sh16": {
  "input_fingerprint": "ae9b42bad8ec3a46",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car has disappeared from the roadside; the living-room window and the second-floor escape window remain open. The female housekeeping robot remains on the floor with torn head skin and exposed metal. 현우: Hyunwoo has a fresh dog-bite wound on his leg and is being led away under arrest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car has disappeared from the roadside; the living-room window and the second-floor escape window remain open. The female housekeeping robot remains on the floor with torn head skin and exposed metal. 현우: Hyunwoo has a fresh dog-bite wound on his leg and is being led away under arrest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car has disappeared from the roadside; the living-room window and the second-floor escape window remain open. The female housekeeping robot remains on the floor with torn head skin and exposed metal. 현우: Hyunwoo has a fresh dog-bite wound on his leg and is being led away under arrest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh16__bgfirst_bg.png",
     "asset_id": "87ad0c21-6893-428f-8870-cce81e611fed",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S5sh16.png",
     "asset_id": "32035306-c502-4408-8bfc-f7eedf49376d",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B01.png",
     "asset_id": "e2cccbcb-d0f6-4b30-ac00-d35ae2ddd10b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우는 오른쪽 어깨 너머로 화면 중좌측 배경에 있는 텅 빈 길가를 돌아보고 있음.",
    "built_space": "왼쪽에 담장, 오른쪽에 주택들이 있는 거리 구조이나, 제공된 로케이션 레퍼런스의 건축 양식을 완전히 무시하고 임의의 건물을 생성함.",
    "entities": "현우의 외모와 의상은 레퍼런스와 일치함. 두 경찰관이 등장함. 프레임 밖인 다리 대신 셔츠 하단에 출처를 알 수 없는 핏자국이 묘사됨.",
    "hard_violations": [],
    "physics": "지면에 서서 두 경찰관의 손에 양팔이 단단히 붙잡힌 상태를 자연스럽게 유지함."
   },
   {
    "label": "B",
    "direction": "현우는 화면 정면 좌측을 응시할 뿐, 배경에 있는 텅 빈 길가 쪽을 돌아보는 시선을 연출하지 않음.",
    "built_space": "콘크리트 담장, 녹슨 철문, 도로 등 로케이션 레퍼런스의 배경 구조와 재질을 완벽하게 재현함.",
    "entities": "현우의 외모는 레퍼런스와 일치하며 오른쪽 허벅지에 개물림 상처가 정확히 묘사됨. 경찰관들에게 제압당한 상태임.",
    "hard_violations": [
     "[gpt-high] 오른쪽 경찰 조끼에 읽을 수 있는 '경' 및 영문 표기가 노출되어, 읽히는 글자를 금지한 조건을 위반한다."
    ],
    "physics": "지면에 서 있으며, 우측 경찰관이 검은 장갑을 낀 손으로 현우의 왼팔을 단단히 쥐고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "로케이션 지시를 어겼으나, 핵심 액션인 뒤돌아보는 시선과 프레임 구도를 정확히 구현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "로케이션과 상처 디테일은 완벽하나, 필수 액션인 뒤돌아보는 시선과 배경 배치를 완전히 실패함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 오른쪽 어깨 너머로 화면 중좌측 배경에 있는 텅 빈 길가를 돌아보고 있음.",
        "built_space": "왼쪽에 담장, 오른쪽에 주택들이 있는 거리 구조이나, 제공된 로케이션 레퍼런스의 건축 양식을 완전히 무시하고 임의의 건물을 생성함.",
        "entities": "현우의 외모와 의상은 레퍼런스와 일치함. 두 경찰관이 등장함. 프레임 밖인 다리 대신 셔츠 하단에 출처를 알 수 없는 핏자국이 묘사됨.",
        "hard_violations": [],
        "physics": "지면에 서서 두 경찰관의 손에 양팔이 단단히 붙잡힌 상태를 자연스럽게 유지함."
       },
       {
        "label": "B",
        "direction": "현우는 화면 정면 좌측을 응시할 뿐, 배경에 있는 텅 빈 길가 쪽을 돌아보는 시선을 연출하지 않음.",
        "built_space": "콘크리트 담장, 녹슨 철문, 도로 등 로케이션 레퍼런스의 배경 구조와 재질을 완벽하게 재현함.",
        "entities": "현우의 외모는 레퍼런스와 일치하며 오른쪽 허벅지에 개물림 상처가 정확히 묘사됨. 경찰관들에게 제압당한 상태임.",
        "hard_violations": [],
        "physics": "지면에 서 있으며, 우측 경찰관이 검은 장갑을 낀 손으로 현우의 왼팔을 단단히 쥐고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "로케이션 지시를 어겼으나, 핵심 액션인 뒤돌아보는 시선과 프레임 구도를 정확히 구현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "로케이션과 상처 디테일은 완벽하나, 필수 액션인 뒤돌아보는 시선과 배경 배치를 완전히 실패함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 오른쪽 어깨 너머로 화면 중좌측 배경에 있는 텅 빈 길가를 돌아보고 있음.",
        "built_space": "왼쪽에 담장, 오른쪽에 주택들이 있는 거리 구조이나, 제공된 로케이션 레퍼런스의 건축 양식을 완전히 무시하고 임의의 건물을 생성함.",
        "entities": "현우의 외모와 의상은 레퍼런스와 일치함. 두 경찰관이 등장함. 프레임 밖인 다리 대신 셔츠 하단에 출처를 알 수 없는 핏자국이 묘사됨.",
        "hard_violations": [],
        "physics": "지면에 서서 두 경찰관의 손에 양팔이 단단히 붙잡힌 상태를 자연스럽게 유지함."
       },
       {
        "label": "B",
        "direction": "현우는 화면 정면 좌측을 응시할 뿐, 배경에 있는 텅 빈 길가 쪽을 돌아보는 시선을 연출하지 않음.",
        "built_space": "콘크리트 담장, 녹슨 철문, 도로 등 로케이션 레퍼런스의 배경 구조와 재질을 완벽하게 재현함.",
        "entities": "현우의 외모는 레퍼런스와 일치하며 오른쪽 허벅지에 개물림 상처가 정확히 묘사됨. 경찰관들에게 제압당한 상태임.",
        "hard_violations": [],
        "physics": "지면에 서 있으며, 우측 경찰관이 검은 장갑을 낀 손으로 현우의 왼팔을 단단히 쥐고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "장소의 낡은 담장과 대문은 잘 맞지만, 현우가 빈 길가 반대쪽을 보고 있으며 허벅지까지 보이는 넓은 구도와 읽히는 경찰 표기가 결정적으로 어긋납니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "양팔을 붙잡힌 채 왼쪽으로 돌아보는 동작과 빈 도로의 배치는 더 충실하지만, 얼굴 클로즈업이 아니며 주택의 외벽·담장·대문이 장소 참조와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 몸은 앞으로 기울어 있고 얼굴과 눈은 화면 오른쪽의 가까운 바깥쪽을 향한다. 비어 있는 길가가 놓인 화면 왼쪽을 바라보지 않아 원망의 대상과 시선이 연결되지 않는다. 뒤쪽 경찰의 눈은 가려져 있고 오른쪽 경찰은 현우 쪽으로 얼굴을 향하지만 눈은 프레임 밖이다.",
        "built_space": "왼쪽에 큰 미닫이창 한 조, 낡은 콘크리트 담장, 녹슨 철문 하나, 우편함 하나와 작은 창의 방범창이 보인다. 기와지붕, 배수로와 전봇대도 장소 참조에 가깝다. 현우와 경찰들은 도로 오른쪽 전경에 있고 왼쪽 길가는 비어 있다. 다만 얼굴이 아닌 허벅지까지 담아 요구한 클로즈업보다 훨씬 넓다. 보이는 큰 창에서는 열린 상태가 명확하지 않다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리, 회색 셔츠와 올리브색 바지가 참조에 대체로 맞는다. 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 얼굴은 찡그려져 있고 허벅지에는 피 묻은 상처가 있으나 개에게 물린 상처인지는 확정하기 어렵다. 경찰 두 명은 숏 텍스트가 요구하는 체포 행위에 참여한다. 페드로와 차량은 없다. 오른쪽 경찰의 가슴에는 '경'과 영문 일부가 읽혀 무문자 조건을 위반한다. 실내 로봇과 위층 탈출창은 구도 밖이므로 평가하지 않는다.",
        "hard_violations": [
         "오른쪽 경찰 조끼에 읽을 수 있는 '경' 및 영문 표기가 노출되어, 읽히는 글자를 금지한 조건을 위반한다."
        ],
        "physics": "오른쪽 경찰의 장갑 낀 손이 현우의 위팔을 잡고 있고 뒤쪽 경찰의 팔도 현우의 반대편 팔 부근에 닿아 있다. 뒤쪽으로 꺾인 팔과 앞으로 기운 상체는 연행 중 버티는 동작으로 가능하다. 발은 화면 밖이지만 다리가 아래로 이어지고 경찰의 붙잡는 접촉이 있어 공중에 뜬 신체로 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 카메라에서 멀어지는 몸통 방향과 반대로 목을 돌려 화면 왼쪽을 곁눈질한다. 왼쪽 배경에는 빈 도로가 있어 A보다 시선과 공간의 연결이 낫다. 다만 눈길은 먼 길가의 특정 지점보다는 카메라 왼쪽 가까운 방향으로도 읽혀, 차가 있던 자리를 정확히 보고 있다고 단정하기는 어렵다. 양쪽 경찰은 현우와 진행 방향 쪽으로 머리를 향한다.",
        "built_space": "오른쪽 주택에는 열린 위층 여닫이창 한 조와 아래층의 큰 창 한 조, 출입문 부근 기둥과 매끈한 담장이 보인다. 왼쪽부터 중앙까지 비어 있는 도로와 연석이 이어진다. 양쪽 경찰 사이에 현우를 둔 배치는 팔을 붙잡고 연행하기에 가능하다. 그러나 참조의 낡은 외벽, 녹슨 철문, 우편함과 담장 대신 다른 형태의 깨끗한 주택을 보여 정확한 장소 일치가 부족하다. 현우의 등과 몸통 대부분을 담아 얼굴 클로즈업에도 미달한다.",
        "entities": "현우는 참조와 유사한 젊은 동아시아계 남성의 얼굴, 헝클어진 검은 머리와 회색 셔츠를 갖고 있으며 미간과 입을 찡그린다. 양쪽에 경찰 두 명이 부분적으로 보이고 각각 현우의 팔을 잡는다. 페드로와 차량은 없으며 도로는 비어 있다. 읽을 수 있는 글자는 보이지 않는다. 다리의 상처, 하의와 실내 로봇은 프레임 밖이므로 불일치로 계산하지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 경찰의 장갑 낀 손이 현우의 한쪽 위팔을 감싸고 오른쪽 경찰도 반대쪽 팔을 붙잡는다. 손과 팔의 접촉이 분명하고 몸통을 앞으로 둔 채 머리만 뒤로 돌리는 자세는 자연스럽게 가능하다. 하체와 발은 잘려 있지만 신체가 정상적으로 프레임 아래로 이어져 지지 없는 부유나 불가능한 관절 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "장소의 낡은 담장과 대문은 잘 맞지만, 현우가 빈 길가 반대쪽을 보고 있으며 허벅지까지 보이는 넓은 구도와 읽히는 경찰 표기가 결정적으로 어긋납니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "양팔을 붙잡힌 채 왼쪽으로 돌아보는 동작과 빈 도로의 배치는 더 충실하지만, 얼굴 클로즈업이 아니며 주택의 외벽·담장·대문이 장소 참조와 다릅니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 몸은 앞으로 기울어 있고 얼굴과 눈은 화면 오른쪽의 가까운 바깥쪽을 향한다. 비어 있는 길가가 놓인 화면 왼쪽을 바라보지 않아 원망의 대상과 시선이 연결되지 않는다. 뒤쪽 경찰의 눈은 가려져 있고 오른쪽 경찰은 현우 쪽으로 얼굴을 향하지만 눈은 프레임 밖이다.",
        "built_space": "왼쪽에 큰 미닫이창 한 조, 낡은 콘크리트 담장, 녹슨 철문 하나, 우편함 하나와 작은 창의 방범창이 보인다. 기와지붕, 배수로와 전봇대도 장소 참조에 가깝다. 현우와 경찰들은 도로 오른쪽 전경에 있고 왼쪽 길가는 비어 있다. 다만 얼굴이 아닌 허벅지까지 담아 요구한 클로즈업보다 훨씬 넓다. 보이는 큰 창에서는 열린 상태가 명확하지 않다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리, 회색 셔츠와 올리브색 바지가 참조에 대체로 맞는다. 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 얼굴은 찡그려져 있고 허벅지에는 피 묻은 상처가 있으나 개에게 물린 상처인지는 확정하기 어렵다. 경찰 두 명은 숏 텍스트가 요구하는 체포 행위에 참여한다. 페드로와 차량은 없다. 오른쪽 경찰의 가슴에는 '경'과 영문 일부가 읽혀 무문자 조건을 위반한다. 실내 로봇과 위층 탈출창은 구도 밖이므로 평가하지 않는다.",
        "hard_violations": [
         "오른쪽 경찰 조끼에 읽을 수 있는 '경' 및 영문 표기가 노출되어, 읽히는 글자를 금지한 조건을 위반한다."
        ],
        "physics": "오른쪽 경찰의 장갑 낀 손이 현우의 위팔을 잡고 있고 뒤쪽 경찰의 팔도 현우의 반대편 팔 부근에 닿아 있다. 뒤쪽으로 꺾인 팔과 앞으로 기운 상체는 연행 중 버티는 동작으로 가능하다. 발은 화면 밖이지만 다리가 아래로 이어지고 경찰의 붙잡는 접촉이 있어 공중에 뜬 신체로 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 카메라에서 멀어지는 몸통 방향과 반대로 목을 돌려 화면 왼쪽을 곁눈질한다. 왼쪽 배경에는 빈 도로가 있어 A보다 시선과 공간의 연결이 낫다. 다만 눈길은 먼 길가의 특정 지점보다는 카메라 왼쪽 가까운 방향으로도 읽혀, 차가 있던 자리를 정확히 보고 있다고 단정하기는 어렵다. 양쪽 경찰은 현우와 진행 방향 쪽으로 머리를 향한다.",
        "built_space": "오른쪽 주택에는 열린 위층 여닫이창 한 조와 아래층의 큰 창 한 조, 출입문 부근 기둥과 매끈한 담장이 보인다. 왼쪽부터 중앙까지 비어 있는 도로와 연석이 이어진다. 양쪽 경찰 사이에 현우를 둔 배치는 팔을 붙잡고 연행하기에 가능하다. 그러나 참조의 낡은 외벽, 녹슨 철문, 우편함과 담장 대신 다른 형태의 깨끗한 주택을 보여 정확한 장소 일치가 부족하다. 현우의 등과 몸통 대부분을 담아 얼굴 클로즈업에도 미달한다.",
        "entities": "현우는 참조와 유사한 젊은 동아시아계 남성의 얼굴, 헝클어진 검은 머리와 회색 셔츠를 갖고 있으며 미간과 입을 찡그린다. 양쪽에 경찰 두 명이 부분적으로 보이고 각각 현우의 팔을 잡는다. 페드로와 차량은 없으며 도로는 비어 있다. 읽을 수 있는 글자는 보이지 않는다. 다리의 상처, 하의와 실내 로봇은 프레임 밖이므로 불일치로 계산하지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 경찰의 장갑 낀 손이 현우의 한쪽 위팔을 감싸고 오른쪽 경찰도 반대쪽 팔을 붙잡는다. 손과 팔의 접촉이 분명하고 몸통을 앞으로 둔 채 머리만 뒤로 돌리는 자세는 자연스럽게 가능하다. 하체와 발은 잘려 있지만 신체가 정상적으로 프레임 아래로 이어져 지지 없는 부유나 불가능한 관절 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.067
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.817
   },
   "violations": {
    "B": [
     "[gpt-high] 오른쪽 경찰 조끼에 읽을 수 있는 '경' 및 영문 표기가 노출되어, 읽히는 글자를 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 817
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "로케이션 지시를 어겼으나, 핵심 액션인 뒤돌아보는 시선과 프레임 구도를 정확히 구현함."
   },
   {
    "label": "B",
    "score": 817,
    "verdict_ko": "로케이션과 상처 디테일은 완벽하나, 필수 액션인 뒤돌아보는 시선과 배경 배치를 완전히 실패함.  ★위반: [gpt-high] 오른쪽 경찰 조끼에 읽을 수 있는 '경' 및 영문 표기가 노출되어, 읽히는 글자를 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B01.png",
    "asset_id": "e2cccbcb-d0f6-4b30-ac00-d35ae2ddd10b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-8efb-76fe-89e7-6052c71ae67a",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh16__bgfirst_bg.png",
   "bg_asset_id": "87ad0c21-6893-428f-8870-cce81e611fed",
   "bg_record_key": "S5sh16::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S5sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:57:07.117965+00:00",
  "fingerprint": "ab1e61464228cf5de73ebc12cff266750efb2d8d9bfd11a33cd748e9f544a12f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S5sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S5sh16_sel.png",
  "source_sha256": "c1fc7f5f50a357c35bb36404585a3b4a60d5e26f85887338b7ffac2cae6cdf7b",
  "file": "S5sh16_cine.png",
  "staged_sha256": "e680ef142c37eed59f0ed698211d69891305615c969e091465597163c76872a0",
  "latency_ms": 10966
 },
 "S6sh15::signage": {
  "fp": "caa95ef3c360673b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh15": {
  "input_fingerprint": "d00d7fa3d6cc804a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The repaired jukebox has its horn-shaped speaker connected and an LP fitted, with tools and additional records scattered around the blanket. Charlie is emerging from the rubbish as a dirt-covered, net-entangled gorilla-shaped robot with blue-lit eyes, stiff joints, and a worn Ubik chest logo.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The repaired jukebox has its horn-shaped speaker connected and an LP fitted, with tools and additional records scattered around the blanket. Charlie is emerging from the rubbish as a dirt-covered, net-entangled gorilla-shaped robot with blue-lit eyes, stiff joints, and a worn Ubik chest logo.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The repaired jukebox has its horn-shaped speaker connected and an LP fitted, with tools and additional records scattered around the blanket. Charlie is emerging from the rubbish as a dirt-covered, net-entangled gorilla-shaped robot with blue-lit eyes, stiff joints, and a worn Ubik chest logo.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15__bgfirst_bg.png",
     "asset_id": "8d75d911-3d18-4135-b6ab-fc26e495c901",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S6sh15.png",
     "asset_id": "c554fa18-115e-4e6b-9462-3d5f65b3a88e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_dump_emergence_3bffab.png",
     "asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 켜져 있음.",
    "built_space": "참조 이미지와 동일한 굴삭기가 보이는 매립지 배경. 찰리는 쓰레기 더미 사이에 위치함.",
    "entities": "머리에 그물과 쓰레기를 쓴 로봇 찰리로, 참조 이미지의 트렌치코트를 입고 있음.",
    "hard_violations": [],
    "physics": "양팔로 쓰레기 더미를 짚고 상체를 반쯤 일으킨 지탱 자세임."
   },
   {
    "label": "B",
    "direction": "시선은 정면 약간 아래를 향하며 두 눈에 파란 불빛이 켜져 있음.",
    "built_space": "폐가전과 쓰레기가 쌓인 매립지이나, 지정된 장소 참조 사진의 지형물이 없음.",
    "entities": "쓰레기를 얹은 로봇이나, 참조 이미지의 필수 복장인 트렌치코트가 없음.",
    "hard_violations": [
     "[gemini-pro] 금지된 읽을 수 있는 텍스트 노출 (가슴의 UBIK 로고)",
     "[gpt-high] 가슴에 ‘Ubik’이라는 읽을 수 있는 문자와 로고가 노출되어, 읽을 수 있는 글자 및 로고를 금지한 조건을 위반한다."
    ],
    "physics": "양손으로 쓰레기 더미를 짚어 상체를 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "캐릭터의 복장과 지정된 매립지 장소를 충실히 재현했으며, 지시된 지탱 자세를 잘 표현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캐릭터의 핵심 복장이 누락되었고, 명시적으로 금지된 읽을 수 있는 텍스트가 렌더링되어 지시를 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 켜져 있음.",
        "built_space": "참조 이미지와 동일한 굴삭기가 보이는 매립지 배경. 찰리는 쓰레기 더미 사이에 위치함.",
        "entities": "머리에 그물과 쓰레기를 쓴 로봇 찰리로, 참조 이미지의 트렌치코트를 입고 있음.",
        "hard_violations": [],
        "physics": "양팔로 쓰레기 더미를 짚고 상체를 반쯤 일으킨 지탱 자세임."
       },
       {
        "label": "B",
        "direction": "시선은 정면 약간 아래를 향하며 두 눈에 파란 불빛이 켜져 있음.",
        "built_space": "폐가전과 쓰레기가 쌓인 매립지이나, 지정된 장소 참조 사진의 지형물이 없음.",
        "entities": "쓰레기를 얹은 로봇이나, 참조 이미지의 필수 복장인 트렌치코트가 없음.",
        "hard_violations": [
         "금지된 읽을 수 있는 텍스트 노출 (가슴의 UBIK 로고)"
        ],
        "physics": "양손으로 쓰레기 더미를 짚어 상체를 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "캐릭터의 복장과 지정된 매립지 장소를 충실히 재현했으며, 지시된 지탱 자세를 잘 표현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캐릭터의 핵심 복장이 누락되었고, 명시적으로 금지된 읽을 수 있는 텍스트가 렌더링되어 지시를 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 켜져 있음.",
        "built_space": "참조 이미지와 동일한 굴삭기가 보이는 매립지 배경. 찰리는 쓰레기 더미 사이에 위치함.",
        "entities": "머리에 그물과 쓰레기를 쓴 로봇 찰리로, 참조 이미지의 트렌치코트를 입고 있음.",
        "hard_violations": [],
        "physics": "양팔로 쓰레기 더미를 짚고 상체를 반쯤 일으킨 지탱 자세임."
       },
       {
        "label": "B",
        "direction": "시선은 정면 약간 아래를 향하며 두 눈에 파란 불빛이 켜져 있음.",
        "built_space": "폐가전과 쓰레기가 쌓인 매립지이나, 지정된 장소 참조 사진의 지형물이 없음.",
        "entities": "쓰레기를 얹은 로봇이나, 참조 이미지의 필수 복장인 트렌치코트가 없음.",
        "hard_violations": [
         "금지된 읽을 수 있는 텍스트 노출 (가슴의 UBIK 로고)"
        ],
        "physics": "양손으로 쓰레기 더미를 짚어 상체를 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "양손으로 버티며 상체를 반쯤 일으킨 순간과 미디엄 구도는 더 정확하지만, 가슴의 읽을 수 있는 ‘Ubik’ 표기가 명시적인 문자·로고 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정 장소의 특징과 찰리의 외형·외투, 머리 위 쓰레기와 파란 두 눈을 충실히 재현했으나, 발까지 보이는 넓은 구도와 더 일으켜진 몸통은 요구한 미디엄 지탱 자세에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 파란 두 눈은 정면에서 약간 아래쪽, 카메라 앞 쓰레기 방향을 향한다. 지정된 응시 대상은 없으므로 방향상 충돌은 없다. 두 팔은 좌우 아래로 뻗어 쓰레기 더미를 짚는다.",
        "built_space": "건물 내부나 고정 설비가 아니라 폐전자제품과 녹슨 금속판이 쌓인 노출된 쓰레기 더미다. 몸 뒤로 높은 쓰레기 층이 이어지고 전경 잔해가 하체를 가린다. 장소 참고의 파란 컨테이너와 오른쪽 굴착기는 이 구도에서 보이지 않아 정확한 지점의 일치는 확인하기 어렵지만, 더 좁은 구도 자체를 결함으로 볼 수는 없다. 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "로봇은 한 대뿐이며 추가 인물은 없다. 육중한 긴 팔, 샌드 베이지 장갑판, 흰 마스크형 얼굴과 점·선 무늬, 파랗게 켜진 두 눈이 보인다. 흙과 그물이 몸에 걸려 있고 머리 위에는 작은 폐기기와 비닐·천 조각이 얹혀 있다. 참고의 올리브색 외투는 보이지 않고 몸통 장갑이 노출되어 있다. 가슴에는 ‘Ubik’이 읽혀 최종 문자 금지 조건과 충돌한다. 주크박스·담요·음반은 보이지 않으며 이 인물 중심 구도에서 반드시 포함할 이유는 없다.",
        "hard_violations": [
         "가슴에 ‘Ubik’이라는 읽을 수 있는 문자와 로고가 노출되어, 읽을 수 있는 글자 및 로고를 금지한 조건을 위반한다."
        ],
        "physics": "양손의 손가락과 주먹 아래쪽이 좌우의 금속 잔해와 쓰레기에 닿아 상체를 지탱한다. 팔꿈치를 굽히고 몸통을 앞으로 기울여 반쯤 일어나는 동작이 물리적으로 읽힌다. 하체는 잔해 사이에 묻혀 있고, 머리 위 폐기물은 머리·어깨와 그물에 받쳐져 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 고개를 화면 오른쪽으로 약간 돌리고 아래쪽 전방을 바라본다. 두 파란 눈 모두 보이며 특정 인물이나 물체를 겨냥하는 동작은 없다. 팔은 좌우 아래의 잔해를 향해 내려가 있다.",
        "built_space": "왼쪽의 파란 컨테이너 한 개, 오른쪽 먼 능선의 집게형 굴착기 한 대, 뒤로 이어지는 금속 폐기물 언덕이 장소 참고와 대응한다. 전경에는 타이어 두 개가 좌우로 보이고 잔해가 찰리의 하체를 둘러싼다. 굴착기는 먼 배경 크기로 유지된다. 다만 하늘과 발까지 포함해 요구한 미디엄보다 넓게 잡혔다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "추가 인물 없이 찰리 한 대만 있다. 긴 기계 팔, 베이지 장갑, 흰 마스크형 얼굴, 올리브색 벨트 외투는 참고와 가깝다. 얼굴의 선 표현은 있으나 참고 및 지시의 점 무늬는 덜 뚜렷하다. 두 눈은 파랗게 빛나며 머리와 어깨에는 그물, 캔, 금속 조각이 걸쳐져 있다. 표면에 먼지와 마모가 보이고, 읽을 수 있는 가슴 글자는 노출되지 않는다. 주크박스와 음반 등은 프레임에 없으나 이 구도 밖 소품의 누락은 감점 사유가 아니다.",
        "hard_violations": [],
        "physics": "화면 오른쪽 손가락 끝은 전경 잔해에 닿고, 왼쪽 손의 접점은 쓰레기에 일부 가려져 있다. 아래쪽의 굽힌 다리와 보이는 발도 쓰레기 표면에 받쳐져 있어 몸이 떠 있지는 않다. 다만 몸통이 비교적 높이 세워져 양팔로 반쯤 일어난 상체를 버티는 순간보다는 더 일어난 단계로 읽힌다. 머리 위 물체들은 그물과 머리·어깨에 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "양손으로 버티며 상체를 반쯤 일으킨 순간과 미디엄 구도는 더 정확하지만, 가슴의 읽을 수 있는 ‘Ubik’ 표기가 명시적인 문자·로고 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정 장소의 특징과 찰리의 외형·외투, 머리 위 쓰레기와 파란 두 눈을 충실히 재현했으나, 발까지 보이는 넓은 구도와 더 일으켜진 몸통은 요구한 미디엄 지탱 자세에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 파란 두 눈은 정면에서 약간 아래쪽, 카메라 앞 쓰레기 방향을 향한다. 지정된 응시 대상은 없으므로 방향상 충돌은 없다. 두 팔은 좌우 아래로 뻗어 쓰레기 더미를 짚는다.",
        "built_space": "건물 내부나 고정 설비가 아니라 폐전자제품과 녹슨 금속판이 쌓인 노출된 쓰레기 더미다. 몸 뒤로 높은 쓰레기 층이 이어지고 전경 잔해가 하체를 가린다. 장소 참고의 파란 컨테이너와 오른쪽 굴착기는 이 구도에서 보이지 않아 정확한 지점의 일치는 확인하기 어렵지만, 더 좁은 구도 자체를 결함으로 볼 수는 없다. 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "로봇은 한 대뿐이며 추가 인물은 없다. 육중한 긴 팔, 샌드 베이지 장갑판, 흰 마스크형 얼굴과 점·선 무늬, 파랗게 켜진 두 눈이 보인다. 흙과 그물이 몸에 걸려 있고 머리 위에는 작은 폐기기와 비닐·천 조각이 얹혀 있다. 참고의 올리브색 외투는 보이지 않고 몸통 장갑이 노출되어 있다. 가슴에는 ‘Ubik’이 읽혀 최종 문자 금지 조건과 충돌한다. 주크박스·담요·음반은 보이지 않으며 이 인물 중심 구도에서 반드시 포함할 이유는 없다.",
        "hard_violations": [
         "가슴에 ‘Ubik’이라는 읽을 수 있는 문자와 로고가 노출되어, 읽을 수 있는 글자 및 로고를 금지한 조건을 위반한다."
        ],
        "physics": "양손의 손가락과 주먹 아래쪽이 좌우의 금속 잔해와 쓰레기에 닿아 상체를 지탱한다. 팔꿈치를 굽히고 몸통을 앞으로 기울여 반쯤 일어나는 동작이 물리적으로 읽힌다. 하체는 잔해 사이에 묻혀 있고, 머리 위 폐기물은 머리·어깨와 그물에 받쳐져 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 고개를 화면 오른쪽으로 약간 돌리고 아래쪽 전방을 바라본다. 두 파란 눈 모두 보이며 특정 인물이나 물체를 겨냥하는 동작은 없다. 팔은 좌우 아래의 잔해를 향해 내려가 있다.",
        "built_space": "왼쪽의 파란 컨테이너 한 개, 오른쪽 먼 능선의 집게형 굴착기 한 대, 뒤로 이어지는 금속 폐기물 언덕이 장소 참고와 대응한다. 전경에는 타이어 두 개가 좌우로 보이고 잔해가 찰리의 하체를 둘러싼다. 굴착기는 먼 배경 크기로 유지된다. 다만 하늘과 발까지 포함해 요구한 미디엄보다 넓게 잡혔다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "추가 인물 없이 찰리 한 대만 있다. 긴 기계 팔, 베이지 장갑, 흰 마스크형 얼굴, 올리브색 벨트 외투는 참고와 가깝다. 얼굴의 선 표현은 있으나 참고 및 지시의 점 무늬는 덜 뚜렷하다. 두 눈은 파랗게 빛나며 머리와 어깨에는 그물, 캔, 금속 조각이 걸쳐져 있다. 표면에 먼지와 마모가 보이고, 읽을 수 있는 가슴 글자는 노출되지 않는다. 주크박스와 음반 등은 프레임에 없으나 이 구도 밖 소품의 누락은 감점 사유가 아니다.",
        "hard_violations": [],
        "physics": "화면 오른쪽 손가락 끝은 전경 잔해에 닿고, 왼쪽 손의 접점은 쓰레기에 일부 가려져 있다. 아래쪽의 굽힌 다리와 보이는 발도 쓰레기 표면에 받쳐져 있어 몸이 떠 있지는 않다. 다만 몸통이 비교적 높이 세워져 양팔로 반쯤 일어난 상체를 버티는 순간보다는 더 일어난 단계로 읽힌다. 머리 위 물체들은 그물과 머리·어깨에 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.071
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.821
   },
   "violations": {
    "B": [
     "[gemini-pro] 금지된 읽을 수 있는 텍스트 노출 (가슴의 UBIK 로고)",
     "[gpt-high] 가슴에 ‘Ubik’이라는 읽을 수 있는 문자와 로고가 노출되어, 읽을 수 있는 글자 및 로고를 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 821
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "캐릭터의 복장과 지정된 매립지 장소를 충실히 재현했으며, 지시된 지탱 자세를 잘 표현함."
   },
   {
    "label": "B",
    "score": 821,
    "verdict_ko": "캐릭터의 핵심 복장이 누락되었고, 명시적으로 금지된 읽을 수 있는 텍스트가 렌더링되어 지시를 위반함.  ★위반: [gemini-pro] 금지된 읽을 수 있는 텍스트 노출 (가슴의 UBIK 로고) / [gpt-high] 가슴에 ‘Ubik’이라는 읽을 수 있는 문자와 로고가 노출되어, 읽을 수 있는 글자 및 로고를 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_dump_emergence_3bffab.png",
    "asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-923c-7a21-adaa-d6ae04dff6f3",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15__bgfirst_bg.png",
   "bg_asset_id": "8d75d911-3d18-4135-b6ab-fc26e495c901",
   "bg_record_key": "S6sh15::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "dump_emergence",
   "groupbg_asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S6sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:03:33.796880+00:00",
  "fingerprint": "1f30672f92248089423b945734847721548c3959c00889a2b1523c0d52a8d2ef",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S6sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S6sh15_sel.png",
  "source_sha256": "200e2edeb244b7dad6d5a3c12b8209cf5d758b077a4da8cead869ea2dcaf29d8",
  "file": "S6sh15_cine.png",
  "staged_sha256": "c3640d58a9174f1e21c9643ad4163e8a5f0dae7660863f7e3959f6170b7c4237",
  "latency_ms": 11818
 },
 "S6sh17::signage": {
  "fp": "456993398ebdc7bf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh17": {
  "input_fingerprint": "3d78fa2985d23aeb",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고릴라 형태의 찰리가 거대한 두 팔로 앰버의 작은 몸을 빈틈없이 감싸 안은 상태.\n\nLOCATION (lock): On the rubbish-strewn ground beside the robot's emergence point in the refugee settlement dump. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse around the pair (Surrounding the place where 찰리 emerged); used as A subdued lower and rear surround that preserves the harsh setting around the intimate contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Steady daytime ambient light and restrained blue eye illumination preserve tenderness without changing the established tonal treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same daylight, exposed refuse, and surrounding scrap heaps at the robot's emergence site. Exclude airborne rubbish and any debris still falling from the emergence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is upright with his arms closed in an embrace, retaining his blue-lit eyes, dirty aged casing, netting, and worn Ubik chest logo. The repaired jukebox, fitted horn speaker, LP, blanket, and scattered tools remain at the rubbish heap. 앰버: Amber has approached the robot's position and is caught in a tight embrace, looking bewildered; her mask and waist tool pouch remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고릴라 형태의 찰리가 거대한 두 팔로 앰버의 작은 몸을 빈틈없이 감싸 안은 상태.\n\nLOCATION (lock): On the rubbish-strewn ground beside the robot's emergence point in the refugee settlement dump. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse around the pair (Surrounding the place where 찰리 emerged); used as A subdued lower and rear surround that preserves the harsh setting around the intimate contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Steady daytime ambient light and restrained blue eye illumination preserve tenderness without changing the established tonal treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same daylight, exposed refuse, and surrounding scrap heaps at the robot's emergence site. Exclude airborne rubbish and any debris still falling from the emergence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is upright with his arms closed in an embrace, retaining his blue-lit eyes, dirty aged casing, netting, and worn Ubik chest logo. The repaired jukebox, fitted horn speaker, LP, blanket, and scattered tools remain at the rubbish heap. 앰버: Amber has approached the robot's position and is caught in a tight embrace, looking bewildered; her mask and waist tool pouch remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고릴라 형태의 찰리가 거대한 두 팔로 앰버의 작은 몸을 빈틈없이 감싸 안은 상태.\n\nLOCATION (lock): On the rubbish-strewn ground beside the robot's emergence point in the refugee settlement dump. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse around the pair (Surrounding the place where 찰리 emerged); used as A subdued lower and rear surround that preserves the harsh setting around the intimate contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Steady daytime ambient light and restrained blue eye illumination preserve tenderness without changing the established tonal treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same daylight, exposed refuse, and surrounding scrap heaps at the robot's emergence site. Exclude airborne rubbish and any debris still falling from the emergence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is upright with his arms closed in an embrace, retaining his blue-lit eyes, dirty aged casing, netting, and worn Ubik chest logo. The repaired jukebox, fitted horn speaker, LP, blanket, and scattered tools remain at the rubbish heap. 앰버: Amber has approached the robot's position and is caught in a tight embrace, looking bewildered; her mask and waist tool pouch remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "찰리는 고개를 숙여 앰버의 머리 쪽을 향하고, 앰버는 찰리를 올려다보지 않고 카메라 왼쪽을 바라본다. 두 팔과 손은 앰버의 가슴과 허리 앞쪽으로 모여 포옹한다. 두 나팔의 개구부는 모두 대체로 카메라 쪽을 향한다.",
    "built_space": "야외 쓰레기 더미와 녹슨 철판, 타이어가 두 인물을 둘러싸며 오른쪽 먼 배경에는 집게 굴착기 한 대가 있다. 오른쪽 전경에는 나팔이 달린 주크박스 한 대 외에 별도 나팔 축음기 한 대까지 놓였다. 두 인물의 발과 넓은 주변 바닥을 모두 보여주는 전신 구도여서, 쓰레기를 낮고 절제된 주변부로 두는 미디엄 숏 지시와 다르다.",
    "entities": "인물은 찰리와 앰버 두 명뿐이다. 앰버는 금발의 어린 여자아이로 둥근 얼굴, 남색 상의, 갈색 작업복, 허리 공구 주머니를 갖췄으나 참조의 이마 위 마스크를 입과 코에 착용했다. 찰리는 흰 기계 얼굴, 푸른 눈, 베이지 장갑과 녹색 코트는 유지하지만 이전 숏의 머리·어깨 그물과 부착된 깡통이 빠지고 외장도 더 깨끗하다. 주크박스, 부착 나팔, 담요, 공구 및 두 장의 음반이 보인다. 가슴 표식은 확인되지 않는다.",
    "hard_violations": [
     "지정된 주크박스와 부착 나팔 외에, 두 번째 나팔을 갖춘 독립 축음기를 추가했다."
    ],
    "physics": "찰리의 넓은 기계 발과 앰버의 부츠가 바닥에 닿아 체중을 지탱한다. 찰리의 양팔은 어깨와 팔꿈치에서 굽혀져 앰버의 몸 앞을 감싸며, 접촉과 관절 연결은 가능해 보인다. 주크박스와 축음기는 바닥에 놓이고 음반 하나는 주크박스에 기대며 다른 하나는 바닥에 놓여 있다. 떠 있거나 낙하하는 쓰레기는 보이지 않는다."
   },
   {
    "label": "A",
    "direction": "앰버는 고개와 눈을 위로 들어 찰리의 얼굴 쪽을 바라보며 당황한 표정을 짓는다. 찰리의 얼굴은 정면에서 약간 아래로 향하지만 눈이 앰버를 직접 응시하는지는 분명하지 않다. 양팔은 앰버의 가슴과 배 앞에서 안쪽으로 닫혀 몸을 밀착해 감싼다.",
    "built_space": "참조와 같은 녹슨 금속판, 폐가전 외함, 타이어가 쌓인 야외 쓰레기 더미다. 두 인물이 화면 중심을 크게 차지하고 하체가 잘리는 미디엄 구도이며, 폐기물은 뒤와 아래 주변부에 남는다. 별도의 건축 설비나 중복된 고정 시설은 보이지 않는다. 주크박스가 있을 주변 바닥은 프레임 밖이라 배치를 확인할 수 없다.",
    "entities": "찰리와 앰버만 등장한다. 찰리는 육중한 팔, 낡은 베이지 장갑판, 녹색 코트, 흰 기계 얼굴, 푸른 눈과 머리·어깨의 그물 및 깡통을 유지한다. 앰버는 참조와 유사한 금발의 어린 여자아이이며 둥근 얼굴, 갈색 작업복과 허리 공구 주머니가 보인다. 다만 이마의 장비는 참조의 호흡 마스크가 아니라 두 렌즈 고글이다. 가슴의 원형 장치가 두드러지지만 읽을 수 있는 로고는 없다. 주크박스, 나팔, 음반, 담요와 바닥 공구는 구도 밖이라 확인할 수 없다.",
    "hard_violations": [],
    "physics": "찰리의 팔은 연결된 관절에서 자연스럽게 굽혀져 앰버의 상체와 복부를 지지하고 감싼다. 앰버의 팔도 포옹 안에서 접혀 있으며 손과 몸의 접촉이 보인다. 두 인물의 하체는 화면 아래로 이어져 발의 접지는 확인할 수 없지만, 공중에 떠 있는 몸으로 묘사되지는 않았다. 그물은 머리와 어깨에 걸리고 깡통은 그물에 매달려 있으며 주변 폐기물은 더미에 얹혀 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전신과 소품까지 펼친 구도가 지정된 미디엄 숏을 벗어나고, 별도 축음기를 추가했으며 찰리의 그물도 누락했다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "미디엄 숏 안에서 두 거대한 팔의 밀착 포옹과 찰리의 낡은 외장·그물을 충실히 살렸지만, 앰버의 마스크를 고글로 바꾼 점은 불일치다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개를 숙여 앰버의 머리 쪽을 향하고, 앰버는 찰리를 올려다보지 않고 카메라 왼쪽을 바라본다. 두 팔과 손은 앰버의 가슴과 허리 앞쪽으로 모여 포옹한다. 두 나팔의 개구부는 모두 대체로 카메라 쪽을 향한다.",
        "built_space": "야외 쓰레기 더미와 녹슨 철판, 타이어가 두 인물을 둘러싸며 오른쪽 먼 배경에는 집게 굴착기 한 대가 있다. 오른쪽 전경에는 나팔이 달린 주크박스 한 대 외에 별도 나팔 축음기 한 대까지 놓였다. 두 인물의 발과 넓은 주변 바닥을 모두 보여주는 전신 구도여서, 쓰레기를 낮고 절제된 주변부로 두는 미디엄 숏 지시와 다르다.",
        "entities": "인물은 찰리와 앰버 두 명뿐이다. 앰버는 금발의 어린 여자아이로 둥근 얼굴, 남색 상의, 갈색 작업복, 허리 공구 주머니를 갖췄으나 참조의 이마 위 마스크를 입과 코에 착용했다. 찰리는 흰 기계 얼굴, 푸른 눈, 베이지 장갑과 녹색 코트는 유지하지만 이전 숏의 머리·어깨 그물과 부착된 깡통이 빠지고 외장도 더 깨끗하다. 주크박스, 부착 나팔, 담요, 공구 및 두 장의 음반이 보인다. 가슴 표식은 확인되지 않는다.",
        "hard_violations": [
         "지정된 주크박스와 부착 나팔 외에, 두 번째 나팔을 갖춘 독립 축음기를 추가했다."
        ],
        "physics": "찰리의 넓은 기계 발과 앰버의 부츠가 바닥에 닿아 체중을 지탱한다. 찰리의 양팔은 어깨와 팔꿈치에서 굽혀져 앰버의 몸 앞을 감싸며, 접촉과 관절 연결은 가능해 보인다. 주크박스와 축음기는 바닥에 놓이고 음반 하나는 주크박스에 기대며 다른 하나는 바닥에 놓여 있다. 떠 있거나 낙하하는 쓰레기는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "앰버는 고개와 눈을 위로 들어 찰리의 얼굴 쪽을 바라보며 당황한 표정을 짓는다. 찰리의 얼굴은 정면에서 약간 아래로 향하지만 눈이 앰버를 직접 응시하는지는 분명하지 않다. 양팔은 앰버의 가슴과 배 앞에서 안쪽으로 닫혀 몸을 밀착해 감싼다.",
        "built_space": "참조와 같은 녹슨 금속판, 폐가전 외함, 타이어가 쌓인 야외 쓰레기 더미다. 두 인물이 화면 중심을 크게 차지하고 하체가 잘리는 미디엄 구도이며, 폐기물은 뒤와 아래 주변부에 남는다. 별도의 건축 설비나 중복된 고정 시설은 보이지 않는다. 주크박스가 있을 주변 바닥은 프레임 밖이라 배치를 확인할 수 없다.",
        "entities": "찰리와 앰버만 등장한다. 찰리는 육중한 팔, 낡은 베이지 장갑판, 녹색 코트, 흰 기계 얼굴, 푸른 눈과 머리·어깨의 그물 및 깡통을 유지한다. 앰버는 참조와 유사한 금발의 어린 여자아이이며 둥근 얼굴, 갈색 작업복과 허리 공구 주머니가 보인다. 다만 이마의 장비는 참조의 호흡 마스크가 아니라 두 렌즈 고글이다. 가슴의 원형 장치가 두드러지지만 읽을 수 있는 로고는 없다. 주크박스, 나팔, 음반, 담요와 바닥 공구는 구도 밖이라 확인할 수 없다.",
        "hard_violations": [],
        "physics": "찰리의 팔은 연결된 관절에서 자연스럽게 굽혀져 앰버의 상체와 복부를 지지하고 감싼다. 앰버의 팔도 포옹 안에서 접혀 있으며 손과 몸의 접촉이 보인다. 두 인물의 하체는 화면 아래로 이어져 발의 접지는 확인할 수 없지만, 공중에 떠 있는 몸으로 묘사되지는 않았다. 그물은 머리와 어깨에 걸리고 깡통은 그물에 매달려 있으며 주변 폐기물은 더미에 얹혀 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전신과 소품까지 펼친 구도가 지정된 미디엄 숏을 벗어나고, 별도 축음기를 추가했으며 찰리의 그물도 누락했다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "미디엄 숏 안에서 두 거대한 팔의 밀착 포옹과 찰리의 낡은 외장·그물을 충실히 살렸지만, 앰버의 마스크를 고글로 바꾼 점은 불일치다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 고개를 숙여 앰버의 머리 쪽을 향하고, 앰버는 찰리를 올려다보지 않고 카메라 왼쪽을 바라본다. 두 팔과 손은 앰버의 가슴과 허리 앞쪽으로 모여 포옹한다. 두 나팔의 개구부는 모두 대체로 카메라 쪽을 향한다.",
        "built_space": "야외 쓰레기 더미와 녹슨 철판, 타이어가 두 인물을 둘러싸며 오른쪽 먼 배경에는 집게 굴착기 한 대가 있다. 오른쪽 전경에는 나팔이 달린 주크박스 한 대 외에 별도 나팔 축음기 한 대까지 놓였다. 두 인물의 발과 넓은 주변 바닥을 모두 보여주는 전신 구도여서, 쓰레기를 낮고 절제된 주변부로 두는 미디엄 숏 지시와 다르다.",
        "entities": "인물은 찰리와 앰버 두 명뿐이다. 앰버는 금발의 어린 여자아이로 둥근 얼굴, 남색 상의, 갈색 작업복, 허리 공구 주머니를 갖췄으나 참조의 이마 위 마스크를 입과 코에 착용했다. 찰리는 흰 기계 얼굴, 푸른 눈, 베이지 장갑과 녹색 코트는 유지하지만 이전 숏의 머리·어깨 그물과 부착된 깡통이 빠지고 외장도 더 깨끗하다. 주크박스, 부착 나팔, 담요, 공구 및 두 장의 음반이 보인다. 가슴 표식은 확인되지 않는다.",
        "hard_violations": [
         "지정된 주크박스와 부착 나팔 외에, 두 번째 나팔을 갖춘 독립 축음기를 추가했다."
        ],
        "physics": "찰리의 넓은 기계 발과 앰버의 부츠가 바닥에 닿아 체중을 지탱한다. 찰리의 양팔은 어깨와 팔꿈치에서 굽혀져 앰버의 몸 앞을 감싸며, 접촉과 관절 연결은 가능해 보인다. 주크박스와 축음기는 바닥에 놓이고 음반 하나는 주크박스에 기대며 다른 하나는 바닥에 놓여 있다. 떠 있거나 낙하하는 쓰레기는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "앰버는 고개와 눈을 위로 들어 찰리의 얼굴 쪽을 바라보며 당황한 표정을 짓는다. 찰리의 얼굴은 정면에서 약간 아래로 향하지만 눈이 앰버를 직접 응시하는지는 분명하지 않다. 양팔은 앰버의 가슴과 배 앞에서 안쪽으로 닫혀 몸을 밀착해 감싼다.",
        "built_space": "참조와 같은 녹슨 금속판, 폐가전 외함, 타이어가 쌓인 야외 쓰레기 더미다. 두 인물이 화면 중심을 크게 차지하고 하체가 잘리는 미디엄 구도이며, 폐기물은 뒤와 아래 주변부에 남는다. 별도의 건축 설비나 중복된 고정 시설은 보이지 않는다. 주크박스가 있을 주변 바닥은 프레임 밖이라 배치를 확인할 수 없다.",
        "entities": "찰리와 앰버만 등장한다. 찰리는 육중한 팔, 낡은 베이지 장갑판, 녹색 코트, 흰 기계 얼굴, 푸른 눈과 머리·어깨의 그물 및 깡통을 유지한다. 앰버는 참조와 유사한 금발의 어린 여자아이이며 둥근 얼굴, 갈색 작업복과 허리 공구 주머니가 보인다. 다만 이마의 장비는 참조의 호흡 마스크가 아니라 두 렌즈 고글이다. 가슴의 원형 장치가 두드러지지만 읽을 수 있는 로고는 없다. 주크박스, 나팔, 음반, 담요와 바닥 공구는 구도 밖이라 확인할 수 없다.",
        "hard_violations": [],
        "physics": "찰리의 팔은 연결된 관절에서 자연스럽게 굽혀져 앰버의 상체와 복부를 지지하고 감싼다. 앰버의 팔도 포옹 안에서 접혀 있으며 손과 몸의 접촉이 보인다. 두 인물의 하체는 화면 아래로 이어져 발의 접지는 확인할 수 없지만, 공중에 떠 있는 몸으로 묘사되지는 않았다. 그물은 머리와 어깨에 걸리고 깡통은 그물에 매달려 있으며 주변 폐기물은 더미에 얹혀 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "전신과 소품까지 펼친 구도가 지정된 미디엄 숏을 벗어나고, 별도 축음기를 추가했으며 찰리의 그물도 누락했다."
   },
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "미디엄 숏 안에서 두 거대한 팔의 밀착 포옹과 찰리의 낡은 외장·그물을 충실히 살렸지만, 앰버의 마스크를 고글로 바꾼 점은 불일치다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15_sel.png",
    "asset_id": "59f3bdaf-9622-4a6f-b71b-edc12df836f8",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-9712-7231-ba17-03c322c370b4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S6sh15"
  }
 },
 "S6sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:04:46.576832+00:00",
  "fingerprint": "d462121612a06ab041df348652b1132fbc86f2c8ed84f009c13086398b6075c6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S6sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S6sh17_sel.png",
  "source_sha256": "4ffa9c82b5a0820e750f63635903447f1bc2bb44a3ee83cabbcc94bdd71b2287",
  "file": "S6sh17_cine.png",
  "staged_sha256": "e70f688adf1477534de31427a4a58453150f89f1f308c5d5faf2ca8db5b8a88c",
  "latency_ms": 10686
 },
 "S6sh26::signage": {
  "fp": "4f5353d86025d5cd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::351d6ebea8b99a6a": {
  "subjects": [],
  "subject_text": "인천 난민촌 쓰레기장\n거대한 쓰레기 산이 솟은 야외 폐기물 지대. 고철과 폐가전, 로봇 부품이 뒤섞여 있고 입구에는 ‘난민 거주지역’ 표지판이 걸려 있다.",
  "identity": "canonical",
  "scope_id": "L154",
  "scope_role": "location_exterior",
  "scope_sha": "476e0d3d155a1022"
 },
 "S6sh26::bgfirst_bg": {
  "input_fingerprint": "2b0b5a907c39d06c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh26__bgfirst_bg.png",
  "asset_id": "dc27e10d-eecc-46d8-bfa1-0d822e7b99c4",
  "input_asset_ids": [
   "e7cb1336-68ab-42f8-8caa-a8db9daa34a5",
   "c7293dea-4a59-4086-b333-889f46e571b8"
  ]
 },
 "S6sh26": {
  "input_fingerprint": "2d3b2a02e1418ce4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is disguised in an old overcoat and hat over his dirty, aged gorilla-shaped body, with the netting and worn chest logo beneath the disguise. His blue-lit eyes remain unchanged, and the repaired jukebox was last left switched off. 앰버: Amber is walking away from the rubbish dump toward Pedro's home, still wearing her mask and waist tool pouch. 라울: Raul is walking toward Pedro's home after escaping notice.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is disguised in an old overcoat and hat over his dirty, aged gorilla-shaped body, with the netting and worn chest logo beneath the disguise. His blue-lit eyes remain unchanged, and the repaired jukebox was last left switched off. 앰버: Amber is walking away from the rubbish dump toward Pedro's home, still wearing her mask and waist tool pouch. 라울: Raul is walking toward Pedro's home after escaping notice.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is disguised in an old overcoat and hat over his dirty, aged gorilla-shaped body, with the netting and worn chest logo beneath the disguise. His blue-lit eyes remain unchanged, and the repaired jukebox was last left switched off. 앰버: Amber is walking away from the rubbish dump toward Pedro's home, still wearing her mask and waist tool pouch. 라울: Raul is walking toward Pedro's home after escaping notice.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh26__bgfirst_bg.png",
     "asset_id": "dc27e10d-eecc-46d8-bfa1-0d822e7b99c4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S6sh26.png",
     "asset_id": "e7cb1336-68ab-42f8-8caa-a8db9daa34a5",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_dump_sel.png",
     "asset_id": "c7293dea-4a59-4086-b333-889f46e571b8",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "B",
    "direction": "세 인물 모두 카메라에서 멀어져 중앙 난민촌 길로 내려간다. 앰버의 머리는 오른쪽 라울 쪽으로 조금 돌아가 있고, 라울과 찰리는 길 전방을 향한다. 페드로의 집 자체는 특정할 수 없지만 길이 주거지 내부로 이어진다. 겨누는 물건은 없다.",
    "built_space": "중앙의 내리막 흙길 양옆에 수십 채의 골판금·천막 판잣집이 빽빽하게 겹치며 전봇대가 길을 따라 이어진다. 세 주인공은 같은 길 위에 왼쪽부터 앰버, 라울, 찰리 순으로 있다. 참조의 한쪽에 집중된 거대한 쓰레기 사면 대신 양쪽 주택 밀집지와 먼 고층 아파트군을 보여 장소 배치의 일치도는 낮다. 참조의 출입문은 이 화면에 보이지 않으며, 중복 설비나 반사 문제는 없다.",
    "entities": "전경의 세 인물은 금발의 어린 앰버, 짙은 피부와 묶은 머리의 어린 라울, 거대한 기계 몸체의 찰리로 구별된다. 앰버는 남색 상의와 갈색 멜빵 작업복, 허리 공구 주머니, 옆으로 보이는 마스크를 착용해 참조와 가깝다. 라울의 회색 티셔츠와 반바지도 맞는다. 찰리는 모자와 낡은 외투를 입었으나 베이지색 어깨·팔 장갑과 등의 망이 크게 노출되어 위장 정도가 약하다. 후면이므로 얼굴·눈·가슴 로고와 두 아이의 세부 혈통은 검증할 수 없다. 주크박스는 없다. 길 깊숙한 중앙에 적어도 두 명의 추가 인물이 있고, 왼쪽 집에는 읽을 수 있는 한글 표기가 있다.",
    "hard_violations": [
     "중앙 원경의 길에 허용된 세 동행자 외의 추가 인물이 적어도 두 명 등장한다.",
     "왼쪽 판잣집 표면에 판독 가능한 한글 표기가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
    ],
    "physics": "앰버와 라울은 각각 한쪽 신발로 지면을 지지하면서 반대 발을 들어 걸어간다. 찰리의 큰 기계 발도 지면에 닿아 하중을 받으며 발밑에 먼지가 일어난다. 찰리의 보폭은 아이들보다 덜 뚜렷하지만 부유하지 않는다. 공구 주머니는 허리띠에, 모자는 머리에 지지되며 외투와 망은 몸에 걸려 있다."
   },
   {
    "label": "A",
    "direction": "세 인물은 등을 보인 채 중앙의 오르막 길을 따라 난민촌 안쪽으로 향한다. 머리와 몸통도 대체로 진행 방향을 향하며 카메라를 돌아보지 않는다. 길은 전경에서 언덕 위 주택 사이로 연속적으로 이어지지만 페드로의 집은 특정되지 않는다. 겨누는 물건은 없다.",
    "built_space": "하나의 넓은 흙길이 화면 아래에서 중앙 언덕 위로 올라가며, 양옆에 수십 채의 낮은 판금 주택과 쓰레기 사면이 있다. 왼쪽에는 골판금 울타리 한 줄, 언덕에는 여러 전봇대와 연결 전선이 보인다. 세 인물은 길 위에 앰버, 라울, 찰리 순으로 떨어져 걸으며 구조물과 충돌하지 않는다. 녹슨 판금, 폐기물과 흙길은 참조에 가깝지만, 참조의 오른쪽 대형 쓰레기 산과 왼쪽 도로라는 구체적 배치를 그대로 확인할 수는 없다. 중복 설비나 불가능한 반사는 없다.",
    "entities": "보이는 인물은 세 명뿐이다. 앰버는 금발의 어린 여자아이로 허리 주머니와 머리 뒤 마스크 끈을 갖췄지만, 회갈색 상의와 일반 바지는 참조의 남색 상의·멜빵 작업복과 다르다. 라울은 짙은 피부, 뒤로 묶은 머리, 회색 티셔츠와 반바지로 식별된다. 찰리는 낡은 긴 외투와 모자로 몸을 감췄고, 아래로 베이지색 기계 다리와 손이 드러난다. 얼굴·눈·가슴 로고·외투 안쪽 망은 후면과 옷에 가려져 검증 대상이 아니다. 주크박스나 판독 가능한 글자는 보이지 않는다.",
    "hard_violations": [],
    "physics": "앰버와 라울은 한쪽 발이 땅을 받치고 반대 발이 들린 보행 중간 자세다. 찰리도 한쪽 큰 발로 체중을 지지하면서 다른 다리를 전진시키며 상체를 약간 숙인다. 세 인물의 발밑 먼지는 걸음과 연결되고, 지지 없이 떠 있는 몸은 없다. 외투는 어깨에서 내려와 자연스럽게 주름지고 허리띠로 묶였으며, 공구 주머니도 허리에 고정되어 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "세 인물의 뒷모습 전신과 보행은 구현했지만, 길 안쪽에 추가 인물이 있고 왼쪽 판잣집에 판독 가능한 글자가 있어 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "세 동행자만 등장하는 후면 와이드숏, 흙먼지를 일으키는 보행과 외투로 위장한 찰리가 충실하며, 다만 앰버의 의상과 장소의 구체적 배치는 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "세 인물 모두 카메라에서 멀어져 중앙 난민촌 길로 내려간다. 앰버의 머리는 오른쪽 라울 쪽으로 조금 돌아가 있고, 라울과 찰리는 길 전방을 향한다. 페드로의 집 자체는 특정할 수 없지만 길이 주거지 내부로 이어진다. 겨누는 물건은 없다.",
        "built_space": "중앙의 내리막 흙길 양옆에 수십 채의 골판금·천막 판잣집이 빽빽하게 겹치며 전봇대가 길을 따라 이어진다. 세 주인공은 같은 길 위에 왼쪽부터 앰버, 라울, 찰리 순으로 있다. 참조의 한쪽에 집중된 거대한 쓰레기 사면 대신 양쪽 주택 밀집지와 먼 고층 아파트군을 보여 장소 배치의 일치도는 낮다. 참조의 출입문은 이 화면에 보이지 않으며, 중복 설비나 반사 문제는 없다.",
        "entities": "전경의 세 인물은 금발의 어린 앰버, 짙은 피부와 묶은 머리의 어린 라울, 거대한 기계 몸체의 찰리로 구별된다. 앰버는 남색 상의와 갈색 멜빵 작업복, 허리 공구 주머니, 옆으로 보이는 마스크를 착용해 참조와 가깝다. 라울의 회색 티셔츠와 반바지도 맞는다. 찰리는 모자와 낡은 외투를 입었으나 베이지색 어깨·팔 장갑과 등의 망이 크게 노출되어 위장 정도가 약하다. 후면이므로 얼굴·눈·가슴 로고와 두 아이의 세부 혈통은 검증할 수 없다. 주크박스는 없다. 길 깊숙한 중앙에 적어도 두 명의 추가 인물이 있고, 왼쪽 집에는 읽을 수 있는 한글 표기가 있다.",
        "hard_violations": [
         "중앙 원경의 길에 허용된 세 동행자 외의 추가 인물이 적어도 두 명 등장한다.",
         "왼쪽 판잣집 표면에 판독 가능한 한글 표기가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "앰버와 라울은 각각 한쪽 신발로 지면을 지지하면서 반대 발을 들어 걸어간다. 찰리의 큰 기계 발도 지면에 닿아 하중을 받으며 발밑에 먼지가 일어난다. 찰리의 보폭은 아이들보다 덜 뚜렷하지만 부유하지 않는다. 공구 주머니는 허리띠에, 모자는 머리에 지지되며 외투와 망은 몸에 걸려 있다."
       },
       {
        "label": "B",
        "direction": "세 인물은 등을 보인 채 중앙의 오르막 길을 따라 난민촌 안쪽으로 향한다. 머리와 몸통도 대체로 진행 방향을 향하며 카메라를 돌아보지 않는다. 길은 전경에서 언덕 위 주택 사이로 연속적으로 이어지지만 페드로의 집은 특정되지 않는다. 겨누는 물건은 없다.",
        "built_space": "하나의 넓은 흙길이 화면 아래에서 중앙 언덕 위로 올라가며, 양옆에 수십 채의 낮은 판금 주택과 쓰레기 사면이 있다. 왼쪽에는 골판금 울타리 한 줄, 언덕에는 여러 전봇대와 연결 전선이 보인다. 세 인물은 길 위에 앰버, 라울, 찰리 순으로 떨어져 걸으며 구조물과 충돌하지 않는다. 녹슨 판금, 폐기물과 흙길은 참조에 가깝지만, 참조의 오른쪽 대형 쓰레기 산과 왼쪽 도로라는 구체적 배치를 그대로 확인할 수는 없다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "보이는 인물은 세 명뿐이다. 앰버는 금발의 어린 여자아이로 허리 주머니와 머리 뒤 마스크 끈을 갖췄지만, 회갈색 상의와 일반 바지는 참조의 남색 상의·멜빵 작업복과 다르다. 라울은 짙은 피부, 뒤로 묶은 머리, 회색 티셔츠와 반바지로 식별된다. 찰리는 낡은 긴 외투와 모자로 몸을 감췄고, 아래로 베이지색 기계 다리와 손이 드러난다. 얼굴·눈·가슴 로고·외투 안쪽 망은 후면과 옷에 가려져 검증 대상이 아니다. 주크박스나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앰버와 라울은 한쪽 발이 땅을 받치고 반대 발이 들린 보행 중간 자세다. 찰리도 한쪽 큰 발로 체중을 지지하면서 다른 다리를 전진시키며 상체를 약간 숙인다. 세 인물의 발밑 먼지는 걸음과 연결되고, 지지 없이 떠 있는 몸은 없다. 외투는 어깨에서 내려와 자연스럽게 주름지고 허리띠로 묶였으며, 공구 주머니도 허리에 고정되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "세 인물의 뒷모습 전신과 보행은 구현했지만, 길 안쪽에 추가 인물이 있고 왼쪽 판잣집에 판독 가능한 글자가 있어 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "세 동행자만 등장하는 후면 와이드숏, 흙먼지를 일으키는 보행과 외투로 위장한 찰리가 충실하며, 다만 앰버의 의상과 장소의 구체적 배치는 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "세 인물 모두 카메라에서 멀어져 중앙 난민촌 길로 내려간다. 앰버의 머리는 오른쪽 라울 쪽으로 조금 돌아가 있고, 라울과 찰리는 길 전방을 향한다. 페드로의 집 자체는 특정할 수 없지만 길이 주거지 내부로 이어진다. 겨누는 물건은 없다.",
        "built_space": "중앙의 내리막 흙길 양옆에 수십 채의 골판금·천막 판잣집이 빽빽하게 겹치며 전봇대가 길을 따라 이어진다. 세 주인공은 같은 길 위에 왼쪽부터 앰버, 라울, 찰리 순으로 있다. 참조의 한쪽에 집중된 거대한 쓰레기 사면 대신 양쪽 주택 밀집지와 먼 고층 아파트군을 보여 장소 배치의 일치도는 낮다. 참조의 출입문은 이 화면에 보이지 않으며, 중복 설비나 반사 문제는 없다.",
        "entities": "전경의 세 인물은 금발의 어린 앰버, 짙은 피부와 묶은 머리의 어린 라울, 거대한 기계 몸체의 찰리로 구별된다. 앰버는 남색 상의와 갈색 멜빵 작업복, 허리 공구 주머니, 옆으로 보이는 마스크를 착용해 참조와 가깝다. 라울의 회색 티셔츠와 반바지도 맞는다. 찰리는 모자와 낡은 외투를 입었으나 베이지색 어깨·팔 장갑과 등의 망이 크게 노출되어 위장 정도가 약하다. 후면이므로 얼굴·눈·가슴 로고와 두 아이의 세부 혈통은 검증할 수 없다. 주크박스는 없다. 길 깊숙한 중앙에 적어도 두 명의 추가 인물이 있고, 왼쪽 집에는 읽을 수 있는 한글 표기가 있다.",
        "hard_violations": [
         "중앙 원경의 길에 허용된 세 동행자 외의 추가 인물이 적어도 두 명 등장한다.",
         "왼쪽 판잣집 표면에 판독 가능한 한글 표기가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "앰버와 라울은 각각 한쪽 신발로 지면을 지지하면서 반대 발을 들어 걸어간다. 찰리의 큰 기계 발도 지면에 닿아 하중을 받으며 발밑에 먼지가 일어난다. 찰리의 보폭은 아이들보다 덜 뚜렷하지만 부유하지 않는다. 공구 주머니는 허리띠에, 모자는 머리에 지지되며 외투와 망은 몸에 걸려 있다."
       },
       {
        "label": "A",
        "direction": "세 인물은 등을 보인 채 중앙의 오르막 길을 따라 난민촌 안쪽으로 향한다. 머리와 몸통도 대체로 진행 방향을 향하며 카메라를 돌아보지 않는다. 길은 전경에서 언덕 위 주택 사이로 연속적으로 이어지지만 페드로의 집은 특정되지 않는다. 겨누는 물건은 없다.",
        "built_space": "하나의 넓은 흙길이 화면 아래에서 중앙 언덕 위로 올라가며, 양옆에 수십 채의 낮은 판금 주택과 쓰레기 사면이 있다. 왼쪽에는 골판금 울타리 한 줄, 언덕에는 여러 전봇대와 연결 전선이 보인다. 세 인물은 길 위에 앰버, 라울, 찰리 순으로 떨어져 걸으며 구조물과 충돌하지 않는다. 녹슨 판금, 폐기물과 흙길은 참조에 가깝지만, 참조의 오른쪽 대형 쓰레기 산과 왼쪽 도로라는 구체적 배치를 그대로 확인할 수는 없다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "보이는 인물은 세 명뿐이다. 앰버는 금발의 어린 여자아이로 허리 주머니와 머리 뒤 마스크 끈을 갖췄지만, 회갈색 상의와 일반 바지는 참조의 남색 상의·멜빵 작업복과 다르다. 라울은 짙은 피부, 뒤로 묶은 머리, 회색 티셔츠와 반바지로 식별된다. 찰리는 낡은 긴 외투와 모자로 몸을 감췄고, 아래로 베이지색 기계 다리와 손이 드러난다. 얼굴·눈·가슴 로고·외투 안쪽 망은 후면과 옷에 가려져 검증 대상이 아니다. 주크박스나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앰버와 라울은 한쪽 발이 땅을 받치고 반대 발이 들린 보행 중간 자세다. 찰리도 한쪽 큰 발로 체중을 지지하면서 다른 다리를 전진시키며 상체를 약간 숙인다. 세 인물의 발밑 먼지는 걸음과 연결되고, 지지 없이 떠 있는 몸은 없다. 외투는 어깨에서 내려와 자연스럽게 주름지고 허리띠로 묶였으며, 공구 주머니도 허리에 고정되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "세 인물의 뒷모습 전신과 보행은 구현했지만, 길 안쪽에 추가 인물이 있고 왼쪽 판잣집에 판독 가능한 글자가 있어 명시적 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "세 동행자만 등장하는 후면 와이드숏, 흙먼지를 일으키는 보행과 외투로 위장한 찰리가 충실하며, 다만 앰버의 의상과 장소의 구체적 배치는 참조와 다르다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_dump_sel.png",
    "asset_id": "c7293dea-4a59-4086-b333-889f46e571b8",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-98c2-7257-ab74-6f0d8b129e34",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh26__bgfirst_bg.png",
   "bg_asset_id": "dc27e10d-eecc-46d8-bfa1-0d822e7b99c4",
   "bg_record_key": "S6sh26::bgfirst_bg",
   "chain_winner": true,
   "authority": "seed_bg"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S6sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:05:54.990612+00:00",
  "fingerprint": "37c516fadb2ea0e6273b43f59a5d51e0528d57ec7e6b2c724cfdfed044add083",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S6sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S6sh26_sel.png",
  "source_sha256": "cc5dff83192b5ba83a331b554a216a88857a14770daf3d48c7fe8fcac221793c",
  "file": "S6sh26_cine.png",
  "staged_sha256": "e0fa23c93301a1f3121ab489b7533d62df333f5f5f27bd9dc0da45bf04f78f00",
  "latency_ms": 11313
 },
 "S7sh11::signage": {
  "fp": "0b0e716578a800e7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::bb81f5972953eb12": {
  "subjects": [],
  "subject_text": "인천 난민촌 인공제방과 공사장\n바다를 막아선 거대한 콘크리트 제방과 보수 공사장. 벽 곳곳의 깊은 균열과 젖은 지지대, 적재된 보수 자재와 돌무더기가 보인다.",
  "identity": "canonical",
  "scope_id": "L156",
  "scope_role": "location_exterior",
  "scope_sha": "2bf570d14c3dfff2"
 },
 "groupbg::seawall_worksite": {
  "input_fingerprint": "aff63496c4b484a8",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "seawall_worksite",
    "tags": [
     "S13sh13",
     "S13sh20",
     "S33sh14",
     "S7sh11",
     "S7sh13",
     "S7sh15"
    ]
   },
   "context_sig": "f22b22c1cb13823e"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n박철진이 탑승한 전투 헬기 내부: 아래 지상을 지휘 통제할 수 있는 장비가 갖춰진 군용 헬기 조종석 옆자리. (특징: 복잡한 계기판과 유리창; 정보가 표시된 휴대용 태블릿 모니터; 아래로 넓게 펼쳐진 지상 시야) / 인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 제방 벽 곳곳이 심하게 금이 가 있다.\n- 어둠 속이지만 여기저기 제방 갈라진 틈이 보인다. 지지대에도 물이 흐르고.\n- 제방 근처에 있던 보수작업을 위한 중장비 기계들과 건설 장비들이 한 순간에 바닷물에 휩쓸려 간다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n박철진이 탑승한 전투 헬기 내부: 아래 지상을 지휘 통제할 수 있는 장비가 갖춰진 군용 헬기 조종석 옆자리. (특징: 복잡한 계기판과 유리창; 정보가 표시된 휴대용 태블릿 모니터; 아래로 넓게 펼쳐진 지상 시야) / 인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 제방 벽 곳곳이 심하게 금이 가 있다.\n- 어둠 속이지만 여기저기 제방 갈라진 틈이 보인다. 지지대에도 물이 흐르고.\n- 제방 근처에 있던 보수작업을 위한 중장비 기계들과 건설 장비들이 한 순간에 바닷물에 휩쓸려 간다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_seawall_worksite_44c24f.png",
  "asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63",
  "input_asset_ids": [
   "d68199d9-a355-482b-9cd4-df9585548d5e"
  ],
  "origin_tag": "S7sh11",
  "place_text": "On the ground beside the cracked seawall at the refugee settlement's suspended repair site.",
  "origin_inputs": {
   "place_text": "On the ground beside the cracked seawall at the refugee settlement's suspended repair site.",
   "time_of_day_en": "day",
   "conti_asset_id": "d68199d9-a355-482b-9cd4-df9585548d5e"
  }
 },
 "S7sh11::bgfirst_bg": {
  "input_fingerprint": "da74c208a9ecfc8e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11__bgfirst_bg.png",
  "asset_id": "0e42b760-1b39-40e2-a02b-2089485b8bc4",
  "input_asset_ids": [
   "d68199d9-a355-482b-9cd4-df9585548d5e",
   "13d1bb52-f02a-4d90-ae80-f1d788351b63"
  ]
 },
 "S7sh11": {
  "input_fingerprint": "15df0b7e5191ce38",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has severe cracks and visible water leakage, with puddles on the worksite ground. The military truck remains stopped in front of the forklift amid the halted repair equipment and materials. 미연: Miyeon is out of the forklift and standing among the gathered workers, urgently indicating the leaking wall.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has severe cracks and visible water leakage, with puddles on the worksite ground. The military truck remains stopped in front of the forklift amid the halted repair equipment and materials. 미연: Miyeon is out of the forklift and standing among the gathered workers, urgently indicating the leaking wall.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has severe cracks and visible water leakage, with puddles on the worksite ground. The military truck remains stopped in front of the forklift amid the halted repair equipment and materials. 미연: Miyeon is out of the forklift and standing among the gathered workers, urgently indicating the leaking wall.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11__bgfirst_bg.png",
     "asset_id": "0e42b760-1b39-40e2-a02b-2089485b8bc4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S7sh11.png",
     "asset_id": "d68199d9-a355-482b-9cd4-df9585548d5e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_seawall_worksite_44c24f.png",
     "asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연은 화면 왼쪽을 절박하게 바라보며, 왼팔을 뻗어 화면 오른쪽의 물이 새는 방벽을 가리킴.",
    "built_space": "물웅덩이가 있는 흙바닥으로 오른쪽에 금이 간 방벽이 있음. 배경에 군용 트럭과 지게차가 레퍼런스의 공간감에 맞게 배치됨.",
    "entities": "미연은 레퍼런스와 의상 및 외모가 일치하며 지시된 절박한 표정을 띠고 있음. 배경 및 전경에 지시된 작업자들이 배치됨.",
    "hard_violations": [
     "[gpt-high] 미연 외에 전경의 대화 상대와 배경 작업자 다섯 명을 추가했다."
    ],
    "physics": "바닥에 안정적으로 서서 체중을 지탱하고 있으며, 손을 뻗는 자세가 자연스러움."
   },
   {
    "label": "B",
    "direction": "미연은 화면 오른쪽을 응시하며, 왼팔을 뻗어 오른쪽의 방벽을 가리킴.",
    "built_space": "오른쪽에 금이 간 방벽이 있으나, 배경에서 지게차가 트럭보다 훨씬 앞에 배치되어 공간 구조가 레퍼런스와 다름.",
    "entities": "미연의 인상착의는 일치하나 눈을 부릅뜬 표정이 부족함. 배경의 작업자들이 경직된 상태로 서 있음.",
    "hard_violations": [
     "[gpt-high] 미연만 등장할 수 있는 장면에 다수의 추가 작업자를 배치했다."
    ],
    "physics": "바닥에 서서 팔을 뻗고 있으나, 배경 인물들의 정자세가 부자연스럽게 연출됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 프레이밍을 정확히 구현했으며, 눈을 부릅뜬 절박한 표정과 배경 요소의 배치가 지시문과 훌륭하게 일치함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업이 아닌 미디엄 샷으로 촬영되었고 절박한 표정 연출이 부족하며, 배경의 트럭과 지게차 배치 순서가 레퍼런스와 어긋남."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 화면 왼쪽을 절박하게 바라보며, 왼팔을 뻗어 화면 오른쪽의 물이 새는 방벽을 가리킴.",
        "built_space": "물웅덩이가 있는 흙바닥으로 오른쪽에 금이 간 방벽이 있음. 배경에 군용 트럭과 지게차가 레퍼런스의 공간감에 맞게 배치됨.",
        "entities": "미연은 레퍼런스와 의상 및 외모가 일치하며 지시된 절박한 표정을 띠고 있음. 배경 및 전경에 지시된 작업자들이 배치됨.",
        "hard_violations": [],
        "physics": "바닥에 안정적으로 서서 체중을 지탱하고 있으며, 손을 뻗는 자세가 자연스러움."
       },
       {
        "label": "B",
        "direction": "미연은 화면 오른쪽을 응시하며, 왼팔을 뻗어 오른쪽의 방벽을 가리킴.",
        "built_space": "오른쪽에 금이 간 방벽이 있으나, 배경에서 지게차가 트럭보다 훨씬 앞에 배치되어 공간 구조가 레퍼런스와 다름.",
        "entities": "미연의 인상착의는 일치하나 눈을 부릅뜬 표정이 부족함. 배경의 작업자들이 경직된 상태로 서 있음.",
        "hard_violations": [],
        "physics": "바닥에 서서 팔을 뻗고 있으나, 배경 인물들의 정자세가 부자연스럽게 연출됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 프레이밍을 정확히 구현했으며, 눈을 부릅뜬 절박한 표정과 배경 요소의 배치가 지시문과 훌륭하게 일치함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업이 아닌 미디엄 샷으로 촬영되었고 절박한 표정 연출이 부족하며, 배경의 트럭과 지게차 배치 순서가 레퍼런스와 어긋남."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 화면 왼쪽을 절박하게 바라보며, 왼팔을 뻗어 화면 오른쪽의 물이 새는 방벽을 가리킴.",
        "built_space": "물웅덩이가 있는 흙바닥으로 오른쪽에 금이 간 방벽이 있음. 배경에 군용 트럭과 지게차가 레퍼런스의 공간감에 맞게 배치됨.",
        "entities": "미연은 레퍼런스와 의상 및 외모가 일치하며 지시된 절박한 표정을 띠고 있음. 배경 및 전경에 지시된 작업자들이 배치됨.",
        "hard_violations": [],
        "physics": "바닥에 안정적으로 서서 체중을 지탱하고 있으며, 손을 뻗는 자세가 자연스러움."
       },
       {
        "label": "B",
        "direction": "미연은 화면 오른쪽을 응시하며, 왼팔을 뻗어 오른쪽의 방벽을 가리킴.",
        "built_space": "오른쪽에 금이 간 방벽이 있으나, 배경에서 지게차가 트럭보다 훨씬 앞에 배치되어 공간 구조가 레퍼런스와 다름.",
        "entities": "미연의 인상착의는 일치하나 눈을 부릅뜬 표정이 부족함. 배경의 작업자들이 경직된 상태로 서 있음.",
        "hard_violations": [],
        "physics": "바닥에 서서 팔을 뻗고 있으나, 배경 인물들의 정자세가 부자연스럽게 연출됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "누수 균열을 가리키는 동작은 맞지만, 금지된 작업자들을 추가했고 허리까지 보이는 넓은 구도로 절박한 얼굴의 클로즈업 지시를 놓쳤습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "더 크게 잡힌 얼굴과 부릅뜬 눈, 균열을 향한 손가락이 상대적으로 충실하지만, 추가 인물과 어깨너머 구도 때문에 이 후보 역시 부적합합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연의 검지는 화면 오른쪽 방벽의 큰 세로 균열을 향하며, 그 균열에서 물이 흐른다. 눈은 손가락과 대체로 같은 화면 오른쪽을 향하지만 정확히 균열을 응시하는지는 불분명하다. 뒤의 작업자들은 대체로 미연과 그녀가 가리키는 쪽을 본다.",
        "built_space": "오른쪽에 상단 난간이 있는 콘크리트 방벽 하나가 길게 이어지고, 하부에는 균열과 낙석, 바닥에는 물웅덩이가 보인다. 배경 중앙에는 지게차 한 대와 그 왼쪽에 일부 가려진 트럭 한 대가 보인다. 지상 시점과 재료는 장소 참조에 부합하지만, 방벽과 작업장 전체가 넓게 드러나 좁은 손상 구간만 얼굴 뒤에 두라는 구도가 아니다. 미연은 허리까지 보이며 작업자 무리가 왼쪽 뒤에 서 있다.",
        "entities": "미연은 검은 머리의 중년 한국인 여성으로 표현되었고, 체크 셔츠와 갈색 앞치마는 참조와 유사하다. 참조의 모자는 없고 얼굴의 세부 인상에도 차이가 있다. 심한 방벽 균열, 누수, 웅덩이와 수리 장비가 보인다. 미연 이외의 작업자 다수가 명확히 등장하여 인물 제한을 위반한다. 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [
         "미연만 등장할 수 있는 장면에 다수의 추가 작업자를 배치했다."
        ],
        "physics": "가리키는 손과 팔은 어깨까지 자연스럽게 이어져 있으며 자체 근육으로 지탱할 수 있는 자세다. 미연의 발은 프레임 밖이므로 접지를 직접 확인할 수 없지만 공중에 뜬 징후는 없다. 뒤쪽 인물들의 보이는 발은 지면에 닿아 있다. 장비와 낙석은 바닥에 놓여 있고, 물은 균열에서 아래로 흘러 웅덩이를 이룬다."
       },
       {
        "label": "B",
        "direction": "미연의 검지는 화면 오른쪽의 누수하는 큰 방벽 균열을 향한다. 부릅뜬 눈은 벽이 아니라 화면 왼쪽 전경 인물의 얼굴을 향하여 그에게 절박하게 호소하는 동작으로 읽힌다. 뒤쪽 작업자들은 미연과 전경 상대방 쪽을 바라본다.",
        "built_space": "오른쪽의 난간 달린 콘크리트 방벽, 하부 낙석, 물웅덩이, 오른쪽 아래의 적재 파이프와 받침대가 장소 참조와 유사하다. 중앙 배경에는 군용 트럭 한 대와 그 오른쪽 지게차 한 대가 정지해 있다. 지상 시점은 맞고 얼굴은 A보다 크게 잡혔지만, 왼쪽 전경 인물의 어깨가 화면을 가리는 대화 구도로 바뀌었다. 벽도 좁은 증거 구간을 넘어 상당히 넓게 노출된다.",
        "entities": "미연은 검은 머리의 중년 한국인 여성으로 보이며 정상적인 눈의 구조를 유지한 채 눈을 크게 뜨고 있다. 갈색 앞치마와 주머니의 연필은 참조와 유사하지만 모자가 없고 셔츠의 체크무늬는 참조보다 훨씬 희미하다. 누수 방벽과 웅덩이, 트럭, 지게차가 보인다. 왼쪽 전경 상대방 한 명과 뒤쪽 안전모 작업자 다섯 명이 추가되어 인물 제한을 위반한다. 의복과 안전모의 작은 표식은 읽을 수 없다.",
        "hard_violations": [
         "미연 외에 전경의 대화 상대와 배경 작업자 다섯 명을 추가했다."
        ],
        "physics": "미연은 상체를 앞으로 기울이고 한쪽 팔을 옆으로 뻗었으며, 손·손목·팔·어깨가 연결된 실행 가능한 자세다. 하체는 잘려 있어 발의 지지는 확인할 수 없지만 부유하는 모습은 아니다. 보이는 작업자의 발과 차량 바퀴는 지면에 놓여 있다. 파이프는 받침대와 적재물에 지지되고 낙석은 바닥에 놓여 있으며, 누수는 중력 방향으로 흐른다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "누수 균열을 가리키는 동작은 맞지만, 금지된 작업자들을 추가했고 허리까지 보이는 넓은 구도로 절박한 얼굴의 클로즈업 지시를 놓쳤습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "더 크게 잡힌 얼굴과 부릅뜬 눈, 균열을 향한 손가락이 상대적으로 충실하지만, 추가 인물과 어깨너머 구도 때문에 이 후보 역시 부적합합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "미연의 검지는 화면 오른쪽 방벽의 큰 세로 균열을 향하며, 그 균열에서 물이 흐른다. 눈은 손가락과 대체로 같은 화면 오른쪽을 향하지만 정확히 균열을 응시하는지는 불분명하다. 뒤의 작업자들은 대체로 미연과 그녀가 가리키는 쪽을 본다.",
        "built_space": "오른쪽에 상단 난간이 있는 콘크리트 방벽 하나가 길게 이어지고, 하부에는 균열과 낙석, 바닥에는 물웅덩이가 보인다. 배경 중앙에는 지게차 한 대와 그 왼쪽에 일부 가려진 트럭 한 대가 보인다. 지상 시점과 재료는 장소 참조에 부합하지만, 방벽과 작업장 전체가 넓게 드러나 좁은 손상 구간만 얼굴 뒤에 두라는 구도가 아니다. 미연은 허리까지 보이며 작업자 무리가 왼쪽 뒤에 서 있다.",
        "entities": "미연은 검은 머리의 중년 한국인 여성으로 표현되었고, 체크 셔츠와 갈색 앞치마는 참조와 유사하다. 참조의 모자는 없고 얼굴의 세부 인상에도 차이가 있다. 심한 방벽 균열, 누수, 웅덩이와 수리 장비가 보인다. 미연 이외의 작업자 다수가 명확히 등장하여 인물 제한을 위반한다. 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [
         "미연만 등장할 수 있는 장면에 다수의 추가 작업자를 배치했다."
        ],
        "physics": "가리키는 손과 팔은 어깨까지 자연스럽게 이어져 있으며 자체 근육으로 지탱할 수 있는 자세다. 미연의 발은 프레임 밖이므로 접지를 직접 확인할 수 없지만 공중에 뜬 징후는 없다. 뒤쪽 인물들의 보이는 발은 지면에 닿아 있다. 장비와 낙석은 바닥에 놓여 있고, 물은 균열에서 아래로 흘러 웅덩이를 이룬다."
       },
       {
        "label": "A",
        "direction": "미연의 검지는 화면 오른쪽의 누수하는 큰 방벽 균열을 향한다. 부릅뜬 눈은 벽이 아니라 화면 왼쪽 전경 인물의 얼굴을 향하여 그에게 절박하게 호소하는 동작으로 읽힌다. 뒤쪽 작업자들은 미연과 전경 상대방 쪽을 바라본다.",
        "built_space": "오른쪽의 난간 달린 콘크리트 방벽, 하부 낙석, 물웅덩이, 오른쪽 아래의 적재 파이프와 받침대가 장소 참조와 유사하다. 중앙 배경에는 군용 트럭 한 대와 그 오른쪽 지게차 한 대가 정지해 있다. 지상 시점은 맞고 얼굴은 A보다 크게 잡혔지만, 왼쪽 전경 인물의 어깨가 화면을 가리는 대화 구도로 바뀌었다. 벽도 좁은 증거 구간을 넘어 상당히 넓게 노출된다.",
        "entities": "미연은 검은 머리의 중년 한국인 여성으로 보이며 정상적인 눈의 구조를 유지한 채 눈을 크게 뜨고 있다. 갈색 앞치마와 주머니의 연필은 참조와 유사하지만 모자가 없고 셔츠의 체크무늬는 참조보다 훨씬 희미하다. 누수 방벽과 웅덩이, 트럭, 지게차가 보인다. 왼쪽 전경 상대방 한 명과 뒤쪽 안전모 작업자 다섯 명이 추가되어 인물 제한을 위반한다. 의복과 안전모의 작은 표식은 읽을 수 없다.",
        "hard_violations": [
         "미연 외에 전경의 대화 상대와 배경 작업자 다섯 명을 추가했다."
        ],
        "physics": "미연은 상체를 앞으로 기울이고 한쪽 팔을 옆으로 뻗었으며, 손·손목·팔·어깨가 연결된 실행 가능한 자세다. 하체는 잘려 있어 발의 지지는 확인할 수 없지만 부유하는 모습은 아니다. 보이는 작업자의 발과 차량 바퀴는 지면에 놓여 있다. 파이프는 받침대와 적재물에 지지되고 낙석은 바닥에 놓여 있으며, 누수는 중력 방향으로 흐른다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.238
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.988
   },
   "violations": {
    "B": [
     "[gpt-high] 미연만 등장할 수 있는 장면에 다수의 추가 작업자를 배치했다."
    ],
    "A": [
     "[gpt-high] 미연 외에 전경의 대화 상대와 배경 작업자 다섯 명을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 988
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 클로즈업 프레이밍을 정확히 구현했으며, 눈을 부릅뜬 절박한 표정과 배경 요소의 배치가 지시문과 훌륭하게 일치함.  ★위반: [gpt-high] 미연 외에 전경의 대화 상대와 배경 작업자 다섯 명을 추가했다."
   },
   {
    "label": "B",
    "score": 988,
    "verdict_ko": "클로즈업이 아닌 미디엄 샷으로 촬영되었고 절박한 표정 연출이 부족하며, 배경의 트럭과 지게차 배치 순서가 레퍼런스와 어긋남.  ★위반: [gpt-high] 미연만 등장할 수 있는 장면에 다수의 추가 작업자를 배치했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_seawall_worksite_44c24f.png",
    "asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-9c28-788b-94fb-832aa3f11bfc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11__bgfirst_bg.png",
   "bg_asset_id": "0e42b760-1b39-40e2-a02b-2089485b8bc4",
   "bg_record_key": "S7sh11::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "seawall_worksite",
   "groupbg_asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S7sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:07:16.420915+00:00",
  "fingerprint": "fbd816e22a3bbfe13f9247a1379d82bab2c09f1004c24ecdcc8f2c5191162295",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S7sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S7sh11_sel.png",
  "source_sha256": "1eff20b21e7ae368ea6ee0f745c62e205f63ec01fb862416eb6b890d3b9b1421",
  "file": "S7sh11_cine.png",
  "staged_sha256": "fa336d0f8e32ec077e9822f3b425ba031ef852e4a93396d8f3b84c445935334d",
  "latency_ms": 10637
 },
 "S7sh13::signage": {
  "fp": "20f2a3458a91bd2a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh13": {
  "input_fingerprint": "0fb371ca48f46f49",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 난민 남자의 어깨를 향해 거칠게 뻗은 경비병의 양손과 그 힘에 밀려 몸이 뒤로 크게 기울어진 난민 남자가 포착된 mid-impact의 정점.\n\nLOCATION (lock): In the muddy gathering area beside the seawall repair works, where guards confront the refugee workers. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Embankment repair area (Repair work has been ordered to stop) — Only a recessed portion of the work area remains visible behind the opposing bodies; used as Maintains location continuity while leaving the hand-to-shoulder contact unobstructed; Ground beneath the confrontation (The refugee has not yet fallen onto it) — Visible along the lower edge beneath the displaced bodies; used as Leaves a readable fall direction without prematurely showing the aftermath.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding daylight and controlled contrast so the impact registers through posture rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The severely cracked seawall continues to leak, and puddles remain on the ground. The military truck is still beside the stopped forklift and repair materials.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 난민 남자의 어깨를 향해 거칠게 뻗은 경비병의 양손과 그 힘에 밀려 몸이 뒤로 크게 기울어진 난민 남자가 포착된 mid-impact의 정점.\n\nLOCATION (lock): In the muddy gathering area beside the seawall repair works, where guards confront the refugee workers. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Embankment repair area (Repair work has been ordered to stop) — Only a recessed portion of the work area remains visible behind the opposing bodies; used as Maintains location continuity while leaving the hand-to-shoulder contact unobstructed; Ground beneath the confrontation (The refugee has not yet fallen onto it) — Visible along the lower edge beneath the displaced bodies; used as Leaves a readable fall direction without prematurely showing the aftermath.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding daylight and controlled contrast so the impact registers through posture rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The severely cracked seawall continues to leak, and puddles remain on the ground. The military truck is still beside the stopped forklift and repair materials.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 난민 남자의 어깨를 향해 거칠게 뻗은 경비병의 양손과 그 힘에 밀려 몸이 뒤로 크게 기울어진 난민 남자가 포착된 mid-impact의 정점.\n\nLOCATION (lock): In the muddy gathering area beside the seawall repair works, where guards confront the refugee workers. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Embankment repair area (Repair work has been ordered to stop) — Only a recessed portion of the work area remains visible behind the opposing bodies; used as Maintains location continuity while leaving the hand-to-shoulder contact unobstructed; Ground beneath the confrontation (The refugee has not yet fallen onto it) — Visible along the lower edge beneath the displaced bodies; used as Leaves a readable fall direction without prematurely showing the aftermath.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding daylight and controlled contrast so the impact registers through posture rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The severely cracked seawall continues to leak, and puddles remain on the ground. The military truck is still beside the stopped forklift and repair materials.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "경비병의 오른손이 난민의 가슴에 닿았으나 왼손은 허공을 향함.",
    "built_space": "방파제 수리 구역이며 지게차가 보이나 군용 트럭이 2대로 복제됨.",
    "entities": "경찰 제복을 입은 경비병과 다국적 난민 남성이 식별됨.",
    "hard_violations": [
     "[gemini-pro] 복제된 사물 (군용 트럭이 2대로 등장)",
     "[gemini-pro] 금지된 텍스트 노출 (제복에 '경찰' 글씨)",
     "[gpt-high] 왼쪽 전경에 장면 문장이 지정하지 않은 제삼자의 몸이 추가되어 있다.",
     "[gpt-high] 참조에서 한 대였던 군용 트럭이 배경에 두 대로 중복되어 있다."
    ],
    "physics": "난민은 한쪽 다리로 땅을 디디고 뒤로 넘어지며 경비병은 두 발로 서 있음."
   },
   {
    "label": "B",
    "direction": "미는 남성의 손이 난민의 어깨와 팔을 향해 뻗어 닿아 있음.",
    "built_space": "방파제 수리 구역이며 군용 트럭 1대와 지게차가 올바르게 배치됨.",
    "entities": "다국적 난민 남성은 나타나지만 상대가 경비병이 아닌 작업복 차림임.",
    "hard_violations": [
     "[gpt-high] 장면 문장이 보여주도록 지정하지 않은 작업자들이 주동작의 두 남자 뒤에 추가되어 있다."
    ],
    "physics": "난민은 다리 하나로 지탱하며 기울어지고 미는 남성은 두 발로 안정적으로 섬."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "경비병과 난민의 대립 구도는 적절하나, 복제된 군용 트럭과 제복의 금지된 텍스트 노출로 인해 실격입니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "공간 구성과 물리적 접촉은 양호하나, 지시된 경비병 대신 일반 작업자가 등장하여 캐릭터 구현이 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "경비병의 오른손이 난민의 가슴에 닿았으나 왼손은 허공을 향함.",
        "built_space": "방파제 수리 구역이며 지게차가 보이나 군용 트럭이 2대로 복제됨.",
        "entities": "경찰 제복을 입은 경비병과 다국적 난민 남성이 식별됨.",
        "hard_violations": [
         "복제된 사물 (군용 트럭이 2대로 등장)",
         "금지된 텍스트 노출 (제복에 '경찰' 글씨)"
        ],
        "physics": "난민은 한쪽 다리로 땅을 디디고 뒤로 넘어지며 경비병은 두 발로 서 있음."
       },
       {
        "label": "B",
        "direction": "미는 남성의 손이 난민의 어깨와 팔을 향해 뻗어 닿아 있음.",
        "built_space": "방파제 수리 구역이며 군용 트럭 1대와 지게차가 올바르게 배치됨.",
        "entities": "다국적 난민 남성은 나타나지만 상대가 경비병이 아닌 작업복 차림임.",
        "hard_violations": [],
        "physics": "난민은 다리 하나로 지탱하며 기울어지고 미는 남성은 두 발로 안정적으로 섬."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "경비병과 난민의 대립 구도는 적절하나, 복제된 군용 트럭과 제복의 금지된 텍스트 노출로 인해 실격입니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "공간 구성과 물리적 접촉은 양호하나, 지시된 경비병 대신 일반 작업자가 등장하여 캐릭터 구현이 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "경비병의 오른손이 난민의 가슴에 닿았으나 왼손은 허공을 향함.",
        "built_space": "방파제 수리 구역이며 지게차가 보이나 군용 트럭이 2대로 복제됨.",
        "entities": "경찰 제복을 입은 경비병과 다국적 난민 남성이 식별됨.",
        "hard_violations": [
         "복제된 사물 (군용 트럭이 2대로 등장)",
         "금지된 텍스트 노출 (제복에 '경찰' 글씨)"
        ],
        "physics": "난민은 한쪽 다리로 땅을 디디고 뒤로 넘어지며 경비병은 두 발로 서 있음."
       },
       {
        "label": "B",
        "direction": "미는 남성의 손이 난민의 어깨와 팔을 향해 뻗어 닿아 있음.",
        "built_space": "방파제 수리 구역이며 군용 트럭 1대와 지게차가 올바르게 배치됨.",
        "entities": "다국적 난민 남성은 나타나지만 상대가 경비병이 아닌 작업복 차림임.",
        "hard_violations": [],
        "physics": "난민은 다리 하나로 지탱하며 기울어지고 미는 남성은 두 발로 안정적으로 섬."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "불필요한 배경 인물 때문에 부적격이지만, B보다 미디엄 숏에 가깝고 양손의 어깨 접촉과 뒤로 밀리는 순간이 분명하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "추가 인물과 군용 트럭 중복이 있으며, 전신에 가까운 구도와 어깨에서 벗어난 손 위치가 지정된 충돌 순간을 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 남자는 오른쪽 난민 남자의 어깨 부위를 보며 두 팔을 뻗고, 손들은 어깨와 쇄골 부근에 닿아 있다. 난민의 머리와 상체는 화면 오른쪽으로 밀려 기울고 시선은 오른쪽 아래를 향한다. 요구된 밀침의 방향과 접촉 대상이 읽힌다.",
        "built_space": "오른쪽에는 균열과 누수가 있는 방조제 한 줄, 상단 난간, 하단의 배관 자재와 받침대가 있다. 뒤쪽에는 군용 트럭 한 대와 지게차 한 대가 보이며 진흙과 물웅덩이가 이어져 장소 연속성이 높다. 다만 두 인물이 왼쪽에 몰려 오른쪽 작업장과 지면이 넓게 드러나므로, 배경을 움푹 들어간 일부만 남기라는 지정보다 개방적이다.",
        "entities": "주동작의 인물은 짧은 검은 머리의 한국인으로 보이는 성인 남성과 짧은 곱슬머리의 어두운 피부를 가진 성인 난민 남성이다. 난민은 회색 작업복을 입었으며, 밀치는 남자도 회색 셔츠와 짙은 바지 차림이라 경비병이라는 신분은 뚜렷하지 않다. 뒤에는 안전모와 작업복을 입은 별도 인물들이 최소 세 명 보인다. 미연과 대머리 남성은 식별되지 않지만, 권위 있는 장면 문장이 요구하는 인물도 아니다. 판독 가능한 글자는 뚜렷하지 않다.",
        "hard_violations": [
         "장면 문장이 보여주도록 지정하지 않은 작업자들이 주동작의 두 남자 뒤에 추가되어 있다."
        ],
        "physics": "난민은 골반과 다리를 아래로 둔 채 상체가 뒤로 기울며, 경비병의 어깨 접촉이 그 움직임의 원인으로 보인다. 발은 프레임 밖이므로 접지를 확인할 수 없지만, 몸 전체가 공중에 뜬 것으로 보이지는 않는다. 밀치는 남자의 팔과 손도 몸에서 접촉점까지 이어져 있어 불가능한 지지나 해부학적 문제는 뚜렷하지 않다."
       },
       {
        "label": "B",
        "direction": "오른쪽 경비병은 왼쪽 난민을 바라보며 팔을 뻗는다. 위쪽 손은 난민의 어깨보다 위팔 쪽을 향하고, 아래쪽 손은 난민의 전완과 손목 부근에 있어 양손으로 어깨를 미는 접촉이 성립하지 않는다. 난민은 경비병 쪽을 보면서 몸을 왼쪽 뒤로 젖힌다.",
        "built_space": "오른쪽의 갈라진 방조제, 상단 난간, 누수, 배관 자재와 물웅덩이는 참조 장소와 이어진다. 그러나 왼쪽 뒤와 중앙 뒤에 군용 트럭이 각각 한 대씩, 총 두 대가 보이고 지게차는 한 대다. 난민의 부츠까지 포함하고 지면과 작업장을 넓게 보여주어 요구된 미디엄 숏보다 넓다.",
        "entities": "남색 제복과 모자를 착용한 한국인으로 보이는 성인 남성은 경비병으로 명확히 읽힌다. 난민은 짧은 곱슬머리와 수염이 있는 어두운 피부의 성인 남성이며 갈색 재킷, 청바지, 장갑과 부츠를 착용한다. 왼쪽 전경에는 회색 옷을 입은 제삼자의 몸이 크게 들어와 있다. 미연이나 대머리 남성으로 식별되는 인물은 없다. 참조에 한 대 있던 군용 트럭이 두 대로 늘었다.",
        "hard_violations": [
         "왼쪽 전경에 장면 문장이 지정하지 않은 제삼자의 몸이 추가되어 있다.",
         "참조에서 한 대였던 군용 트럭이 배경에 두 대로 중복되어 있다."
        ],
        "physics": "난민은 골반을 뒤로 빼고 무릎을 굽혔으며 오른쪽으로 뻗은 부츠는 들려 있다. 아래쪽 부츠는 진흙 표면에 매우 가깝지만 확실한 접지는 판별하기 어렵다. 경비병의 밀침이 후퇴의 원인으로 제시되므로 공중 부양이라고 단정할 수는 없으나, 어깨의 확실한 접촉과 하체의 지지가 모두 A보다 불명확하다. 경비병은 하체를 낮추고 앞으로 체중을 싣고 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "불필요한 배경 인물 때문에 부적격이지만, B보다 미디엄 숏에 가깝고 양손의 어깨 접촉과 뒤로 밀리는 순간이 분명하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "추가 인물과 군용 트럭 중복이 있으며, 전신에 가까운 구도와 어깨에서 벗어난 손 위치가 지정된 충돌 순간을 약화한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 남자는 오른쪽 난민 남자의 어깨 부위를 보며 두 팔을 뻗고, 손들은 어깨와 쇄골 부근에 닿아 있다. 난민의 머리와 상체는 화면 오른쪽으로 밀려 기울고 시선은 오른쪽 아래를 향한다. 요구된 밀침의 방향과 접촉 대상이 읽힌다.",
        "built_space": "오른쪽에는 균열과 누수가 있는 방조제 한 줄, 상단 난간, 하단의 배관 자재와 받침대가 있다. 뒤쪽에는 군용 트럭 한 대와 지게차 한 대가 보이며 진흙과 물웅덩이가 이어져 장소 연속성이 높다. 다만 두 인물이 왼쪽에 몰려 오른쪽 작업장과 지면이 넓게 드러나므로, 배경을 움푹 들어간 일부만 남기라는 지정보다 개방적이다.",
        "entities": "주동작의 인물은 짧은 검은 머리의 한국인으로 보이는 성인 남성과 짧은 곱슬머리의 어두운 피부를 가진 성인 난민 남성이다. 난민은 회색 작업복을 입었으며, 밀치는 남자도 회색 셔츠와 짙은 바지 차림이라 경비병이라는 신분은 뚜렷하지 않다. 뒤에는 안전모와 작업복을 입은 별도 인물들이 최소 세 명 보인다. 미연과 대머리 남성은 식별되지 않지만, 권위 있는 장면 문장이 요구하는 인물도 아니다. 판독 가능한 글자는 뚜렷하지 않다.",
        "hard_violations": [
         "장면 문장이 보여주도록 지정하지 않은 작업자들이 주동작의 두 남자 뒤에 추가되어 있다."
        ],
        "physics": "난민은 골반과 다리를 아래로 둔 채 상체가 뒤로 기울며, 경비병의 어깨 접촉이 그 움직임의 원인으로 보인다. 발은 프레임 밖이므로 접지를 확인할 수 없지만, 몸 전체가 공중에 뜬 것으로 보이지는 않는다. 밀치는 남자의 팔과 손도 몸에서 접촉점까지 이어져 있어 불가능한 지지나 해부학적 문제는 뚜렷하지 않다."
       },
       {
        "label": "A",
        "direction": "오른쪽 경비병은 왼쪽 난민을 바라보며 팔을 뻗는다. 위쪽 손은 난민의 어깨보다 위팔 쪽을 향하고, 아래쪽 손은 난민의 전완과 손목 부근에 있어 양손으로 어깨를 미는 접촉이 성립하지 않는다. 난민은 경비병 쪽을 보면서 몸을 왼쪽 뒤로 젖힌다.",
        "built_space": "오른쪽의 갈라진 방조제, 상단 난간, 누수, 배관 자재와 물웅덩이는 참조 장소와 이어진다. 그러나 왼쪽 뒤와 중앙 뒤에 군용 트럭이 각각 한 대씩, 총 두 대가 보이고 지게차는 한 대다. 난민의 부츠까지 포함하고 지면과 작업장을 넓게 보여주어 요구된 미디엄 숏보다 넓다.",
        "entities": "남색 제복과 모자를 착용한 한국인으로 보이는 성인 남성은 경비병으로 명확히 읽힌다. 난민은 짧은 곱슬머리와 수염이 있는 어두운 피부의 성인 남성이며 갈색 재킷, 청바지, 장갑과 부츠를 착용한다. 왼쪽 전경에는 회색 옷을 입은 제삼자의 몸이 크게 들어와 있다. 미연이나 대머리 남성으로 식별되는 인물은 없다. 참조에 한 대 있던 군용 트럭이 두 대로 늘었다.",
        "hard_violations": [
         "왼쪽 전경에 장면 문장이 지정하지 않은 제삼자의 몸이 추가되어 있다.",
         "참조에서 한 대였던 군용 트럭이 배경에 두 대로 중복되어 있다."
        ],
        "physics": "난민은 골반을 뒤로 빼고 무릎을 굽혔으며 오른쪽으로 뻗은 부츠는 들려 있다. 아래쪽 부츠는 진흙 표면에 매우 가깝지만 확실한 접지는 판별하기 어렵다. 경비병의 밀침이 후퇴의 원인으로 제시되므로 공중 부양이라고 단정할 수는 없으나, 어깨의 확실한 접촉과 하체의 지지가 모두 A보다 불명확하다. 경비병은 하체를 낮추고 앞으로 체중을 싣고 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.1,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.85,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 복제된 사물 (군용 트럭이 2대로 등장)",
     "[gemini-pro] 금지된 텍스트 노출 (제복에 '경찰' 글씨)",
     "[gpt-high] 왼쪽 전경에 장면 문장이 지정하지 않은 제삼자의 몸이 추가되어 있다.",
     "[gpt-high] 참조에서 한 대였던 군용 트럭이 배경에 두 대로 중복되어 있다."
    ],
    "B": [
     "[gpt-high] 장면 문장이 보여주도록 지정하지 않은 작업자들이 주동작의 두 남자 뒤에 추가되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 850,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 850,
    "verdict_ko": "경비병과 난민의 대립 구도는 적절하나, 복제된 군용 트럭과 제복의 금지된 텍스트 노출로 인해 실격입니다.  ★위반: [gemini-pro] 복제된 사물 (군용 트럭이 2대로 등장) / [gemini-pro] 금지된 텍스트 노출 (제복에 '경찰' 글씨) / [gpt-high] 왼쪽 전경에 장면 문장이 지정하지 않은 제삼자의 몸이 추가되어 있다. / [gpt-high] 참조에서 한 대였던 군용 트럭이 배경에 두 대로 중복되어 있다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "공간 구성과 물리적 접촉은 양호하나, 지시된 경비병 대신 일반 작업자가 등장하여 캐릭터 구현이 아쉽습니다.  ★위반: [gpt-high] 장면 문장이 보여주도록 지정하지 않은 작업자들이 주동작의 두 남자 뒤에 추가되어 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11_sel.png",
    "asset_id": "92376a28-6a24-41e0-8ef5-2c96ea1c4239",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 50대 대머리 한국인 남성: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1399111>",
    "asset_id": "41215473-9904-471e-81f7-2f6c01467eee",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-a0f1-7a4b-a7df-414533825302",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S7sh11"
  }
 },
 "S7sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:08:55.948418+00:00",
  "fingerprint": "1cc4d5f97473944ce78cf15529f9053443fe8b4c0915c237634177e89e77b6cb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S7sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S7sh13_sel.png",
  "source_sha256": "888a59a0284b48716d54012b8fdfb82cd26966af0bf28d8c7119e4022822b467",
  "file": "S7sh13_cine.png",
  "staged_sha256": "0885d74d35623fea2dd3620e0fb071da307d4cb6e4d872c60566881291fe80c9",
  "latency_ms": 10589
 },
 "S7sh15::signage": {
  "fp": "a5dd6b91acf402d1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh15": {
  "input_fingerprint": "a662c22d0832fd46",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성난 표정으로 몰려드는 사람들을 등진 채, 트럭 조수석 안으로 다급히 상체를 반쯤 밀어 넣은 대머리 남자의 뒷모습.\n\nLOCATION (lock): At the open passenger doorway of a military truck parked on the muddy seawall repair site, viewed from outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open military truck passenger entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Military truck passenger entrance (Open for the administrator to climb inside) — The passenger-side opening is seen diagonally past his back, with the cabin extending beyond the right frame edge; used as Creates the destination of his retreat while remaining a cropped, subordinate portion of the frame; Approach beside the truck (People are converging on the administrator) — The open approach occupies the left side and leads diagonally toward the passenger entrance; used as Keeps the pursuing crowd and the retreat connected within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same daytime ambient illumination with readable separation between the retreating figure, the cabin opening, and the crowd.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The military truck remains at the departure point beside the halted construction site. The seawall's severe cracks, continuing leaks, and ground puddles remain unchanged. 50대 대머리 한국인 남성: The frightened administrator is hurriedly climbing into the truck; his shoes remain puddle-soiled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성난 표정으로 몰려드는 사람들을 등진 채, 트럭 조수석 안으로 다급히 상체를 반쯤 밀어 넣은 대머리 남자의 뒷모습.\n\nLOCATION (lock): At the open passenger doorway of a military truck parked on the muddy seawall repair site, viewed from outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open military truck passenger entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Military truck passenger entrance (Open for the administrator to climb inside) — The passenger-side opening is seen diagonally past his back, with the cabin extending beyond the right frame edge; used as Creates the destination of his retreat while remaining a cropped, subordinate portion of the frame; Approach beside the truck (People are converging on the administrator) — The open approach occupies the left side and leads diagonally toward the passenger entrance; used as Keeps the pursuing crowd and the retreat connected within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same daytime ambient illumination with readable separation between the retreating figure, the cabin opening, and the crowd.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The military truck remains at the departure point beside the halted construction site. The seawall's severe cracks, continuing leaks, and ground puddles remain unchanged. 50대 대머리 한국인 남성: The frightened administrator is hurriedly climbing into the truck; his shoes remain puddle-soiled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성난 표정으로 몰려드는 사람들을 등진 채, 트럭 조수석 안으로 다급히 상체를 반쯤 밀어 넣은 대머리 남자의 뒷모습.\n\nLOCATION (lock): At the open passenger doorway of a military truck parked on the muddy seawall repair site, viewed from outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open military truck passenger entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Military truck passenger entrance (Open for the administrator to climb inside) — The passenger-side opening is seen diagonally past his back, with the cabin extending beyond the right frame edge; used as Creates the destination of his retreat while remaining a cropped, subordinate portion of the frame; Approach beside the truck (People are converging on the administrator) — The open approach occupies the left side and leads diagonally toward the passenger entrance; used as Keeps the pursuing crowd and the retreat connected within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same daytime ambient illumination with readable separation between the retreating figure, the cabin opening, and the crowd.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The military truck remains at the departure point beside the halted construction site. The seawall's severe cracks, continuing leaks, and ground puddles remain unchanged. 50대 대머리 한국인 남성: The frightened administrator is hurriedly climbing into the truck; his shoes remain puddle-soiled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "군중의 시선과 움직임이 트럭 앞의 남자를 향하고, 남자는 트럭 내부를 향하고 있음.",
    "built_space": "우측에 트럭이 위치하나, 트럭의 앞부분이 우측 바위벽의 텍스처와 완전히 융합되어 공간이 왜곡됨.",
    "entities": "50대 대머리 남성의 외모와 작업복을 입은 다국적 노동자 무리가 지시대로 등장함.",
    "hard_violations": [
     "[gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 인물",
     "[gemini-pro] 트럭과 바위벽이 융합된 물리적으로 불가능한 구조",
     "[gpt-high] 오른쪽 방조제·바위·배관 영역을 트럭 측면과 바퀴 옆에 붙여 넣은 듯한 합성 경계가 있어, 단일 실사 프레임이 아닌 콜라주성 공간으로 나타난다."
    ],
    "physics": "남자가 양손으로 문틀과 내부를 잡고 있으나, 두 발이 허공에 뜬 채로 체중을 지탱하는 곳이 없음."
   },
   {
    "label": "B",
    "direction": "성난 군중이 남자를 향해 몰려오고, 남자는 다급히 트럭 안쪽으로 몸을 기울이고 있음.",
    "built_space": "트럭과 방파제가 분리되어 배치되었으나, 열린 문 안쪽에 운전대가 위치해 조수석이라는 설정과 어긋남.",
    "entities": "프롬프트의 대머리 남성과 이전 샷의 노동자(특히 앞쪽 인물)가 인상착의에 맞춰 정확히 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 인물",
     "[gpt-high] 기존 장면의 군용 트럭 외에 왼쪽 배경에 두 번째 군용 트럭을 추가해 차량을 중복 생성했다."
    ],
    "physics": "남자가 손잡이를 잡고 있으나, 두 발이 지면이나 트럭 발판에 닿지 않고 허공에 매달려 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물이 허공에 떠 있고 트럭이 바위벽과 융합되는 등 치명적인 물리적 오류로 인해 프롬프트를 실패했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물과 군중의 묘사는 우수하나, 남성의 발이 지탱될 곳 없이 허공에 떠 있는 치명적인 물리적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "군중의 시선과 움직임이 트럭 앞의 남자를 향하고, 남자는 트럭 내부를 향하고 있음.",
        "built_space": "우측에 트럭이 위치하나, 트럭의 앞부분이 우측 바위벽의 텍스처와 완전히 융합되어 공간이 왜곡됨.",
        "entities": "50대 대머리 남성의 외모와 작업복을 입은 다국적 노동자 무리가 지시대로 등장함.",
        "hard_violations": [
         "아무것도 지탱하지 않는 허공에 뜬 인물",
         "트럭과 바위벽이 융합된 물리적으로 불가능한 구조"
        ],
        "physics": "남자가 양손으로 문틀과 내부를 잡고 있으나, 두 발이 허공에 뜬 채로 체중을 지탱하는 곳이 없음."
       },
       {
        "label": "B",
        "direction": "성난 군중이 남자를 향해 몰려오고, 남자는 다급히 트럭 안쪽으로 몸을 기울이고 있음.",
        "built_space": "트럭과 방파제가 분리되어 배치되었으나, 열린 문 안쪽에 운전대가 위치해 조수석이라는 설정과 어긋남.",
        "entities": "프롬프트의 대머리 남성과 이전 샷의 노동자(특히 앞쪽 인물)가 인상착의에 맞춰 정확히 묘사됨.",
        "hard_violations": [
         "아무것도 지탱하지 않는 허공에 뜬 인물"
        ],
        "physics": "남자가 손잡이를 잡고 있으나, 두 발이 지면이나 트럭 발판에 닿지 않고 허공에 매달려 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물이 허공에 떠 있고 트럭이 바위벽과 융합되는 등 치명적인 물리적 오류로 인해 프롬프트를 실패했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물과 군중의 묘사는 우수하나, 남성의 발이 지탱될 곳 없이 허공에 떠 있는 치명적인 물리적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "군중의 시선과 움직임이 트럭 앞의 남자를 향하고, 남자는 트럭 내부를 향하고 있음.",
        "built_space": "우측에 트럭이 위치하나, 트럭의 앞부분이 우측 바위벽의 텍스처와 완전히 융합되어 공간이 왜곡됨.",
        "entities": "50대 대머리 남성의 외모와 작업복을 입은 다국적 노동자 무리가 지시대로 등장함.",
        "hard_violations": [
         "아무것도 지탱하지 않는 허공에 뜬 인물",
         "트럭과 바위벽이 융합된 물리적으로 불가능한 구조"
        ],
        "physics": "남자가 양손으로 문틀과 내부를 잡고 있으나, 두 발이 허공에 뜬 채로 체중을 지탱하는 곳이 없음."
       },
       {
        "label": "B",
        "direction": "성난 군중이 남자를 향해 몰려오고, 남자는 다급히 트럭 안쪽으로 몸을 기울이고 있음.",
        "built_space": "트럭과 방파제가 분리되어 배치되었으나, 열린 문 안쪽에 운전대가 위치해 조수석이라는 설정과 어긋남.",
        "entities": "프롬프트의 대머리 남성과 이전 샷의 노동자(특히 앞쪽 인물)가 인상착의에 맞춰 정확히 묘사됨.",
        "hard_violations": [
         "아무것도 지탱하지 않는 허공에 뜬 인물"
        ],
        "physics": "남자가 손잡이를 잡고 있으나, 두 발이 지면이나 트럭 발판에 닿지 않고 허공에 매달려 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "군중의 추격과 승차 지지는 비교적 자연스럽지만, 군용 트럭을 한 대 더 만들었고 남자의 상체가 조수석 안으로 반쯤 들어간 핵심 순간도 부족하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "남자의 뒷모습과 접근하는 군중은 맞지만, 방조제와 배관이 트럭 오른쪽에 붙여진 합성 경계로 나타나 단일한 실제 공간이라는 조건을 크게 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 군중은 오른쪽 남자를 바라보며 손을 뻗고 접근한다. 남자의 몸은 차량 입구를 향하지만 머리는 왼쪽 아래 군중 쪽으로 돌아가 옆얼굴이 보인다. 등을 완전히 돌린 채 상체를 실내로 밀어 넣는 동작보다는 문밖 발판에서 뒤를 살피는 순간에 가깝다.",
        "built_space": "오른쪽에 열린 차량 문 하나, 출입구 하나, 거울 하나, 좌석 일부와 운전대 하나가 보인다. 출입구는 중경의 작은 요소가 아니라 전경을 크게 차지한다. 운전대가 입구 바로 안쪽에 있어 조수석보다 운전석 입구처럼 읽힌다. 왼쪽 뒤에는 별도의 군용 트럭 한 대와 지게차 한 대가 있으며, 갈라진 콘크리트 방조제와 진흙 웅덩이는 참고 장소의 재질을 따른다.",
        "entities": "주인공은 정수리가 벗겨지고 옆머리가 남은 중년 동아시아계 남성으로, 참고의 남색 체크 재킷과 어두운 바지를 유지한다. 얼굴이 일부만 보여 정확한 동일 인물 여부는 제한적으로만 확인된다. 신발과 바짓단에는 진흙이 묻어 있다. 회색·남색 작업복의 성난 군중은 샷 텍스트에 명시된 사람들에 해당한다. 군용 트럭은 두 대로 늘어났으며, 지게차는 참고와 같은 종류다. 뚜렷하게 판독되는 문구는 보이지 않는다.",
        "hard_violations": [
         "기존 장면의 군용 트럭 외에 왼쪽 배경에 두 번째 군용 트럭을 추가해 차량을 중복 생성했다."
        ],
        "physics": "남자는 오른손으로 출입구 손잡이를 잡고 왼손으로 열린 문의 윗부분을 짚는다. 오른발은 낮은 차량 발판에 걸려 있고 두 손도 몸을 지탱하므로 공중에 무지지 상태로 떠 있는 것은 아니다. 다만 두 발을 모은 자세와 거의 실외에 남은 상체는 급히 몸을 반쯤 밀어 넣는 동작을 약하게 표현한다. 군중은 진흙 바닥에 발을 딛거나 보행 중이며 지지 관계가 자연스럽다."
       },
       {
        "label": "B",
        "direction": "남자는 군중을 등지고 머리와 양팔을 차량 내부 쪽으로 향한다. 왼쪽 군중도 남자와 출입구를 바라보며 다가와 이동 목표는 맞는다. 앞줄 일부는 입을 벌리고 있으나, 전체적으로 격하게 몰려들기보다는 걸어오는 모습이다.",
        "built_space": "열린 문 하나와 거울 하나, 출입구 하나, 앞바퀴 하나가 보이며 차량은 오른쪽 경계에서 잘린다. 그러나 출입구와 남자가 전경에서 크게 보여 지정된 중경 배치와는 차이가 있다. 오른쪽에는 참고 이미지의 방조제 균열·바위·배관 구역이 차량 측면을 덮듯 이어지고, 바퀴와 차체 옆에서 서로 다른 공간의 경계가 생긴다. 이를 실제 차체 표면이나 자연스러운 반사로 설명하기 어렵다.",
        "entities": "대머리 중년 남성의 후두부, 남색 체크 재킷, 어두운 바지와 진흙 묻은 신발은 참고 인물의 보이는 특징과 부합한다. 얼굴은 뒤돌아 있어 나이와 동일 얼굴을 세밀하게 검증할 수 없다. 왼쪽에는 회색·남색 작업복과 일부 안전모를 착용한 군중이 보인다. 차량과 방조제의 개별 질감은 사실적이지만 두 대상이 하나의 물리적 공간에 놓인 방식은 성립하지 않는다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "오른쪽 방조제·바위·배관 영역을 트럭 측면과 바퀴 옆에 붙여 넣은 듯한 합성 경계가 있어, 단일 실사 프레임이 아닌 콜라주성 공간으로 나타난다."
        ],
        "physics": "남자의 오른발은 차량 발판에 걸려 있고 왼손은 문 윗부분을 짚으며 오른팔은 실내로 뻗어 있어, 왼발이 공중에 들린 승차 동작 자체에는 지지가 있다. 군중의 발도 지면과 접촉한다. 다만 오른쪽 차량과 방조제가 겹친 구역은 고체 차체, 벽, 배관이 차지하는 공간을 일관되게 설명할 수 없어 물리적 배치가 깨진다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "군중의 추격과 승차 지지는 비교적 자연스럽지만, 군용 트럭을 한 대 더 만들었고 남자의 상체가 조수석 안으로 반쯤 들어간 핵심 순간도 부족하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "남자의 뒷모습과 접근하는 군중은 맞지만, 방조제와 배관이 트럭 오른쪽에 붙여진 합성 경계로 나타나 단일한 실제 공간이라는 조건을 크게 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 군중은 오른쪽 남자를 바라보며 손을 뻗고 접근한다. 남자의 몸은 차량 입구를 향하지만 머리는 왼쪽 아래 군중 쪽으로 돌아가 옆얼굴이 보인다. 등을 완전히 돌린 채 상체를 실내로 밀어 넣는 동작보다는 문밖 발판에서 뒤를 살피는 순간에 가깝다.",
        "built_space": "오른쪽에 열린 차량 문 하나, 출입구 하나, 거울 하나, 좌석 일부와 운전대 하나가 보인다. 출입구는 중경의 작은 요소가 아니라 전경을 크게 차지한다. 운전대가 입구 바로 안쪽에 있어 조수석보다 운전석 입구처럼 읽힌다. 왼쪽 뒤에는 별도의 군용 트럭 한 대와 지게차 한 대가 있으며, 갈라진 콘크리트 방조제와 진흙 웅덩이는 참고 장소의 재질을 따른다.",
        "entities": "주인공은 정수리가 벗겨지고 옆머리가 남은 중년 동아시아계 남성으로, 참고의 남색 체크 재킷과 어두운 바지를 유지한다. 얼굴이 일부만 보여 정확한 동일 인물 여부는 제한적으로만 확인된다. 신발과 바짓단에는 진흙이 묻어 있다. 회색·남색 작업복의 성난 군중은 샷 텍스트에 명시된 사람들에 해당한다. 군용 트럭은 두 대로 늘어났으며, 지게차는 참고와 같은 종류다. 뚜렷하게 판독되는 문구는 보이지 않는다.",
        "hard_violations": [
         "기존 장면의 군용 트럭 외에 왼쪽 배경에 두 번째 군용 트럭을 추가해 차량을 중복 생성했다."
        ],
        "physics": "남자는 오른손으로 출입구 손잡이를 잡고 왼손으로 열린 문의 윗부분을 짚는다. 오른발은 낮은 차량 발판에 걸려 있고 두 손도 몸을 지탱하므로 공중에 무지지 상태로 떠 있는 것은 아니다. 다만 두 발을 모은 자세와 거의 실외에 남은 상체는 급히 몸을 반쯤 밀어 넣는 동작을 약하게 표현한다. 군중은 진흙 바닥에 발을 딛거나 보행 중이며 지지 관계가 자연스럽다."
       },
       {
        "label": "A",
        "direction": "남자는 군중을 등지고 머리와 양팔을 차량 내부 쪽으로 향한다. 왼쪽 군중도 남자와 출입구를 바라보며 다가와 이동 목표는 맞는다. 앞줄 일부는 입을 벌리고 있으나, 전체적으로 격하게 몰려들기보다는 걸어오는 모습이다.",
        "built_space": "열린 문 하나와 거울 하나, 출입구 하나, 앞바퀴 하나가 보이며 차량은 오른쪽 경계에서 잘린다. 그러나 출입구와 남자가 전경에서 크게 보여 지정된 중경 배치와는 차이가 있다. 오른쪽에는 참고 이미지의 방조제 균열·바위·배관 구역이 차량 측면을 덮듯 이어지고, 바퀴와 차체 옆에서 서로 다른 공간의 경계가 생긴다. 이를 실제 차체 표면이나 자연스러운 반사로 설명하기 어렵다.",
        "entities": "대머리 중년 남성의 후두부, 남색 체크 재킷, 어두운 바지와 진흙 묻은 신발은 참고 인물의 보이는 특징과 부합한다. 얼굴은 뒤돌아 있어 나이와 동일 얼굴을 세밀하게 검증할 수 없다. 왼쪽에는 회색·남색 작업복과 일부 안전모를 착용한 군중이 보인다. 차량과 방조제의 개별 질감은 사실적이지만 두 대상이 하나의 물리적 공간에 놓인 방식은 성립하지 않는다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "오른쪽 방조제·바위·배관 영역을 트럭 측면과 바퀴 옆에 붙여 넣은 듯한 합성 경계가 있어, 단일 실사 프레임이 아닌 콜라주성 공간으로 나타난다."
        ],
        "physics": "남자의 오른발은 차량 발판에 걸려 있고 왼손은 문 윗부분을 짚으며 오른팔은 실내로 뻗어 있어, 왼발이 공중에 들린 승차 동작 자체에는 지지가 있다. 군중의 발도 지면과 접촉한다. 다만 오른쪽 차량과 방조제가 겹친 구역은 고체 차체, 벽, 배관이 차지하는 공간을 일관되게 설명할 수 없어 물리적 배치가 깨진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.25,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.0,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 인물",
     "[gemini-pro] 트럭과 바위벽이 융합된 물리적으로 불가능한 구조",
     "[gpt-high] 오른쪽 방조제·바위·배관 영역을 트럭 측면과 바퀴 옆에 붙여 넣은 듯한 합성 경계가 있어, 단일 실사 프레임이 아닌 콜라주성 공간으로 나타난다."
    ],
    "B": [
     "[gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 인물",
     "[gpt-high] 기존 장면의 군용 트럭 외에 왼쪽 배경에 두 번째 군용 트럭을 추가해 차량을 중복 생성했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 1000,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1000,
    "verdict_ko": "인물이 허공에 떠 있고 트럭이 바위벽과 융합되는 등 치명적인 물리적 오류로 인해 프롬프트를 실패했습니다.  ★위반: [gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 인물 / [gemini-pro] 트럭과 바위벽이 융합된 물리적으로 불가능한 구조 / [gpt-high] 오른쪽 방조제·바위·배관 영역을 트럭 측면과 바퀴 옆에 붙여 넣은 듯한 합성 경계가 있어, 단일 실사 프레임이 아닌 콜라주성 공간으로 나타난다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "인물과 군중의 묘사는 우수하나, 남성의 발이 지탱될 곳 없이 허공에 떠 있는 치명적인 물리적 오류가 있습니다.  ★위반: [gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 인물 / [gpt-high] 기존 장면의 군용 트럭 외에 왼쪽 배경에 두 번째 군용 트럭을 추가해 차량을 중복 생성했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh13_sel.png",
    "asset_id": "fbe0e5bc-3af1-4026-b52d-831b7703896e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 50대 대머리 한국인 남성: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1399111>",
    "asset_id": "41215473-9904-471e-81f7-2f6c01467eee",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-a29f-78a0-99e9-b08247d56468",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S7sh13"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S7sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:10:01.656751+00:00",
  "fingerprint": "e635eff53c23585d3323f6b74c61a2ab8ebe93569cc9939e01d4f8a31da7288a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S7sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S7sh15_sel.png",
  "source_sha256": "ceaeada66472ad8ea14460820fa6156674f82b44b062675abfd92f5cdda64bcb",
  "file": "S7sh15_cine.png",
  "staged_sha256": "924e32559bfb74325252c826f314beb758db696bba045a78115122c0d76594df",
  "latency_ms": 10273
 },
 "S8sh5::signage": {
  "fp": "fb070707f47a136b",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "붉은 도장"
   },
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "서류 종이"
   }
  ],
  "dropped": []
 },
 "era_assess::a338852cede772ba": {
  "subjects": [],
  "subject_text": "난민 재판소 내부\n낡고 허름한 실내 공간. 중앙의 초라한 탁자와 의자, 탁자 위에 높이 쌓인 서류 더미가 두드러진다.",
  "identity": "canonical",
  "scope_id": "L159",
  "scope_role": "location_interior",
  "scope_sha": "cf0a4752f24c4151"
 },
 "S8sh5::bgfirst_bg": {
  "input_fingerprint": "7a76f2066210a064",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5__bgfirst_bg.png",
  "asset_id": "70a6fcd0-a558-4ba0-8445-4a35efe45243",
  "input_asset_ids": [
   "5bd46cd4-f353-4051-a92b-c244f42f383d",
   "8c2a0cd6-3979-4553-8bb3-16258809b12f"
  ]
 },
 "S8sh5": {
  "input_fingerprint": "c40f92405782f6bc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A shabby table and chair serve as the judge's station, with a backlog of documents piled on the table.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 재판관 right now, so 재판관's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 재판관: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A shabby table and chair serve as the judge's station, with a backlog of documents piled on the table.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 재판관 right now, so 재판관's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 재판관: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A shabby table and chair serve as the judge's station, with a backlog of documents piled on the table.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 재판관 right now, so 재판관's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 재판관: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5__bgfirst_bg.png",
     "asset_id": "70a6fcd0-a558-4ba0-8445-4a35efe45243",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S8sh5.png",
     "asset_id": "5bd46cd4-f353-4051-a92b-c244f42f383d",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B01.png",
     "asset_id": "8c2a0cd6-3979-4553-8bb3-16258809b12f",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "재판관의 시선은 아래쪽의 도장을 향하고 있으며, 오른손이 도장을 쥐고 책상 위 서류를 향해 누르고 있다.",
    "built_space": "레퍼런스 사진의 법정 내부 요소(창문, 태극기, 캐비닛 등)가 보이지만, 지시된 클로즈업 구도가 아닌 미디엄 샷으로 넓게 촬영되어 카메라 앵글 지시를 위반했다.",
    "entities": "법복을 입은 재판관의 상체와 얼굴, 붉은색 사각형 도장, 서류 더미. 도장 밑으로 붉은 액체가 과도하게 번져 있다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
    ],
    "physics": "손이 도장을 잡고 서류 표면을 누르고 있으나, 도장에서 뿜어져 나온 듯한 액체의 번짐 현상은 인주의 물리적 반응으로 볼 수 없다."
   },
   {
    "label": "B",
    "direction": "손이 붉은 도장을 단단히 쥐고 서류 위를 향해 수직에 가깝게 누르고 있다.",
    "built_space": "낡은 나무 책상의 상판과 가장자리, 쌓여 있는 서류 더미가 하향 카메라 앵글을 통해 클로즈업으로 정확하게 묘사되었다.",
    "entities": "옷소매가 보이는 굵은 손, 둥근 형태의 붉은 도장, 텍스트가 읽히지 않도록 흐릿하게 처리된 서류 용지들.",
    "hard_violations": [],
    "physics": "손아귀가 도장의 윗부분을 단단하게 지지하며 쥐고 있고, 도장의 밑면이 종이 표면에 자연스럽고 안정적으로 맞닿아 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 클로즈업 샷과 하향 카메라 앵글을 정확히 구현하여, 서류에 맞닿은 붉은 도장과 굵은 손의 물리적 접촉을 완벽하게 묘사한 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레이밍 지시를 어기고 인물의 상체와 배경 전체를 넓게 보여주었으며, 도장에서 피처럼 번지는 비현실적인 붉은 액체가 생성되어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "재판관의 시선은 아래쪽의 도장을 향하고 있으며, 오른손이 도장을 쥐고 책상 위 서류를 향해 누르고 있다.",
        "built_space": "레퍼런스 사진의 법정 내부 요소(창문, 태극기, 캐비닛 등)가 보이지만, 지시된 클로즈업 구도가 아닌 미디엄 샷으로 넓게 촬영되어 카메라 앵글 지시를 위반했다.",
        "entities": "법복을 입은 재판관의 상체와 얼굴, 붉은색 사각형 도장, 서류 더미. 도장 밑으로 붉은 액체가 과도하게 번져 있다.",
        "hard_violations": [
         "물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
        ],
        "physics": "손이 도장을 잡고 서류 표면을 누르고 있으나, 도장에서 뿜어져 나온 듯한 액체의 번짐 현상은 인주의 물리적 반응으로 볼 수 없다."
       },
       {
        "label": "B",
        "direction": "손이 붉은 도장을 단단히 쥐고 서류 위를 향해 수직에 가깝게 누르고 있다.",
        "built_space": "낡은 나무 책상의 상판과 가장자리, 쌓여 있는 서류 더미가 하향 카메라 앵글을 통해 클로즈업으로 정확하게 묘사되었다.",
        "entities": "옷소매가 보이는 굵은 손, 둥근 형태의 붉은 도장, 텍스트가 읽히지 않도록 흐릿하게 처리된 서류 용지들.",
        "hard_violations": [],
        "physics": "손아귀가 도장의 윗부분을 단단하게 지지하며 쥐고 있고, 도장의 밑면이 종이 표면에 자연스럽고 안정적으로 맞닿아 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 클로즈업 샷과 하향 카메라 앵글을 정확히 구현하여, 서류에 맞닿은 붉은 도장과 굵은 손의 물리적 접촉을 완벽하게 묘사한 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레이밍 지시를 어기고 인물의 상체와 배경 전체를 넓게 보여주었으며, 도장에서 피처럼 번지는 비현실적인 붉은 액체가 생성되어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "재판관의 시선은 아래쪽의 도장을 향하고 있으며, 오른손이 도장을 쥐고 책상 위 서류를 향해 누르고 있다.",
        "built_space": "레퍼런스 사진의 법정 내부 요소(창문, 태극기, 캐비닛 등)가 보이지만, 지시된 클로즈업 구도가 아닌 미디엄 샷으로 넓게 촬영되어 카메라 앵글 지시를 위반했다.",
        "entities": "법복을 입은 재판관의 상체와 얼굴, 붉은색 사각형 도장, 서류 더미. 도장 밑으로 붉은 액체가 과도하게 번져 있다.",
        "hard_violations": [
         "물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
        ],
        "physics": "손이 도장을 잡고 서류 표면을 누르고 있으나, 도장에서 뿜어져 나온 듯한 액체의 번짐 현상은 인주의 물리적 반응으로 볼 수 없다."
       },
       {
        "label": "B",
        "direction": "손이 붉은 도장을 단단히 쥐고 서류 위를 향해 수직에 가깝게 누르고 있다.",
        "built_space": "낡은 나무 책상의 상판과 가장자리, 쌓여 있는 서류 더미가 하향 카메라 앵글을 통해 클로즈업으로 정확하게 묘사되었다.",
        "entities": "옷소매가 보이는 굵은 손, 둥근 형태의 붉은 도장, 텍스트가 읽히지 않도록 흐릿하게 처리된 서류 용지들.",
        "hard_violations": [],
        "physics": "손아귀가 도장의 윗부분을 단단하게 지지하며 쥐고 있고, 도장의 밑면이 종이 표면에 자연스럽고 안정적으로 맞닿아 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "굵은 손과 도장의 종이 접촉점을 중심으로 한 클로즈업이 정확하며, 낡은 책상 가장자리와 서류 더미도 유지한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "도장의 접촉 방향과 장소는 맞지만, 얼굴·상체·방까지 넓힌 구도가 손 중심 클로즈업에서 벗어나고 첫 접촉치고 잉크 번짐이 과하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 위에서 들어온 손이 도장 손잡이를 감싸 쥐고 아래쪽 서류를 누른다. 도장의 날인면은 종이를 향하고 실제로 맞닿아 있으며, 서류의 작업면은 비스듬히 내려다보는 카메라 쪽으로 열려 있다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "낡은 목제 책상 하나의 상판과 오른쪽 바깥 가장자리가 보인다. 서류는 전경, 왼쪽 중경, 왼쪽 후경, 중앙 후경, 오른쪽 후경에 대략 다섯 무더기로 놓여 있다. 배경에는 흐릿한 창과 수납장 일부가 보이며 참고 장소의 목재와 주간 채광에 부합한다. 의자와 재판관의 착석 위치는 이 클로즈업으로 확인할 수 없고, 중복 설비나 반사는 없다.",
        "entities": "보이는 인물 부분은 검은 소매에서 이어지는 굵고 주름진 성인 손과 손목·팔뚝 하나뿐이다. 재판관의 굵은 손이라는 지시에 부합하며, 손만으로 정확한 나이나 한국인 정체성을 확정할 수는 없다. 붉은 갈색 목제 도장 하나, 접촉 대상 서류, 쌓인 문서와 닳은 책상이 모두 있다. 도장은 선명한 빨강보다는 어두운 적갈색이고, 문서의 문자 흔적은 판독되지 않는다. 현우·앰버·미연은 등장하지 않는다.",
        "hard_violations": [],
        "physics": "손가락이 손잡이를 확실히 잡고 있으며 도장 밑면은 책상이 받치는 종이에 밀착한다. 손목에서 도장으로 내려가는 힘의 연결이 자연스럽다. 문서 더미도 상판 위에 놓여 있어 지지 없는 물체는 없다. 강한 타격의 속도감은 절제되어 있지만 접촉 순간의 자세는 가능하다."
       },
       {
        "label": "B",
        "direction": "재판관은 앞으로 숙여 도장과 서류 쪽을 내려다본다. 전경의 손이 붉은 사각 도장을 아래로 누르고, 날인면은 위로 펼쳐진 문서에 닿아 있다. 시선과 도장의 작동 방향 모두 대상 서류로 향한다.",
        "built_space": "책상 하나의 상판과 전면 가장자리, 왼쪽의 큰 서류 더미와 그 뒤의 작은 문서 묶음들, 오른쪽 필기구통 하나와 받침이 보인다. 배경에는 왼쪽 창 하나와 수납장 하나, 태극기 하나, 벽 액자 하나, 오른쪽 문 하나가 보여 참고 장소의 배치와 대체로 맞는다. 다만 손뿐 아니라 재판관의 얼굴·상체와 실내 대부분을 포함하여 지정된 클로즈업보다 넓다. 의자는 가려져 착석 상태를 확인할 수 없다.",
        "entities": "검은 법복과 흰 옷깃을 입은 중년 이상 동아시아계로 보이는 남성 재판관 한 명이 있으며, 굵은 손과 얼굴이 함께 보인다. 붉은 도장 하나, 날인 대상 문서, 낡은 목제 책상과 밀린 서류가 있다. 장면에 필요 없는 다른 인물은 없다. 문서 오른쪽에 작은 문자 형태가 남아 있으나 확실히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "손이 도장 윗부분을 잡고 밑면을 종이에 누르며, 종이와 서류 더미는 책상이 받친다. 몸은 책상 뒤에서 앞으로 기울어져 있고 팔과 손의 연결에 명백한 물리적 불가능은 없다. 다만 도장 둘레로 넓게 퍼진 붉은 잉크는 통상적인 도장의 첫 접촉보다 과장되어, 막 닿은 순간의 재료 반응으로는 설득력이 떨어진다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "굵은 손과 도장의 종이 접촉점을 중심으로 한 클로즈업이 정확하며, 낡은 책상 가장자리와 서류 더미도 유지한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "도장의 접촉 방향과 장소는 맞지만, 얼굴·상체·방까지 넓힌 구도가 손 중심 클로즈업에서 벗어나고 첫 접촉치고 잉크 번짐이 과하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 위에서 들어온 손이 도장 손잡이를 감싸 쥐고 아래쪽 서류를 누른다. 도장의 날인면은 종이를 향하고 실제로 맞닿아 있으며, 서류의 작업면은 비스듬히 내려다보는 카메라 쪽으로 열려 있다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "낡은 목제 책상 하나의 상판과 오른쪽 바깥 가장자리가 보인다. 서류는 전경, 왼쪽 중경, 왼쪽 후경, 중앙 후경, 오른쪽 후경에 대략 다섯 무더기로 놓여 있다. 배경에는 흐릿한 창과 수납장 일부가 보이며 참고 장소의 목재와 주간 채광에 부합한다. 의자와 재판관의 착석 위치는 이 클로즈업으로 확인할 수 없고, 중복 설비나 반사는 없다.",
        "entities": "보이는 인물 부분은 검은 소매에서 이어지는 굵고 주름진 성인 손과 손목·팔뚝 하나뿐이다. 재판관의 굵은 손이라는 지시에 부합하며, 손만으로 정확한 나이나 한국인 정체성을 확정할 수는 없다. 붉은 갈색 목제 도장 하나, 접촉 대상 서류, 쌓인 문서와 닳은 책상이 모두 있다. 도장은 선명한 빨강보다는 어두운 적갈색이고, 문서의 문자 흔적은 판독되지 않는다. 현우·앰버·미연은 등장하지 않는다.",
        "hard_violations": [],
        "physics": "손가락이 손잡이를 확실히 잡고 있으며 도장 밑면은 책상이 받치는 종이에 밀착한다. 손목에서 도장으로 내려가는 힘의 연결이 자연스럽다. 문서 더미도 상판 위에 놓여 있어 지지 없는 물체는 없다. 강한 타격의 속도감은 절제되어 있지만 접촉 순간의 자세는 가능하다."
       },
       {
        "label": "A",
        "direction": "재판관은 앞으로 숙여 도장과 서류 쪽을 내려다본다. 전경의 손이 붉은 사각 도장을 아래로 누르고, 날인면은 위로 펼쳐진 문서에 닿아 있다. 시선과 도장의 작동 방향 모두 대상 서류로 향한다.",
        "built_space": "책상 하나의 상판과 전면 가장자리, 왼쪽의 큰 서류 더미와 그 뒤의 작은 문서 묶음들, 오른쪽 필기구통 하나와 받침이 보인다. 배경에는 왼쪽 창 하나와 수납장 하나, 태극기 하나, 벽 액자 하나, 오른쪽 문 하나가 보여 참고 장소의 배치와 대체로 맞는다. 다만 손뿐 아니라 재판관의 얼굴·상체와 실내 대부분을 포함하여 지정된 클로즈업보다 넓다. 의자는 가려져 착석 상태를 확인할 수 없다.",
        "entities": "검은 법복과 흰 옷깃을 입은 중년 이상 동아시아계로 보이는 남성 재판관 한 명이 있으며, 굵은 손과 얼굴이 함께 보인다. 붉은 도장 하나, 날인 대상 문서, 낡은 목제 책상과 밀린 서류가 있다. 장면에 필요 없는 다른 인물은 없다. 문서 오른쪽에 작은 문자 형태가 남아 있으나 확실히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "손이 도장 윗부분을 잡고 밑면을 종이에 누르며, 종이와 서류 더미는 책상이 받친다. 몸은 책상 뒤에서 앞으로 기울어져 있고 팔과 손의 연결에 명백한 물리적 불가능은 없다. 다만 도장 둘레로 넓게 퍼진 붉은 잉크는 통상적인 도장의 첫 접촉보다 과장되어, 막 닿은 순간의 재료 반응으로는 설득력이 떨어진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 샷과 하향 카메라 앵글을 정확히 구현하여, 서류에 맞닿은 붉은 도장과 굵은 손의 물리적 접촉을 완벽하게 묘사한 결과물입니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "클로즈업 프레이밍 지시를 어기고 인물의 상체와 배경 전체를 넓게 보여주었으며, 도장에서 피처럼 번지는 비현실적인 붉은 액체가 생성되어 감점되었습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B01.png",
    "asset_id": "8c2a0cd6-3979-4553-8bb3-16258809b12f",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-a451-7b2e-8238-676064e39784",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5__bgfirst_bg.png",
   "bg_asset_id": "70a6fcd0-a558-4ba0-8445-4a35efe45243",
   "bg_record_key": "S8sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S8sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:07:47.044483+00:00",
  "fingerprint": "64f8d6a5d9b9d016ffe436413ad1cf793ae464664936b8e476f54cb3ecfefba0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S8sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S8sh5_sel.png",
  "source_sha256": "8a695eac10e1cd607fd3a0e6a8032c0dcedbb44e8a4ac45969db9565945bf7fa",
  "file": "S8sh5_cine.png",
  "staged_sha256": "d3eb702a4d68819e8fb8f9471c5f868c76710931c252e7930140c9bd40b6f05a",
  "latency_ms": 11305
 },
 "S8sh8::signage": {
  "fp": "06c16cf5d24a5c12",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S8sh8": {
  "input_fingerprint": "24cf0690fb0a88e4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러진 재판관의 멱살을 양손으로 꽉 틀어쥔 채 핏발 선 눈으로 노려보는 현우의 상체.\n\nLOCATION (lock): On the floor beside the judge's table and fallen chair inside the shabby refugee tribunal, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Courtroom floor (The judge has fallen onto it) — A narrow strip remains visible around the judge's shoulder; used as Confirms the low physical relationship between the figures; Shabby table (Remains beside the confrontation) — Its outer end is cropped at the background edge, viewed from below tabletop height; used as Connects the floor confrontation to the preceding document shot without obstructing either hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the courtroom's ambient illumination and controlled contrast, keeping this physical confrontation visually distinct from the later black-and-white nightmare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shabby courtroom table, chair, and piled case documents remain in place. 현우: Hyunwoo is bent over at floor level with both hands clenched in a collar-gripping posture; his dog-bitten leg remains injured and he is still feverish.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러진 재판관의 멱살을 양손으로 꽉 틀어쥔 채 핏발 선 눈으로 노려보는 현우의 상체.\n\nLOCATION (lock): On the floor beside the judge's table and fallen chair inside the shabby refugee tribunal, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Courtroom floor (The judge has fallen onto it) — A narrow strip remains visible around the judge's shoulder; used as Confirms the low physical relationship between the figures; Shabby table (Remains beside the confrontation) — Its outer end is cropped at the background edge, viewed from below tabletop height; used as Connects the floor confrontation to the preceding document shot without obstructing either hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the courtroom's ambient illumination and controlled contrast, keeping this physical confrontation visually distinct from the later black-and-white nightmare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shabby courtroom table, chair, and piled case documents remain in place. 현우: Hyunwoo is bent over at floor level with both hands clenched in a collar-gripping posture; his dog-bitten leg remains injured and he is still feverish.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러진 재판관의 멱살을 양손으로 꽉 틀어쥔 채 핏발 선 눈으로 노려보는 현우의 상체.\n\nLOCATION (lock): On the floor beside the judge's table and fallen chair inside the shabby refugee tribunal, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Courtroom floor (The judge has fallen onto it) — A narrow strip remains visible around the judge's shoulder; used as Confirms the low physical relationship between the figures; Shabby table (Remains beside the confrontation) — Its outer end is cropped at the background edge, viewed from below tabletop height; used as Connects the floor confrontation to the preceding document shot without obstructing either hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the courtroom's ambient illumination and controlled contrast, keeping this physical confrontation visually distinct from the later black-and-white nightmare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shabby courtroom table, chair, and piled case documents remain in place. 현우: Hyunwoo is bent over at floor level with both hands clenched in a collar-gripping posture; his dog-bitten leg remains injured and he is still feverish.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 강렬한 시선이 바닥에 쓰러진 재판관의 얼굴을 향해 똑바로 꽂혀 있습니다.",
    "built_space": "법정 바닥에서 벌어지는 상황이며, 프롬프트의 지시대로 우측 배경 가장자리에 낡은 나무 책상의 일부가 프레임에 잘린 채(cropped) 상판 아래 높이의 앵글로 배치되어 있습니다.",
    "entities": "현우는 참조 이미지의 앳된 얼굴, 헝클어진 머리, 회색 셔츠 및 카고 바지와 일치하며 핏발 선 눈을 명확히 보여줍니다. 바닥에 쓰러진 재판관 역시 확인됩니다.",
    "hard_violations": [],
    "physics": "현우는 바닥에 무릎을 대고 체중을 실은 안정적인 자세로, 두 손(양손)을 모두 사용하여 재판관의 옷깃을 강하게 움켜쥐고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 바닥에 누운 재판관을 향해 내려다보고 있습니다.",
    "built_space": "법정 바닥과 뒤편에 나무 책상, 넘어진 의자가 보입니다. 하지만 책상이 배경 가장자리에 잘리지 않고 완전히 드러나 있으며, 카메라 앵글이 상판보다 높아 구도 지시를 어겼습니다.",
    "entities": "현우의 인상착의와 눈가 표현은 참조 이미지와 일치하나, 바닥의 재판관이 법복 대신 일반 정장을 입은 것처럼 보입니다.",
    "hard_violations": [],
    "physics": "현우는 엎드린 자세를 유지하고 있으나, 프롬프트에 명시된 '양손으로 꽉 틀어쥔' 상태가 아니라 오른손 하나로만 옷깃을 쥐고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 낮은 카메라 앵글, 우측 배경 가장자리에 잘린 책상의 구도, 그리고 양손으로 멱살을 쥐는 동작과 핏발 선 눈을 모두 완벽하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "카메라 앵글이 높아 책상 윗면이 보이고 책상이 프레임에 걸치지 않았으며, 양손이 아닌 한 손으로만 옷깃을 쥐고 있어 핵심 지시사항들을 다수 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 강렬한 시선이 바닥에 쓰러진 재판관의 얼굴을 향해 똑바로 꽂혀 있습니다.",
        "built_space": "법정 바닥에서 벌어지는 상황이며, 프롬프트의 지시대로 우측 배경 가장자리에 낡은 나무 책상의 일부가 프레임에 잘린 채(cropped) 상판 아래 높이의 앵글로 배치되어 있습니다.",
        "entities": "현우는 참조 이미지의 앳된 얼굴, 헝클어진 머리, 회색 셔츠 및 카고 바지와 일치하며 핏발 선 눈을 명확히 보여줍니다. 바닥에 쓰러진 재판관 역시 확인됩니다.",
        "hard_violations": [],
        "physics": "현우는 바닥에 무릎을 대고 체중을 실은 안정적인 자세로, 두 손(양손)을 모두 사용하여 재판관의 옷깃을 강하게 움켜쥐고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 바닥에 누운 재판관을 향해 내려다보고 있습니다.",
        "built_space": "법정 바닥과 뒤편에 나무 책상, 넘어진 의자가 보입니다. 하지만 책상이 배경 가장자리에 잘리지 않고 완전히 드러나 있으며, 카메라 앵글이 상판보다 높아 구도 지시를 어겼습니다.",
        "entities": "현우의 인상착의와 눈가 표현은 참조 이미지와 일치하나, 바닥의 재판관이 법복 대신 일반 정장을 입은 것처럼 보입니다.",
        "hard_violations": [],
        "physics": "현우는 엎드린 자세를 유지하고 있으나, 프롬프트에 명시된 '양손으로 꽉 틀어쥔' 상태가 아니라 오른손 하나로만 옷깃을 쥐고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 낮은 카메라 앵글, 우측 배경 가장자리에 잘린 책상의 구도, 그리고 양손으로 멱살을 쥐는 동작과 핏발 선 눈을 모두 완벽하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "카메라 앵글이 높아 책상 윗면이 보이고 책상이 프레임에 걸치지 않았으며, 양손이 아닌 한 손으로만 옷깃을 쥐고 있어 핵심 지시사항들을 다수 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 강렬한 시선이 바닥에 쓰러진 재판관의 얼굴을 향해 똑바로 꽂혀 있습니다.",
        "built_space": "법정 바닥에서 벌어지는 상황이며, 프롬프트의 지시대로 우측 배경 가장자리에 낡은 나무 책상의 일부가 프레임에 잘린 채(cropped) 상판 아래 높이의 앵글로 배치되어 있습니다.",
        "entities": "현우는 참조 이미지의 앳된 얼굴, 헝클어진 머리, 회색 셔츠 및 카고 바지와 일치하며 핏발 선 눈을 명확히 보여줍니다. 바닥에 쓰러진 재판관 역시 확인됩니다.",
        "hard_violations": [],
        "physics": "현우는 바닥에 무릎을 대고 체중을 실은 안정적인 자세로, 두 손(양손)을 모두 사용하여 재판관의 옷깃을 강하게 움켜쥐고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 바닥에 누운 재판관을 향해 내려다보고 있습니다.",
        "built_space": "법정 바닥과 뒤편에 나무 책상, 넘어진 의자가 보입니다. 하지만 책상이 배경 가장자리에 잘리지 않고 완전히 드러나 있으며, 카메라 앵글이 상판보다 높아 구도 지시를 어겼습니다.",
        "entities": "현우의 인상착의와 눈가 표현은 참조 이미지와 일치하나, 바닥의 재판관이 법복 대신 일반 정장을 입은 것처럼 보입니다.",
        "hard_violations": [],
        "physics": "현우는 엎드린 자세를 유지하고 있으나, 프롬프트에 명시된 '양손으로 꽉 틀어쥔' 상태가 아니라 오른손 하나로만 옷깃을 쥐고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "바닥의 대치와 현우의 외형은 맞지만, 양손 멱살잡기가 명확하지 않고 책상과 바닥을 지나치게 넓게 보여 지정된 상체 중심 구도에서 벗어난다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "재판관을 향한 아래쪽 시선, 양손으로 움켜쥔 옷깃, 바닥에 지지된 자세와 가장자리에서 잘린 낡은 책상이 핵심 연출에 가장 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개를 숙이고 화면 오른쪽 아래를 노려보지만, 눈동자의 방향은 재판관의 얼굴보다 조금 위를 향하는 듯해 정확한 눈맞춤은 약하다. 재판관의 얼굴은 현우 쪽으로 돌아가 있다. 앞쪽 주먹은 재판관의 목깃에 닿지만 반대쪽 손의 멱살잡기는 확인되지 않는다.",
        "built_space": "뒤쪽에 목제 책상 한 개, 왼쪽 바닥에 넘어진 금속 프레임 의자 한 개가 보인다. 책상 위에는 여러 문서 더미가 있고 오른쪽 바닥에도 서류가 놓여 있다. 현우는 누운 재판관 위로 몸을 굽히고 있다. 책상 전면과 상판 양 끝이 거의 모두 보여, 배경 가장자리에서 끝부분만 잘려 보이라는 요구보다 책상의 비중이 크다. 어깨 주변 바닥도 좁은 띠 이상으로 드러난다. 반사는 없다.",
        "entities": "두 사람만 보인다. 현우는 앳된 동아시아계 남성으로 검은 헝클어진 머리, 회색 셔츠와 올리브색 바지가 참조와 대체로 맞는다. 눈 주변이 붉고 얼굴에 피로감이 있으나 안구는 정상적인 형태다. 재판관은 중년 이상으로 보이는 남성이며 검은 겉옷과 흰 셔츠를 입었다. 낡은 목재 책상과 종이 문서의 재질은 이전 장면과 통한다. 다리의 부상은 확인할 수 없으며 이를 감점할 이유는 없다. 명확하게 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "재판관의 등과 몸통은 바닥에 놓여 있고 머리도 바닥 가까이에 있어 공중에 뜬 상태가 아니다. 현우의 하체 접점은 상당 부분 가려졌지만 바닥 가까이 몸을 낮춘 자세는 가능하다. 보이는 주먹은 옷깃을 실제로 움켜쥐고 있다. 반대쪽 팔과 손은 재판관 머리 부근에 가려져 양손의 당기는 힘은 확인하기 어렵다. 책상은 다리로, 넘어진 의자는 옆면 프레임으로 바닥에 지지된다."
       },
       {
        "label": "B",
        "direction": "현우의 고개와 눈동자가 모두 아래쪽 재판관의 얼굴을 향한다. 양손은 재판관 목 양옆의 검은 옷깃을 각각 움켜쥐어 위로 당긴다. 재판관은 얼굴을 위로 둔 채 누워 있으며 눈의 초점은 프레임 아래쪽 잘림 때문에 판별하기 어렵다.",
        "built_space": "오른쪽에 낡은 목제 책상 한 개가 있고 상판 위로 문서 더미들이 보인다. 책상의 바깥쪽 끝은 오른쪽 화면 경계에서 잘리며 카메라는 상판 아래 높이에 있다. 뒤에는 낮빛이 들어오는 창과 낡은 벽이 보인다. 재판관은 책상 옆 바닥에 누워 있고 현우는 그 위로 몸을 굽힌다. 어깨 옆 바닥이 드러나 낮은 위치 관계를 설명한다. 넘어진 의자는 이 구도 안에서는 확인되지 않는다. 중복 가구나 반사는 없다.",
        "entities": "현우와 재판관 두 사람만 등장한다. 현우의 앳된 동아시아계 얼굴, 헝클어진 검은 머리, 마른 체격, 낡은 회색 셔츠와 올리브색 카고 바지는 참조에 가깝다. 붉은 눈 주변과 땀은 열이 오른 상태를 자연스럽게 표현하며 눈 자체의 비정상적인 변형은 없다. 재판관은 성인 남성으로 검은 겉옷과 흰 셔츠를 입고 있다. 낡은 목재 책상과 누렇게 바랜 문서 더미는 장소 참조의 재질과 분위기를 유지한다. 다리의 물린 상처는 노출되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 화면 왼쪽 무릎과 정강이가 바닥에 닿고 반대쪽 다리는 굽혀져 있어 앞으로 기울인 상체를 지지한다. 두 손가락이 각각 옷깃을 감싸며 천이 주먹 쪽으로 당겨져 실제 접촉과 장력이 보인다. 재판관의 등과 하체는 바닥에 지지되고 목 부근만 옷깃에 의해 당겨지는 자세라 실행 가능한 동작이다. 책상 다리는 바닥에 닿고 문서들은 상판에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "바닥의 대치와 현우의 외형은 맞지만, 양손 멱살잡기가 명확하지 않고 책상과 바닥을 지나치게 넓게 보여 지정된 상체 중심 구도에서 벗어난다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "재판관을 향한 아래쪽 시선, 양손으로 움켜쥔 옷깃, 바닥에 지지된 자세와 가장자리에서 잘린 낡은 책상이 핵심 연출에 가장 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개를 숙이고 화면 오른쪽 아래를 노려보지만, 눈동자의 방향은 재판관의 얼굴보다 조금 위를 향하는 듯해 정확한 눈맞춤은 약하다. 재판관의 얼굴은 현우 쪽으로 돌아가 있다. 앞쪽 주먹은 재판관의 목깃에 닿지만 반대쪽 손의 멱살잡기는 확인되지 않는다.",
        "built_space": "뒤쪽에 목제 책상 한 개, 왼쪽 바닥에 넘어진 금속 프레임 의자 한 개가 보인다. 책상 위에는 여러 문서 더미가 있고 오른쪽 바닥에도 서류가 놓여 있다. 현우는 누운 재판관 위로 몸을 굽히고 있다. 책상 전면과 상판 양 끝이 거의 모두 보여, 배경 가장자리에서 끝부분만 잘려 보이라는 요구보다 책상의 비중이 크다. 어깨 주변 바닥도 좁은 띠 이상으로 드러난다. 반사는 없다.",
        "entities": "두 사람만 보인다. 현우는 앳된 동아시아계 남성으로 검은 헝클어진 머리, 회색 셔츠와 올리브색 바지가 참조와 대체로 맞는다. 눈 주변이 붉고 얼굴에 피로감이 있으나 안구는 정상적인 형태다. 재판관은 중년 이상으로 보이는 남성이며 검은 겉옷과 흰 셔츠를 입었다. 낡은 목재 책상과 종이 문서의 재질은 이전 장면과 통한다. 다리의 부상은 확인할 수 없으며 이를 감점할 이유는 없다. 명확하게 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "재판관의 등과 몸통은 바닥에 놓여 있고 머리도 바닥 가까이에 있어 공중에 뜬 상태가 아니다. 현우의 하체 접점은 상당 부분 가려졌지만 바닥 가까이 몸을 낮춘 자세는 가능하다. 보이는 주먹은 옷깃을 실제로 움켜쥐고 있다. 반대쪽 팔과 손은 재판관 머리 부근에 가려져 양손의 당기는 힘은 확인하기 어렵다. 책상은 다리로, 넘어진 의자는 옆면 프레임으로 바닥에 지지된다."
       },
       {
        "label": "A",
        "direction": "현우의 고개와 눈동자가 모두 아래쪽 재판관의 얼굴을 향한다. 양손은 재판관 목 양옆의 검은 옷깃을 각각 움켜쥐어 위로 당긴다. 재판관은 얼굴을 위로 둔 채 누워 있으며 눈의 초점은 프레임 아래쪽 잘림 때문에 판별하기 어렵다.",
        "built_space": "오른쪽에 낡은 목제 책상 한 개가 있고 상판 위로 문서 더미들이 보인다. 책상의 바깥쪽 끝은 오른쪽 화면 경계에서 잘리며 카메라는 상판 아래 높이에 있다. 뒤에는 낮빛이 들어오는 창과 낡은 벽이 보인다. 재판관은 책상 옆 바닥에 누워 있고 현우는 그 위로 몸을 굽힌다. 어깨 옆 바닥이 드러나 낮은 위치 관계를 설명한다. 넘어진 의자는 이 구도 안에서는 확인되지 않는다. 중복 가구나 반사는 없다.",
        "entities": "현우와 재판관 두 사람만 등장한다. 현우의 앳된 동아시아계 얼굴, 헝클어진 검은 머리, 마른 체격, 낡은 회색 셔츠와 올리브색 카고 바지는 참조에 가깝다. 붉은 눈 주변과 땀은 열이 오른 상태를 자연스럽게 표현하며 눈 자체의 비정상적인 변형은 없다. 재판관은 성인 남성으로 검은 겉옷과 흰 셔츠를 입고 있다. 낡은 목재 책상과 누렇게 바랜 문서 더미는 장소 참조의 재질과 분위기를 유지한다. 다리의 물린 상처는 노출되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 화면 왼쪽 무릎과 정강이가 바닥에 닿고 반대쪽 다리는 굽혀져 있어 앞으로 기울인 상체를 지지한다. 두 손가락이 각각 옷깃을 감싸며 천이 주먹 쪽으로 당겨져 실제 접촉과 장력이 보인다. 재판관의 등과 하체는 바닥에 지지되고 목 부근만 옷깃에 의해 당겨지는 자세라 실행 가능한 동작이다. 책상 다리는 바닥에 닿고 문서들은 상판에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.111
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.111
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1111
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 낮은 카메라 앵글, 우측 배경 가장자리에 잘린 책상의 구도, 그리고 양손으로 멱살을 쥐는 동작과 핏발 선 눈을 모두 완벽하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1111,
    "verdict_ko": "카메라 앵글이 높아 책상 윗면이 보이고 책상이 프레임에 걸치지 않았으며, 양손이 아닌 한 손으로만 옷깃을 쥐고 있어 핵심 지시사항들을 다수 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5_sel.png",
    "asset_id": "f229b1cd-f3f6-4c3b-b29c-a990d8901eb1",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-a7b3-71f4-8b93-cd23ad7cfc40",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S8sh5"
  }
 },
 "S8sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:11:09.912486+00:00",
  "fingerprint": "db40af74d7a8f7d46721dd99f035ed207eb01a0cd30458c195494b0d7add96ca",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S8sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S8sh8_sel.png",
  "source_sha256": "0f7f279146bbd0d99a50b30a490cbb8bb5e057d55d269ef22b8cb434d9cafa19",
  "file": "S8sh8_cine.png",
  "staged_sha256": "30b97887c2f586b208ce976f28fa26b2e03e6b3c302c18c425e77d94721469af",
  "latency_ms": 10888
 },
 "S8sh13::signage": {
  "fp": "b17dd516f0137658",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S8sh13::bgfirst_bg": {
  "input_fingerprint": "3b44ab324d5bcb68",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh13__bgfirst_bg.png",
  "asset_id": "8bcd8334-9652-44d4-9e5a-7ec29199889a",
  "input_asset_ids": [
   "863e681c-c732-49e9-9c06-7c759918fd73",
   "dcf94519-9d42-4399-a0db-17ea44888173"
  ]
 },
 "S8sh13": {
  "input_fingerprint": "5a7c9dbb5c0df435",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): This is a black-and-white nighttime memory of a U.S. street, with intermittent gunfire flashes and a large American flag. 미연: Miyeon is kneeling in terror, hunched with both arms wrapped protectively inward. 현우: Hyunwoo appears as a younger child, huddled low within a protective embrace. 앰버: Amber appears as a younger child, huddled low within a protective embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): This is a black-and-white nighttime memory of a U.S. street, with intermittent gunfire flashes and a large American flag. 미연: Miyeon is kneeling in terror, hunched with both arms wrapped protectively inward. 현우: Hyunwoo appears as a younger child, huddled low within a protective embrace. 앰버: Amber appears as a younger child, huddled low within a protective embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): This is a black-and-white nighttime memory of a U.S. street, with intermittent gunfire flashes and a large American flag. 미연: Miyeon is kneeling in terror, hunched with both arms wrapped protectively inward. 현우: Hyunwoo appears as a younger child, huddled low within a protective embrace. 앰버: Amber appears as a younger child, huddled low within a protective embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh13__bgfirst_bg.png",
     "asset_id": "8bcd8334-9652-44d4-9e5a-7ec29199889a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S8sh13.png",
     "asset_id": "863e681c-c732-49e9-9c06-7c759918fd73",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B02.png",
     "asset_id": "dcf94519-9d42-4399-a0db-17ea44888173",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "세 인물 모두 화면 우측 밖을 불안하게 응시하고 있음.",
    "built_space": "참조 이미지의 야간 거리. 좌측 배경에 지시된 미국 국기가 배치됨.",
    "entities": "미연과 어린 현우, 앰버만 정확히 등장함. 샷 텍스트에 없는 인물은 없음.",
    "hard_violations": [
     "[gemini-pro] 콜라주/스티커 아티팩트 (인물 그룹 좌측의 부자연스러운 흰색 컷아웃 윤곽선)"
    ],
    "physics": "바닥에 웅크린 자세는 지탱되나, 윤곽선 결함으로 인해 배경에서 물리적으로 붕 떠 보임."
   },
   {
    "label": "B",
    "direction": "미연은 정면을 바라보며, 배경의 추가 인물들이 양옆으로 총구를 겨누고 있음.",
    "built_space": "참조 이미지의 상점과 야간 거리 구조를 잘 반영함.",
    "entities": "미연, 어린 현우, 앰버 외에 텍스트에 없는 무장 인물들이 다수 등장함. 미국 국기 누락.",
    "hard_violations": [
     "[gemini-pro] 발명된 인물 (샷 텍스트에 명시되지 않은 배경의 무장 인물 4명 이상)",
     "[gpt-high] 숏 텍스트가 허용한 가족 세 명 외에 추가 인물 네 명과 이들이 사용하는 총기를 배치했다."
    ],
    "physics": "인물들 모두 지면 위에 안정적으로 무릎을 꿇고 웅크린 상태로 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 인물만 배치하고 미국 국기를 포함했으나, 인물 좌측에 스티커를 오려 붙인 듯한 뚜렷한 흰색 윤곽선이 발생해 하드 위반입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지정되지 않은 다수의 무장 인물을 배경에 임의로 추가하여 엄격한 인물 제한 규칙을 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "세 인물 모두 화면 우측 밖을 불안하게 응시하고 있음.",
        "built_space": "참조 이미지의 야간 거리. 좌측 배경에 지시된 미국 국기가 배치됨.",
        "entities": "미연과 어린 현우, 앰버만 정확히 등장함. 샷 텍스트에 없는 인물은 없음.",
        "hard_violations": [
         "콜라주/스티커 아티팩트 (인물 그룹 좌측의 부자연스러운 흰색 컷아웃 윤곽선)"
        ],
        "physics": "바닥에 웅크린 자세는 지탱되나, 윤곽선 결함으로 인해 배경에서 물리적으로 붕 떠 보임."
       },
       {
        "label": "B",
        "direction": "미연은 정면을 바라보며, 배경의 추가 인물들이 양옆으로 총구를 겨누고 있음.",
        "built_space": "참조 이미지의 상점과 야간 거리 구조를 잘 반영함.",
        "entities": "미연, 어린 현우, 앰버 외에 텍스트에 없는 무장 인물들이 다수 등장함. 미국 국기 누락.",
        "hard_violations": [
         "발명된 인물 (샷 텍스트에 명시되지 않은 배경의 무장 인물 4명 이상)"
        ],
        "physics": "인물들 모두 지면 위에 안정적으로 무릎을 꿇고 웅크린 상태로 지탱됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 인물만 배치하고 미국 국기를 포함했으나, 인물 좌측에 스티커를 오려 붙인 듯한 뚜렷한 흰색 윤곽선이 발생해 하드 위반입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지정되지 않은 다수의 무장 인물을 배경에 임의로 추가하여 엄격한 인물 제한 규칙을 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "세 인물 모두 화면 우측 밖을 불안하게 응시하고 있음.",
        "built_space": "참조 이미지의 야간 거리. 좌측 배경에 지시된 미국 국기가 배치됨.",
        "entities": "미연과 어린 현우, 앰버만 정확히 등장함. 샷 텍스트에 없는 인물은 없음.",
        "hard_violations": [
         "콜라주/스티커 아티팩트 (인물 그룹 좌측의 부자연스러운 흰색 컷아웃 윤곽선)"
        ],
        "physics": "바닥에 웅크린 자세는 지탱되나, 윤곽선 결함으로 인해 배경에서 물리적으로 붕 떠 보임."
       },
       {
        "label": "B",
        "direction": "미연은 정면을 바라보며, 배경의 추가 인물들이 양옆으로 총구를 겨누고 있음.",
        "built_space": "참조 이미지의 상점과 야간 거리 구조를 잘 반영함.",
        "entities": "미연, 어린 현우, 앰버 외에 텍스트에 없는 무장 인물들이 다수 등장함. 미국 국기 누락.",
        "hard_violations": [
         "발명된 인물 (샷 텍스트에 명시되지 않은 배경의 무장 인물 4명 이상)"
        ],
        "physics": "인물들 모두 지면 위에 안정적으로 무릎을 꿇고 웅크린 상태로 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 추가 인물과 총기를 배치했으며, 가족의 다리와 신발까지 보여주는 넓은 구도로 미연의 상체 중심 미디엄 숏 지시를 벗어났다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "세 사람만을 밀착된 상체 중심으로 담아 공포와 보호 포옹을 구현하고 장소도 더 충실하지만, 의상이 참조와 다르고 얼굴 질감이 다소 회화적이다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 정면보다 조금 위쪽의 화면 밖을 두려운 표정으로 바라본다. 앰버는 아래로 얼굴을 묻고, 현우는 미연의 가슴 쪽으로 고개를 돌려 포옹 안으로 파고든다. 왼쪽 총잡이들의 총구는 화면 오른쪽을, 오른쪽 가장자리 총구는 왼쪽을 향한다. 총격의 정확한 표적은 확인되지 않으며, 이 총잡이들은 애초에 숏에 허용된 인물이 아니다.",
        "built_space": "가족은 차도 위에 무릎을 꿇고 있고, 뒤에는 연석과 보도, 벽돌 기둥 사이의 판자로 막힌 큰 창들이 이어진다. 왼쪽에는 가로등들이 보인다. 참조의 벽돌 상가 재료는 따르지만, 가족 전신과 주변 총잡이들까지 포함해 제한된 배경의 상체 숏보다 훨씬 넓다. 참조 오른쪽의 차양 달린 문은 이 구도에서 확인되지 않으며, 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "중앙에는 검은 머리의 중년 동아시아계 여성, 어린 검은 머리 남자아이, 밝은 머리 여자아이가 있어 세 주인공의 기본 외형과 어린 시절 설정은 대체로 맞는다. 미연과 앰버의 어두운 상의는 참조와 유사하지만, 현우는 참조의 단추 셔츠 대신 둥근 목 상의를 입었다. 가족 외에 왼쪽 세 명과 오른쪽 가장자리 한 명이 추가되어 있고 여러 총기가 등장한다. 요구된 성조기는 보이지 않으며 읽을 수 있는 글자는 없다. 흑백 야간 표현으로 기억 장면 지시는 따르지만, 별도로 적힌 낮 시간 잠금과는 충돌한다.",
        "hard_violations": [
         "숏 텍스트가 허용한 가족 세 명 외에 추가 인물 네 명과 이들이 사용하는 총기를 배치했다."
        ],
        "physics": "아이들의 무릎과 접힌 정강이, 신발이 노면에 닿아 체중을 지지한다. 미연도 아이들 사이에서 하체를 접고 앉아 있으며 양손이 각각 아이의 어깨와 등을 감싼다. 추가 총잡이들은 무릎이나 발로 지지되고 총기를 손으로 잡고 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "미연은 화면 오른쪽 위의 화면 밖을 눈을 크게 뜨고 바라보며, 두 아이도 같은 오른쪽 방향을 경계한다. 보이는 오른쪽 섬광 쪽으로 주의가 모이지만 실제 위협 인물은 화면에 없다. 미연의 양팔은 바깥으로 뻗지 않고 두 아이를 몸 안쪽으로 끌어안는다. 조준 방향을 검증할 총기나 총구는 보이지 않는다.",
        "built_space": "가족은 왼쪽 연석 바로 옆 차도에 낮게 웅크려 있다. 오른쪽에는 참조와 대응하는 낡은 벽면, 문 한 개와 그 위 차양 한 개, 뒤쪽 벽돌 상가가 보인다. 왼쪽 보도에는 가로등과 나무가 원근에 맞게 줄지어 있고, 작은 차량들이 멀리 있다. 가족의 상체와 포옹이 화면 대부분을 차지하며 하체는 아래에서 잘려 A보다 미디엄 숏에 가깝다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "추가 인물 없이 미연과 어린 현우, 앰버 세 명만 보인다. 미연의 중년 얼굴과 검은 머리, 현우의 어린 동아시아계 남아 외형과 헝클어진 검은 머리, 앰버의 밝은 머리와 둥근 아동 얼굴은 설정에 대체로 부합한다. 다만 세 사람의 겉옷과 긴소매 차림은 제공된 의상 참조와 다르다. 왼쪽 위에는 성조기 한 개가 걸려 있고 오른쪽에는 강한 섬광이 보이지만 발사원은 확인되지 않는다. 읽을 수 있는 글자는 없다. 흑백 야간 기억 장면을 따르므로 별도의 낮 시간 잠금과는 충돌한다. 피부와 옷의 명암은 실사보다 다소 매끈하게 가공된 인상이다.",
        "hard_violations": [],
        "physics": "미연의 팔은 두 아이의 바깥쪽 어깨를 돌아 앞가슴 쪽을 단단히 감싸며 손과 옷의 접촉이 보인다. 아이들은 무릎을 몸 가까이 접고 미연에게 기대어 있다. 지면 접촉부는 하단 크롭에 가려져 무릎을 꿇었는지 완전히 확인할 수 없지만, 노면에 낮게 앉거나 쪼그린 자세로 성립하며 공중에 떠 있는 징후는 없다. 성조기는 건물 쪽 지지대에 걸려 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 추가 인물과 총기를 배치했으며, 가족의 다리와 신발까지 보여주는 넓은 구도로 미연의 상체 중심 미디엄 숏 지시를 벗어났다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "세 사람만을 밀착된 상체 중심으로 담아 공포와 보호 포옹을 구현하고 장소도 더 충실하지만, 의상이 참조와 다르고 얼굴 질감이 다소 회화적이다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 정면보다 조금 위쪽의 화면 밖을 두려운 표정으로 바라본다. 앰버는 아래로 얼굴을 묻고, 현우는 미연의 가슴 쪽으로 고개를 돌려 포옹 안으로 파고든다. 왼쪽 총잡이들의 총구는 화면 오른쪽을, 오른쪽 가장자리 총구는 왼쪽을 향한다. 총격의 정확한 표적은 확인되지 않으며, 이 총잡이들은 애초에 숏에 허용된 인물이 아니다.",
        "built_space": "가족은 차도 위에 무릎을 꿇고 있고, 뒤에는 연석과 보도, 벽돌 기둥 사이의 판자로 막힌 큰 창들이 이어진다. 왼쪽에는 가로등들이 보인다. 참조의 벽돌 상가 재료는 따르지만, 가족 전신과 주변 총잡이들까지 포함해 제한된 배경의 상체 숏보다 훨씬 넓다. 참조 오른쪽의 차양 달린 문은 이 구도에서 확인되지 않으며, 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "중앙에는 검은 머리의 중년 동아시아계 여성, 어린 검은 머리 남자아이, 밝은 머리 여자아이가 있어 세 주인공의 기본 외형과 어린 시절 설정은 대체로 맞는다. 미연과 앰버의 어두운 상의는 참조와 유사하지만, 현우는 참조의 단추 셔츠 대신 둥근 목 상의를 입었다. 가족 외에 왼쪽 세 명과 오른쪽 가장자리 한 명이 추가되어 있고 여러 총기가 등장한다. 요구된 성조기는 보이지 않으며 읽을 수 있는 글자는 없다. 흑백 야간 표현으로 기억 장면 지시는 따르지만, 별도로 적힌 낮 시간 잠금과는 충돌한다.",
        "hard_violations": [
         "숏 텍스트가 허용한 가족 세 명 외에 추가 인물 네 명과 이들이 사용하는 총기를 배치했다."
        ],
        "physics": "아이들의 무릎과 접힌 정강이, 신발이 노면에 닿아 체중을 지지한다. 미연도 아이들 사이에서 하체를 접고 앉아 있으며 양손이 각각 아이의 어깨와 등을 감싼다. 추가 총잡이들은 무릎이나 발로 지지되고 총기를 손으로 잡고 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "미연은 화면 오른쪽 위의 화면 밖을 눈을 크게 뜨고 바라보며, 두 아이도 같은 오른쪽 방향을 경계한다. 보이는 오른쪽 섬광 쪽으로 주의가 모이지만 실제 위협 인물은 화면에 없다. 미연의 양팔은 바깥으로 뻗지 않고 두 아이를 몸 안쪽으로 끌어안는다. 조준 방향을 검증할 총기나 총구는 보이지 않는다.",
        "built_space": "가족은 왼쪽 연석 바로 옆 차도에 낮게 웅크려 있다. 오른쪽에는 참조와 대응하는 낡은 벽면, 문 한 개와 그 위 차양 한 개, 뒤쪽 벽돌 상가가 보인다. 왼쪽 보도에는 가로등과 나무가 원근에 맞게 줄지어 있고, 작은 차량들이 멀리 있다. 가족의 상체와 포옹이 화면 대부분을 차지하며 하체는 아래에서 잘려 A보다 미디엄 숏에 가깝다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "추가 인물 없이 미연과 어린 현우, 앰버 세 명만 보인다. 미연의 중년 얼굴과 검은 머리, 현우의 어린 동아시아계 남아 외형과 헝클어진 검은 머리, 앰버의 밝은 머리와 둥근 아동 얼굴은 설정에 대체로 부합한다. 다만 세 사람의 겉옷과 긴소매 차림은 제공된 의상 참조와 다르다. 왼쪽 위에는 성조기 한 개가 걸려 있고 오른쪽에는 강한 섬광이 보이지만 발사원은 확인되지 않는다. 읽을 수 있는 글자는 없다. 흑백 야간 기억 장면을 따르므로 별도의 낮 시간 잠금과는 충돌한다. 피부와 옷의 명암은 실사보다 다소 매끈하게 가공된 인상이다.",
        "hard_violations": [],
        "physics": "미연의 팔은 두 아이의 바깥쪽 어깨를 돌아 앞가슴 쪽을 단단히 감싸며 손과 옷의 접촉이 보인다. 아이들은 무릎을 몸 가까이 접고 미연에게 기대어 있다. 지면 접촉부는 하단 크롭에 가려져 무릎을 꿇었는지 완전히 확인할 수 없지만, 노면에 낮게 앉거나 쪼그린 자세로 성립하며 공중에 떠 있는 징후는 없다. 성조기는 건물 쪽 지지대에 걸려 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.917
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.667
   },
   "violations": {
    "A": [
     "[gemini-pro] 콜라주/스티커 아티팩트 (인물 그룹 좌측의 부자연스러운 흰색 컷아웃 윤곽선)"
    ],
    "B": [
     "[gemini-pro] 발명된 인물 (샷 텍스트에 명시되지 않은 배경의 무장 인물 4명 이상)",
     "[gpt-high] 숏 텍스트가 허용한 가족 세 명 외에 추가 인물 네 명과 이들이 사용하는 총기를 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 인물만 배치하고 미국 국기를 포함했으나, 인물 좌측에 스티커를 오려 붙인 듯한 뚜렷한 흰색 윤곽선이 발생해 하드 위반입니다.  ★위반: [gemini-pro] 콜라주/스티커 아티팩트 (인물 그룹 좌측의 부자연스러운 흰색 컷아웃 윤곽선)"
   },
   {
    "label": "B",
    "score": 667,
    "verdict_ko": "지정되지 않은 다수의 무장 인물을 배경에 임의로 추가하여 엄격한 인물 제한 규칙을 심각하게 위반했습니다.  ★위반: [gemini-pro] 발명된 인물 (샷 텍스트에 명시되지 않은 배경의 무장 인물 4명 이상) / [gpt-high] 숏 텍스트가 허용한 가족 세 명 외에 추가 인물 네 명과 이들이 사용하는 총기를 배치했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B02.png",
    "asset_id": "dcf94519-9d42-4399-a0db-17ea44888173",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-a959-7dba-a174-5a7ac5f299bf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh13__bgfirst_bg.png",
   "bg_asset_id": "8bcd8334-9652-44d4-9e5a-7ec29199889a",
   "bg_record_key": "S8sh13::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S8sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:12:26.037971+00:00",
  "fingerprint": "c015d164c8387a54d2842885035e722e9bd896cc75e838bc09f0e99810a8adf2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S8sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S8sh13_sel.png",
  "source_sha256": "683e14c1a28952267ecab98ab73c9fb8ccec255842c2efd870f897b301886290",
  "file": "S8sh13_cine.png",
  "staged_sha256": "6ab891f40732b447b71ca11586065b74cfedf876fe52f14752807d0c3426249a",
  "latency_ms": 11351
 },
 "S9sh2::signage": {
  "fp": "315f6cf4fa6208e8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::6fab8fc81ab4f62a": {
  "subjects": [],
  "subject_text": "난민 재판소 밖 쓰레기 더미\n허름한 건물 밖으로 각종 폐기물이 불규칙하게 쌓인 황폐한 공간. 쓰레기 더미가 주변 바닥을 뒤덮고 있다.",
  "identity": "canonical",
  "scope_id": "L161",
  "scope_role": "location_exterior",
  "scope_sha": "d71d8f809e2bbef2"
 },
 "S9sh2::bgfirst_bg": {
  "input_fingerprint": "941c58ae665c2cb8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2__bgfirst_bg.png",
  "asset_id": "74193d07-10b1-46b2-a9cb-c0e151291eb3",
  "input_asset_ids": [
   "e2420bad-c4d6-4722-a5e2-886e38397c1a",
   "595d4bdd-0afb-4da7-95e6-a460d0259575"
  ]
 },
 "S9sh2": {
  "input_fingerprint": "c1fe56d558ed994f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night outside the refugee court, surrounded by desolate rubbish heaps. 현우: Hyunwoo is waking outside with his dog-bitten leg and injuries from the baton blow still untreated, though his awareness is returning. The stolen wallet containing the judge's identification and cash is already in his pocket.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night outside the refugee court, surrounded by desolate rubbish heaps. 현우: Hyunwoo is waking outside with his dog-bitten leg and injuries from the baton blow still untreated, though his awareness is returning. The stolen wallet containing the judge's identification and cash is already in his pocket.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night outside the refugee court, surrounded by desolate rubbish heaps. 현우: Hyunwoo is waking outside with his dog-bitten leg and injuries from the baton blow still untreated, though his awareness is returning. The stolen wallet containing the judge's identification and cash is already in his pocket.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2__bgfirst_bg.png",
     "asset_id": "74193d07-10b1-46b2-a9cb-c0e151291eb3",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S9sh2.png",
     "asset_id": "e2420bad-c4d6-4722-a5e2-886e38397c1a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L161B01.png",
     "asset_id": "595d4bdd-0afb-4da7-95e6-a460d0259575",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 아래쪽을 향하고 있으며, 입을 크게 벌리고 가쁜 숨을 쉬는 표정임.",
    "built_space": "레퍼런스 이미지에 있는 우측의 콘크리트 건물 외벽과 바닥의 통로, 좌측의 쓰레기 더미가 정확한 위치에 존재함.",
    "entities": "현우의 얼굴, 헤어스타일, 회색 셔츠 및 바지가 캐릭터 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "바닥에 주저앉아 한 손은 가슴의 옷깃을 강하게 쥐고 있으며, 다른 한 손은 바닥을 짚어 체중을 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 정면 약간 좌측을 향하고, 입을 벌려 숨을 쉬고 있음.",
    "built_space": "쓰레기 더미는 보이나, 위치 레퍼런스의 핵심인 우측 콘크리트 건물과 통로가 보이지 않고 흐릿한 배경으로 대체됨.",
    "entities": "현우의 외형과 의상이 캐릭터 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "바닥에 앉아있으며, 오른손은 가슴 위에 펴진 상태로 얹혀 있고 강하게 쥐는 물리적 작용이 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴을 강하게 움켜쥔 손의 동작과 고통스럽게 숨을 들이마시는 표정이 텍스트와 정확히 일치하며, 지정된 장소의 건축물과 배경을 완벽하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지정된 배경 레퍼런스의 주요 건축물이 누락되어 장소의 일치도가 떨어지며, 가슴을 강하게 움켜쥐라는 지시와 달리 손을 그저 가슴 위에 평평하게 얹고만 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 아래쪽을 향하고 있으며, 입을 크게 벌리고 가쁜 숨을 쉬는 표정임.",
        "built_space": "레퍼런스 이미지에 있는 우측의 콘크리트 건물 외벽과 바닥의 통로, 좌측의 쓰레기 더미가 정확한 위치에 존재함.",
        "entities": "현우의 얼굴, 헤어스타일, 회색 셔츠 및 바지가 캐릭터 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "바닥에 주저앉아 한 손은 가슴의 옷깃을 강하게 쥐고 있으며, 다른 한 손은 바닥을 짚어 체중을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 정면 약간 좌측을 향하고, 입을 벌려 숨을 쉬고 있음.",
        "built_space": "쓰레기 더미는 보이나, 위치 레퍼런스의 핵심인 우측 콘크리트 건물과 통로가 보이지 않고 흐릿한 배경으로 대체됨.",
        "entities": "현우의 외형과 의상이 캐릭터 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "바닥에 앉아있으며, 오른손은 가슴 위에 펴진 상태로 얹혀 있고 강하게 쥐는 물리적 작용이 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴을 강하게 움켜쥔 손의 동작과 고통스럽게 숨을 들이마시는 표정이 텍스트와 정확히 일치하며, 지정된 장소의 건축물과 배경을 완벽하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지정된 배경 레퍼런스의 주요 건축물이 누락되어 장소의 일치도가 떨어지며, 가슴을 강하게 움켜쥐라는 지시와 달리 손을 그저 가슴 위에 평평하게 얹고만 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 아래쪽을 향하고 있으며, 입을 크게 벌리고 가쁜 숨을 쉬는 표정임.",
        "built_space": "레퍼런스 이미지에 있는 우측의 콘크리트 건물 외벽과 바닥의 통로, 좌측의 쓰레기 더미가 정확한 위치에 존재함.",
        "entities": "현우의 얼굴, 헤어스타일, 회색 셔츠 및 바지가 캐릭터 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "바닥에 주저앉아 한 손은 가슴의 옷깃을 강하게 쥐고 있으며, 다른 한 손은 바닥을 짚어 체중을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 정면 약간 좌측을 향하고, 입을 벌려 숨을 쉬고 있음.",
        "built_space": "쓰레기 더미는 보이나, 위치 레퍼런스의 핵심인 우측 콘크리트 건물과 통로가 보이지 않고 흐릿한 배경으로 대체됨.",
        "entities": "현우의 외형과 의상이 캐릭터 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "바닥에 앉아있으며, 오른손은 가슴 위에 펴진 상태로 얹혀 있고 강하게 쥐는 물리적 작용이 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "상체 중심 구도와 벌어진 입은 맞지만, 가슴을 움켜쥐기보다 손바닥으로 누르고 있으며 앉은 몸 아래의 지면이 거의 드러나지 않는다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "가슴의 옷감을 강하게 움켜쥔 손과 고통으로 일그러진 얼굴, 지면을 짚고 일어난 자세가 핵심 순간을 충실히 구현하며 법원 외벽도 참조와 부합한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이고 시선은 화면 왼쪽의 프레임 밖으로 향한다. 특정 응시 대상은 보이지 않으며 지시된 대상도 없다. 오른손은 자기 왼쪽 가슴에 닿지만 손가락을 펼쳐 누르는 형태라 강하게 움켜쥐는 동작과는 차이가 있다.",
        "built_space": "인물의 머리부터 허리 부근까지 담은 상체 중심 구도다. 양어깨 뒤와 화면 가장자리에 쓰레기 더미가 있고, 뒤쪽에는 빈 마당과 울타리, 나무, 건물 일부가 보인다. 참조의 오른쪽 법원 출입문과 차양은 프레임에 없어 그 설비의 일치는 확인할 수 없다. 앉은 몸 바로 아래 지면은 거의 가려져 있다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "검은 머리의 동아시아계 외형을 가진 젊은 남성 한 명만 있다. 회색 단추 셔츠와 하단에 조금 보이는 올리브색 바지는 참조와 부합한다. 얼굴은 참조보다 다소 성숙하고 각져 보인다. 얼굴과 손에 상처가 보이고 입은 크게 벌어져 있으나, 눈을 뜬 표정은 극심한 통증보다 놀람도 강하게 읽힌다. 다리의 물린 상처와 주머니 속 지갑·신분증·현금은 보이지 않아 판단하지 않는다. 읽을 수 있는 글자는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "하단의 굽힌 무릎과 낮게 놓인 몸통은 지면에 앉은 자세로 읽힌다. 엉덩이와 반대 손의 접촉점은 화면 밖이지만 공중에 떠 있다는 징후는 없다. 가슴에 댄 손은 몸과 정상적으로 접촉하고 팔의 연결도 자연스럽다. 다만 펼친 손 주변의 옷감에는 강한 움켜쥠에 따른 당김이 뚜렷하지 않다."
       },
       {
        "label": "B",
        "direction": "고개가 앞으로 숙여지고 눈은 통증으로 감겨 있어 외부의 응시 대상은 없다. 오른손은 자기 왼쪽 가슴의 셔츠를 향해 모여 옷감을 움켜쥐고, 왼손은 몸 옆뒤 지면을 짚는다. 벌어진 입과 수축한 얼굴 근육이 고통 속에 크게 숨을 들이마시는 순간에 부합한다.",
        "built_space": "상체를 중심으로 골반과 굽힌 다리 일부까지 포함한 미디엄 구도다. 몸 아래와 오른쪽에 거친 지면이 보이며, 쓰레기 더미는 왼쪽 가장자리와 어깨 뒤 먼 배경에 남아 있다. 오른쪽에는 참조와 대응하는 콘크리트 외벽, 닫힌 금속문 한 개, 문 위 차양 한 개, 세로 배관이 보인다. 벽을 따라 이어지는 포장 띠와 뒤편 울타리도 참조 공간에 부합한다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "검은 머리의 동아시아계 외형을 가진 젊은 남성 한 명만 보인다. 얼굴은 통증으로 찡그려져 정확한 동일인 비교에 제한이 있으나 참조의 머리색과 체격, 회색 셔츠와 올리브색 바지는 대체로 일치한다. 옷은 더러워져 있고 바지 무릎 부근에 손상이 보이지만 개에게 물린 상처인지는 단정할 수 없다. 주머니 속 지갑과 내용물은 노출되지 않는다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "골반과 굽힌 다리가 지면에 놓이고 왼손 손바닥이 바닥을 짚어 몸통을 지탱한다. 누웠다가 상체를 일으킨 직후의 체중 배분으로 가능하다. 오른손 손가락은 셔츠를 실제로 잡고 있으며 가슴 쪽 옷감에 당김과 주름이 생긴다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "상체 중심 구도와 벌어진 입은 맞지만, 가슴을 움켜쥐기보다 손바닥으로 누르고 있으며 앉은 몸 아래의 지면이 거의 드러나지 않는다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "가슴의 옷감을 강하게 움켜쥔 손과 고통으로 일그러진 얼굴, 지면을 짚고 일어난 자세가 핵심 순간을 충실히 구현하며 법원 외벽도 참조와 부합한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이고 시선은 화면 왼쪽의 프레임 밖으로 향한다. 특정 응시 대상은 보이지 않으며 지시된 대상도 없다. 오른손은 자기 왼쪽 가슴에 닿지만 손가락을 펼쳐 누르는 형태라 강하게 움켜쥐는 동작과는 차이가 있다.",
        "built_space": "인물의 머리부터 허리 부근까지 담은 상체 중심 구도다. 양어깨 뒤와 화면 가장자리에 쓰레기 더미가 있고, 뒤쪽에는 빈 마당과 울타리, 나무, 건물 일부가 보인다. 참조의 오른쪽 법원 출입문과 차양은 프레임에 없어 그 설비의 일치는 확인할 수 없다. 앉은 몸 바로 아래 지면은 거의 가려져 있다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "검은 머리의 동아시아계 외형을 가진 젊은 남성 한 명만 있다. 회색 단추 셔츠와 하단에 조금 보이는 올리브색 바지는 참조와 부합한다. 얼굴은 참조보다 다소 성숙하고 각져 보인다. 얼굴과 손에 상처가 보이고 입은 크게 벌어져 있으나, 눈을 뜬 표정은 극심한 통증보다 놀람도 강하게 읽힌다. 다리의 물린 상처와 주머니 속 지갑·신분증·현금은 보이지 않아 판단하지 않는다. 읽을 수 있는 글자는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "하단의 굽힌 무릎과 낮게 놓인 몸통은 지면에 앉은 자세로 읽힌다. 엉덩이와 반대 손의 접촉점은 화면 밖이지만 공중에 떠 있다는 징후는 없다. 가슴에 댄 손은 몸과 정상적으로 접촉하고 팔의 연결도 자연스럽다. 다만 펼친 손 주변의 옷감에는 강한 움켜쥠에 따른 당김이 뚜렷하지 않다."
       },
       {
        "label": "A",
        "direction": "고개가 앞으로 숙여지고 눈은 통증으로 감겨 있어 외부의 응시 대상은 없다. 오른손은 자기 왼쪽 가슴의 셔츠를 향해 모여 옷감을 움켜쥐고, 왼손은 몸 옆뒤 지면을 짚는다. 벌어진 입과 수축한 얼굴 근육이 고통 속에 크게 숨을 들이마시는 순간에 부합한다.",
        "built_space": "상체를 중심으로 골반과 굽힌 다리 일부까지 포함한 미디엄 구도다. 몸 아래와 오른쪽에 거친 지면이 보이며, 쓰레기 더미는 왼쪽 가장자리와 어깨 뒤 먼 배경에 남아 있다. 오른쪽에는 참조와 대응하는 콘크리트 외벽, 닫힌 금속문 한 개, 문 위 차양 한 개, 세로 배관이 보인다. 벽을 따라 이어지는 포장 띠와 뒤편 울타리도 참조 공간에 부합한다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "검은 머리의 동아시아계 외형을 가진 젊은 남성 한 명만 보인다. 얼굴은 통증으로 찡그려져 정확한 동일인 비교에 제한이 있으나 참조의 머리색과 체격, 회색 셔츠와 올리브색 바지는 대체로 일치한다. 옷은 더러워져 있고 바지 무릎 부근에 손상이 보이지만 개에게 물린 상처인지는 단정할 수 없다. 주머니 속 지갑과 내용물은 노출되지 않는다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "골반과 굽힌 다리가 지면에 놓이고 왼손 손바닥이 바닥을 짚어 몸통을 지탱한다. 누웠다가 상체를 일으킨 직후의 체중 배분으로 가능하다. 오른손 손가락은 셔츠를 실제로 잡고 있으며 가슴 쪽 옷감에 당김과 주름이 생긴다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.349
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.349
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1349
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "가슴을 강하게 움켜쥔 손의 동작과 고통스럽게 숨을 들이마시는 표정이 텍스트와 정확히 일치하며, 지정된 장소의 건축물과 배경을 완벽하게 재현했습니다."
   },
   {
    "label": "B",
    "score": 1349,
    "verdict_ko": "지정된 배경 레퍼런스의 주요 건축물이 누락되어 장소의 일치도가 떨어지며, 가슴을 강하게 움켜쥐라는 지시와 달리 손을 그저 가슴 위에 평평하게 얹고만 있습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L161B01.png",
    "asset_id": "595d4bdd-0afb-4da7-95e6-a460d0259575",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-acb5-7e74-9036-039070d279a9",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2__bgfirst_bg.png",
   "bg_asset_id": "74193d07-10b1-46b2-a9cb-c0e151291eb3",
   "bg_record_key": "S9sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S9sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:13:31.498571+00:00",
  "fingerprint": "6dcc350310c59cfb3e7b9901ceb69ef005b5e2ee6fd6f721da8b760a70ea6d2d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S9sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S9sh2_sel.png",
  "source_sha256": "ae90b30fd69f1a84122f997edab2f45df90bdaf773a4aa94f20cb6107e52b0a0",
  "file": "S9sh2_cine.png",
  "staged_sha256": "e855137ba1a576b434a3902f3d9cf67c57696809b68cc6fb951d6403e0179c67",
  "latency_ms": 11065
 },
 "S9sh5::signage": {
  "fp": "cc021fcecb804d26",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "신분증"
   }
  ],
  "dropped": []
 },
 "S9sh5": {
  "input_fingerprint": "3abb8a6de94a2a92",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지폐가 빠져나가 신분증만 덩그러니 남은 빈 가죽 지갑이 허공을 향해 뻗은 현우의 손끝에서 막 떨어져 날아가는 mid-action 순간.\n\nLOCATION (lock): Beside the desolate rubbish heaps outside the refugee tribunal at night, where the emptied wallet is discarded. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Leather wallet (Released from the fingertips with the money removed and identification remaining) — Its open inner side is angled toward the camera during the first instant of separation; used as Marks the release while remaining small enough for hand and body scale to read naturally; Judge's identification (Still inside the discarded wallet) — The identification-bearing face is partly visible within the open wallet; no specific text is supplied; used as Distinguishes the retained identification from the removed money; Rubbish heaps (Remain around the court exterior) — Visible beyond the hand and the wallet's open travel space; used as Maintains exterior continuity and prevents the action from becoming an isolated product image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient treatment, using controlled tonal separation to distinguish the fingers, leather wallet, and remaining identification.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The discarded wallet still contains the judge's identification but no longer contains its cash. Nighttime rubbish heaps surround the exterior of the refugee court. 현우: Hyunwoo has retained the cash removed from the wallet and is alert, but still limps on his dog-bitten leg and retains the earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지폐가 빠져나가 신분증만 덩그러니 남은 빈 가죽 지갑이 허공을 향해 뻗은 현우의 손끝에서 막 떨어져 날아가는 mid-action 순간.\n\nLOCATION (lock): Beside the desolate rubbish heaps outside the refugee tribunal at night, where the emptied wallet is discarded. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Leather wallet (Released from the fingertips with the money removed and identification remaining) — Its open inner side is angled toward the camera during the first instant of separation; used as Marks the release while remaining small enough for hand and body scale to read naturally; Judge's identification (Still inside the discarded wallet) — The identification-bearing face is partly visible within the open wallet; no specific text is supplied; used as Distinguishes the retained identification from the removed money; Rubbish heaps (Remain around the court exterior) — Visible beyond the hand and the wallet's open travel space; used as Maintains exterior continuity and prevents the action from becoming an isolated product image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient treatment, using controlled tonal separation to distinguish the fingers, leather wallet, and remaining identification.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The discarded wallet still contains the judge's identification but no longer contains its cash. Nighttime rubbish heaps surround the exterior of the refugee court. 현우: Hyunwoo has retained the cash removed from the wallet and is alert, but still limps on his dog-bitten leg and retains the earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지폐가 빠져나가 신분증만 덩그러니 남은 빈 가죽 지갑이 허공을 향해 뻗은 현우의 손끝에서 막 떨어져 날아가는 mid-action 순간.\n\nLOCATION (lock): Beside the desolate rubbish heaps outside the refugee tribunal at night, where the emptied wallet is discarded. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Leather wallet (Released from the fingertips with the money removed and identification remaining) — Its open inner side is angled toward the camera during the first instant of separation; used as Marks the release while remaining small enough for hand and body scale to read naturally; Judge's identification (Still inside the discarded wallet) — The identification-bearing face is partly visible within the open wallet; no specific text is supplied; used as Distinguishes the retained identification from the removed money; Rubbish heaps (Remain around the court exterior) — Visible beyond the hand and the wallet's open travel space; used as Maintains exterior continuity and prevents the action from becoming an isolated product image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient treatment, using controlled tonal separation to distinguish the fingers, leather wallet, and remaining identification.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The discarded wallet still contains the judge's identification but no longer contains its cash. Nighttime rubbish heaps surround the exterior of the refugee court. 현우: Hyunwoo has retained the cash removed from the wallet and is alert, but still limps on his dog-bitten leg and retains the earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 아래로 떨어지는 지갑을 향하고 있으며, 손은 지갑을 방금 놓은 형태를 취하고 있음.",
    "built_space": "우측에 콘크리트 건물 벽면, 좌측에 쓰레기 더미가 배치되어 이전 샷의 야외 배경과 구조적으로 일치함.",
    "entities": "현우의 얼굴, 헤어스타일, 의상(오염된 셔츠)이 레퍼런스와 일치함. 열린 가죽 지갑 내부에 신분증이 있으며, 글씨는 의도적으로 읽을 수 없게 처리됨.",
    "hard_violations": [],
    "physics": "지갑은 손가락 바로 아래 허공에 떠 있으며, 방금 놓아서 중력에 의해 낙하하는 움직임을 자연스럽게 보여줌."
   },
   {
    "label": "B",
    "direction": "시선은 지갑 쪽을 향하나 약간 엇나가 있으며, 팔을 뻗고 손을 넓게 펴고 있음.",
    "built_space": "우측의 건물과 좌측의 쓰레기 더미 등 참조된 배경의 구조적 요소를 적절히 재현함.",
    "entities": "현우의 외모와 의상은 레퍼런스와 일치함. 지갑 안의 신분증에 'Judge's ID'라는 프롬프트 텍스트가 명확하게 렌더링됨.",
    "hard_violations": [
     "[gemini-pro] 프롬프트 지시어('Judge's ID')가 소품에 읽을 수 있는 텍스트로 직접 유출됨 (텍스트 유출 및 읽을 수 있는 글씨 금지 위반)."
    ],
    "physics": "지갑이 허공에 떠 있으나 손에서 상당히 멀리 떨어져 있어 '막 떨어져 날아가는 첫 순간'의 물리적 거리와 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레임과 '손끝에서 막 떨어진 첫 순간'을 훌륭하게 포착했으며 식별할 수 없는 텍스트 지침도 잘 준수했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지갑이 손에서 너무 멀리 떨어져 타이밍 지침에 어긋나며, 신분증에 프롬프트 텍스트가 영어로 유출되어 명백한 위반입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 아래로 떨어지는 지갑을 향하고 있으며, 손은 지갑을 방금 놓은 형태를 취하고 있음.",
        "built_space": "우측에 콘크리트 건물 벽면, 좌측에 쓰레기 더미가 배치되어 이전 샷의 야외 배경과 구조적으로 일치함.",
        "entities": "현우의 얼굴, 헤어스타일, 의상(오염된 셔츠)이 레퍼런스와 일치함. 열린 가죽 지갑 내부에 신분증이 있으며, 글씨는 의도적으로 읽을 수 없게 처리됨.",
        "hard_violations": [],
        "physics": "지갑은 손가락 바로 아래 허공에 떠 있으며, 방금 놓아서 중력에 의해 낙하하는 움직임을 자연스럽게 보여줌."
       },
       {
        "label": "B",
        "direction": "시선은 지갑 쪽을 향하나 약간 엇나가 있으며, 팔을 뻗고 손을 넓게 펴고 있음.",
        "built_space": "우측의 건물과 좌측의 쓰레기 더미 등 참조된 배경의 구조적 요소를 적절히 재현함.",
        "entities": "현우의 외모와 의상은 레퍼런스와 일치함. 지갑 안의 신분증에 'Judge's ID'라는 프롬프트 텍스트가 명확하게 렌더링됨.",
        "hard_violations": [
         "프롬프트 지시어('Judge's ID')가 소품에 읽을 수 있는 텍스트로 직접 유출됨 (텍스트 유출 및 읽을 수 있는 글씨 금지 위반)."
        ],
        "physics": "지갑이 허공에 떠 있으나 손에서 상당히 멀리 떨어져 있어 '막 떨어져 날아가는 첫 순간'의 물리적 거리와 맞지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레임과 '손끝에서 막 떨어진 첫 순간'을 훌륭하게 포착했으며 식별할 수 없는 텍스트 지침도 잘 준수했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지갑이 손에서 너무 멀리 떨어져 타이밍 지침에 어긋나며, 신분증에 프롬프트 텍스트가 영어로 유출되어 명백한 위반입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 아래로 떨어지는 지갑을 향하고 있으며, 손은 지갑을 방금 놓은 형태를 취하고 있음.",
        "built_space": "우측에 콘크리트 건물 벽면, 좌측에 쓰레기 더미가 배치되어 이전 샷의 야외 배경과 구조적으로 일치함.",
        "entities": "현우의 얼굴, 헤어스타일, 의상(오염된 셔츠)이 레퍼런스와 일치함. 열린 가죽 지갑 내부에 신분증이 있으며, 글씨는 의도적으로 읽을 수 없게 처리됨.",
        "hard_violations": [],
        "physics": "지갑은 손가락 바로 아래 허공에 떠 있으며, 방금 놓아서 중력에 의해 낙하하는 움직임을 자연스럽게 보여줌."
       },
       {
        "label": "B",
        "direction": "시선은 지갑 쪽을 향하나 약간 엇나가 있으며, 팔을 뻗고 손을 넓게 펴고 있음.",
        "built_space": "우측의 건물과 좌측의 쓰레기 더미 등 참조된 배경의 구조적 요소를 적절히 재현함.",
        "entities": "현우의 외모와 의상은 레퍼런스와 일치함. 지갑 안의 신분증에 'Judge's ID'라는 프롬프트 텍스트가 명확하게 렌더링됨.",
        "hard_violations": [
         "프롬프트 지시어('Judge's ID')가 소품에 읽을 수 있는 텍스트로 직접 유출됨 (텍스트 유출 및 읽을 수 있는 글씨 금지 위반)."
        ],
        "physics": "지갑이 허공에 떠 있으나 손에서 상당히 멀리 떨어져 있어 '막 떨어져 날아가는 첫 순간'의 물리적 거리와 맞지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지갑 내부와 남은 신분증, 야간 쓰레기장은 구현했지만 허벅지까지 포함한 넓은 구도와 손에서 이미 멀어진 지갑이 지정된 클로즈업·분리 첫 순간에 덜 충실합니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "손·열린 지갑·얼굴 일부를 묶은 클로즈업과 쓰레기 더미로 향하는 투척 동작이 더 충실하지만, 지갑은 손끝에서 막 떨어진 순간치고 이미 다소 멀리 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 인물의 팔과 펼친 손가락은 화면 왼쪽으로 뻗어 있고, 지갑은 그보다 왼쪽 아래의 쓰레기 더미 쪽으로 날아가는 배치입니다. 지갑의 열린 안쪽과 신분증 면은 카메라를 향합니다. 현우의 시선은 왼쪽으로 향하지만 낮게 놓인 지갑 자체보다는 그 위쪽이나 화면 밖을 보는 듯합니다.",
        "built_space": "왼쪽에 쓰레기 더미, 중앙에 빈 통로와 뒤쪽 울타리·나무, 오른쪽 가장자리에 콘크리트 외벽과 수직 배관이 보입니다. 문과 차양은 이 구도에서 확인되지 않습니다. 중앙 뒤쪽에 밝은 외부등 하나가 두드러집니다. 참조의 외부 공간 배치는 대체로 이어지지만 인물의 허벅지까지 보여 지정된 클로즈업보다 넓습니다.",
        "entities": "검은 머리의 젊은 동아시아계 남성 한 명이며, 회색의 때 탄 셔츠와 올리브색 바지는 참조와 부합합니다. 얼굴에 상처와 오염이 보입니다. 갈색 가죽 지갑 하나에 사진이 있는 신분증 한 장이 남아 있고 지폐는 보이지 않습니다. 신분증의 구체적인 신원은 확인할 수 없으며 글자는 명확히 판독되지 않습니다. 지갑이 손에 비해 시각적으로 크게 강조됩니다. 보유한 현금과 다리 부상은 이 화면만으로 확인되지 않습니다.",
        "hard_violations": [],
        "physics": "팔은 몸통에 자연스럽게 연결되고 손가락은 물건을 놓은 뒤 펼쳐진 형태입니다. 지갑은 손의 투척으로 왼쪽 아래로 이동하여 쓰레기 더미나 지면에 떨어질 수 있으므로 근거 없는 공중 부양은 아닙니다. 신분증은 지갑의 수납창에 고정되어 있습니다. 다만 손끝과 지갑 사이의 간격은 분리 직후보다 투척이 조금 진행된 순간처럼 보입니다."
       },
       {
        "label": "B",
        "direction": "들린 손가락은 왼쪽 위로 펼쳐져 있고, 열린 지갑은 손의 왼쪽 아래에서 쓰레기 더미 쪽으로 떨어지는 위치입니다. 현우의 고개와 눈은 왼쪽 아래의 지갑 이동 공간을 향합니다. 지갑 안쪽과 신분증의 사진 면이 비스듬히 카메라를 향해 요구된 방향을 충족합니다.",
        "built_space": "손과 지갑 뒤로 검은 쓰레기봉투, 상자와 폐자재가 쌓여 있고 그 뒤에 낡은 콘크리트 외벽이 있습니다. 왼쪽의 어두운 출입구 하나, 중앙의 회색 문 하나, 오른쪽의 차양 하나와 좁은 창 하나가 구별되며 상단에도 창들이 일부 보입니다. 참조와 같은 재료와 황폐한 야간 외부 환경이 유지됩니다. 손·지갑·얼굴 일부를 중심으로 잘라 클로즈업 요구에 더 가깝습니다.",
        "entities": "검은 머리의 젊은 동아시아계 남성 한 명의 얼굴 일부와 상체가 보이고, 낡은 회색 셔츠와 얼굴의 오염은 참조에 부합합니다. 갈색 가죽 지갑 하나의 내부에 사진 신분증 한 장이 남아 있으며 지폐는 보이지 않습니다. 신분증 글자는 흐려 구체적인 신원을 판독할 수 없습니다. 다른 실제 인물이나 자막은 없습니다. 다리와 보유 현금은 구도 밖이므로 확인 대상이 아닙니다.",
        "hard_violations": [],
        "physics": "굽힌 팔과 놓은 뒤 벌어진 손가락이 투척 동작의 출발을 설명합니다. 지갑은 손에서 왼쪽 아래로 날아가 아래 쓰레기 더미에 떨어질 수 있는 위치여서 지지 없는 정지 부양으로 보이지 않습니다. 신분증은 지갑 수납부에 끼워져 있습니다. 손과 지갑의 간격은 분리 첫 순간을 엄밀하게 표현하기에는 다소 큽니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지갑 내부와 남은 신분증, 야간 쓰레기장은 구현했지만 허벅지까지 포함한 넓은 구도와 손에서 이미 멀어진 지갑이 지정된 클로즈업·분리 첫 순간에 덜 충실합니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "손·열린 지갑·얼굴 일부를 묶은 클로즈업과 쓰레기 더미로 향하는 투척 동작이 더 충실하지만, 지갑은 손끝에서 막 떨어진 순간치고 이미 다소 멀리 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 인물의 팔과 펼친 손가락은 화면 왼쪽으로 뻗어 있고, 지갑은 그보다 왼쪽 아래의 쓰레기 더미 쪽으로 날아가는 배치입니다. 지갑의 열린 안쪽과 신분증 면은 카메라를 향합니다. 현우의 시선은 왼쪽으로 향하지만 낮게 놓인 지갑 자체보다는 그 위쪽이나 화면 밖을 보는 듯합니다.",
        "built_space": "왼쪽에 쓰레기 더미, 중앙에 빈 통로와 뒤쪽 울타리·나무, 오른쪽 가장자리에 콘크리트 외벽과 수직 배관이 보입니다. 문과 차양은 이 구도에서 확인되지 않습니다. 중앙 뒤쪽에 밝은 외부등 하나가 두드러집니다. 참조의 외부 공간 배치는 대체로 이어지지만 인물의 허벅지까지 보여 지정된 클로즈업보다 넓습니다.",
        "entities": "검은 머리의 젊은 동아시아계 남성 한 명이며, 회색의 때 탄 셔츠와 올리브색 바지는 참조와 부합합니다. 얼굴에 상처와 오염이 보입니다. 갈색 가죽 지갑 하나에 사진이 있는 신분증 한 장이 남아 있고 지폐는 보이지 않습니다. 신분증의 구체적인 신원은 확인할 수 없으며 글자는 명확히 판독되지 않습니다. 지갑이 손에 비해 시각적으로 크게 강조됩니다. 보유한 현금과 다리 부상은 이 화면만으로 확인되지 않습니다.",
        "hard_violations": [],
        "physics": "팔은 몸통에 자연스럽게 연결되고 손가락은 물건을 놓은 뒤 펼쳐진 형태입니다. 지갑은 손의 투척으로 왼쪽 아래로 이동하여 쓰레기 더미나 지면에 떨어질 수 있으므로 근거 없는 공중 부양은 아닙니다. 신분증은 지갑의 수납창에 고정되어 있습니다. 다만 손끝과 지갑 사이의 간격은 분리 직후보다 투척이 조금 진행된 순간처럼 보입니다."
       },
       {
        "label": "A",
        "direction": "들린 손가락은 왼쪽 위로 펼쳐져 있고, 열린 지갑은 손의 왼쪽 아래에서 쓰레기 더미 쪽으로 떨어지는 위치입니다. 현우의 고개와 눈은 왼쪽 아래의 지갑 이동 공간을 향합니다. 지갑 안쪽과 신분증의 사진 면이 비스듬히 카메라를 향해 요구된 방향을 충족합니다.",
        "built_space": "손과 지갑 뒤로 검은 쓰레기봉투, 상자와 폐자재가 쌓여 있고 그 뒤에 낡은 콘크리트 외벽이 있습니다. 왼쪽의 어두운 출입구 하나, 중앙의 회색 문 하나, 오른쪽의 차양 하나와 좁은 창 하나가 구별되며 상단에도 창들이 일부 보입니다. 참조와 같은 재료와 황폐한 야간 외부 환경이 유지됩니다. 손·지갑·얼굴 일부를 중심으로 잘라 클로즈업 요구에 더 가깝습니다.",
        "entities": "검은 머리의 젊은 동아시아계 남성 한 명의 얼굴 일부와 상체가 보이고, 낡은 회색 셔츠와 얼굴의 오염은 참조에 부합합니다. 갈색 가죽 지갑 하나의 내부에 사진 신분증 한 장이 남아 있으며 지폐는 보이지 않습니다. 신분증 글자는 흐려 구체적인 신원을 판독할 수 없습니다. 다른 실제 인물이나 자막은 없습니다. 다리와 보유 현금은 구도 밖이므로 확인 대상이 아닙니다.",
        "hard_violations": [],
        "physics": "굽힌 팔과 놓은 뒤 벌어진 손가락이 투척 동작의 출발을 설명합니다. 지갑은 손에서 왼쪽 아래로 날아가 아래 쓰레기 더미에 떨어질 수 있는 위치여서 지지 없는 정지 부양으로 보이지 않습니다. 신분증은 지갑 수납부에 끼워져 있습니다. 손과 지갑의 간격은 분리 첫 순간을 엄밀하게 표현하기에는 다소 큽니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.179
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.929
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트 지시어('Judge's ID')가 소품에 읽을 수 있는 텍스트로 직접 유출됨 (텍스트 유출 및 읽을 수 있는 글씨 금지 위반)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 프레임과 '손끝에서 막 떨어진 첫 순간'을 훌륭하게 포착했으며 식별할 수 없는 텍스트 지침도 잘 준수했습니다."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "지갑이 손에서 너무 멀리 떨어져 타이밍 지침에 어긋나며, 신분증에 프롬프트 텍스트가 영어로 유출되어 명백한 위반입니다.  ★위반: [gemini-pro] 프롬프트 지시어('Judge's ID')가 소품에 읽을 수 있는 텍스트로 직접 유출됨 (텍스트 유출 및 읽을 수 있는 글씨 금지 위반)."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2_sel.png",
    "asset_id": "d892a0c9-711d-41d4-aab7-ce078b31fa44",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-b008-72ad-81e5-c5b987c31ba6",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S9sh2"
  }
 },
 "S9sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:14:27.712744+00:00",
  "fingerprint": "25c1a22ff9f0566bec93d531b48126b8e5243c48e36777e887356b5c18347ca2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S9sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S9sh5_sel.png",
  "source_sha256": "595d597621a81b9d7bcc0c4cc7632593eeea1d622f9f8099234fd6f7124b56d4",
  "file": "S9sh5_cine.png",
  "staged_sha256": "17ef31b663a85fb533110fb026d809419025d6a8180df983cd18e0a25997ee48",
  "latency_ms": 11310
 },
 "S10sh3::signage": {
  "fp": "c5a6300222d7094b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::f65f2d6badb3b97d": {
  "subjects": [],
  "subject_text": "허름한 골목 식당 내부\n낡은 탁자와 의자가 놓인 소규모 식당. 출입구에서 떨어진 안쪽 구석에 외진 좌석이 있다.",
  "identity": "canonical",
  "scope_id": "L162",
  "scope_role": "location_interior",
  "scope_sha": "e4027065f7112ec7"
 },
 "S10sh3::bgfirst_bg": {
  "input_fingerprint": "8929de6d3f312bbf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3__bgfirst_bg.png",
  "asset_id": "33a03a51-7f51-4fab-9373-3f1601df1a45",
  "input_asset_ids": [
   "2c3ae081-487d-467f-b6d8-4c65de90d776",
   "7a16f2f4-24b6-4a89-a184-9204ad75cbe4"
  ]
 },
 "S10sh3": {
  "input_fingerprint": "3ccf993e82f5c258",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A cash-filled envelope is on the table in an out-of-the-way part of the shabby restaurant. 현우: Hyunwoo is seated at the table and has taken hold of the money envelope again; his dog-bitten leg and earlier head injury remain untreated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A cash-filled envelope is on the table in an out-of-the-way part of the shabby restaurant. 현우: Hyunwoo is seated at the table and has taken hold of the money envelope again; his dog-bitten leg and earlier head injury remain untreated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A cash-filled envelope is on the table in an out-of-the-way part of the shabby restaurant. 현우: Hyunwoo is seated at the table and has taken hold of the money envelope again; his dog-bitten leg and earlier head injury remain untreated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3__bgfirst_bg.png",
     "asset_id": "33a03a51-7f51-4fab-9373-3f1601df1a45",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S10sh3.png",
     "asset_id": "2c3ae081-487d-467f-b6d8-4c65de90d776",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L162B01.png",
     "asset_id": "7a16f2f4-24b6-4a89-a184-9204ad75cbe4",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물들의 시선은 프레임 밖이며, 두 팔이 테이블 중앙의 한 지점을 향해 뻗어 있음.",
    "built_space": "식당 내부. 나무 테이블과 의자, 배경의 주방 입구 등 위치 레퍼런스의 공간 형태가 반영됨.",
    "entities": "두 사람의 팔, 테이블 위에 놓인 두꺼운 종이 패키지 형태의 봉투(지폐 여부 불분명). 손만 등장하여 신원 확인 불가.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy: 맞잡은 두 손의 손가락들이 서로 융합되고 관절 구조가 뭉개져 손의 형태가 해부학적으로 불가능함"
    ],
    "physics": "왼쪽에서 온 손이 오른쪽에서 온 손의 바닥과 손가락을 부드럽게 감싸 쥔 상태로 테이블 위에 기대어 있으며, 상대방을 다급하게 제압하는 힘이나 동작이 전혀 읽히지 않음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 테이블 위의 맞잡은 손과 돈 봉투를 향해 아래로 고정되어 있음.",
    "built_space": "식당 내부. 나무 테이블, 빨간색 쿠션의 금속 의자, 벽면의 한국어 메뉴판 등 위치 레퍼런스의 공간 구조와 기물들이 정확히 배치됨.",
    "entities": "현우(레퍼런스와 일치하는 얼굴, 헤어스타일, 흙 묻은 회색 셔츠 및 얼굴의 상처), 지폐가 드러난 돈 봉투, 뻗어 나온 상대방의 왼팔.",
    "hard_violations": [
     "[gpt-high] 벽 메뉴에 읽을 수 있는 한글과 가격 숫자가 남아 있어 가독성 있는 문자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "현우의 왼손이 돈 봉투를 덮고 있는 상대방의 왼쪽 손목을 위에서 아래로 체중을 실어 강하게 짓누르고 있으며, 동작과 지지면이 물리적으로 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 손 클로즈업 숏보다 화각이 넓어져 인물의 상반신과 얼굴까지 노출되었으나, 돈 봉투 위로 뻗은 손등을 강하게 내리누르는 핵심 액션과 위치 레퍼런스를 정확하게 구현하였습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "손 클로즈업이라는 프레임 스케일은 맞추었으나, 두 손의 구조가 해부학적으로 붕괴되었고 봉투를 제압하는 대신 손을 쥐고 있는 듯한 엉뚱한 액션을 연출하여 탈락입니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 테이블 위의 맞잡은 손과 돈 봉투를 향해 아래로 고정되어 있음.",
        "built_space": "식당 내부. 나무 테이블, 빨간색 쿠션의 금속 의자, 벽면의 한국어 메뉴판 등 위치 레퍼런스의 공간 구조와 기물들이 정확히 배치됨.",
        "entities": "현우(레퍼런스와 일치하는 얼굴, 헤어스타일, 흙 묻은 회색 셔츠 및 얼굴의 상처), 지폐가 드러난 돈 봉투, 뻗어 나온 상대방의 왼팔.",
        "hard_violations": [],
        "physics": "현우의 왼손이 돈 봉투를 덮고 있는 상대방의 왼쪽 손목을 위에서 아래로 체중을 실어 강하게 짓누르고 있으며, 동작과 지지면이 물리적으로 자연스러움."
       },
       {
        "label": "A",
        "direction": "인물들의 시선은 프레임 밖이며, 두 팔이 테이블 중앙의 한 지점을 향해 뻗어 있음.",
        "built_space": "식당 내부. 나무 테이블과 의자, 배경의 주방 입구 등 위치 레퍼런스의 공간 형태가 반영됨.",
        "entities": "두 사람의 팔, 테이블 위에 놓인 두꺼운 종이 패키지 형태의 봉투(지폐 여부 불분명). 손만 등장하여 신원 확인 불가.",
        "hard_violations": [
         "physically impossible anatomy: 맞잡은 두 손의 손가락들이 서로 융합되고 관절 구조가 뭉개져 손의 형태가 해부학적으로 불가능함"
        ],
        "physics": "왼쪽에서 온 손이 오른쪽에서 온 손의 바닥과 손가락을 부드럽게 감싸 쥔 상태로 테이블 위에 기대어 있으며, 상대방을 다급하게 제압하는 힘이나 동작이 전혀 읽히지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 손 클로즈업 숏보다 화각이 넓어져 인물의 상반신과 얼굴까지 노출되었으나, 돈 봉투 위로 뻗은 손등을 강하게 내리누르는 핵심 액션과 위치 레퍼런스를 정확하게 구현하였습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "손 클로즈업이라는 프레임 스케일은 맞추었으나, 두 손의 구조가 해부학적으로 붕괴되었고 봉투를 제압하는 대신 손을 쥐고 있는 듯한 엉뚱한 액션을 연출하여 탈락입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 테이블 위의 맞잡은 손과 돈 봉투를 향해 아래로 고정되어 있음.",
        "built_space": "식당 내부. 나무 테이블, 빨간색 쿠션의 금속 의자, 벽면의 한국어 메뉴판 등 위치 레퍼런스의 공간 구조와 기물들이 정확히 배치됨.",
        "entities": "현우(레퍼런스와 일치하는 얼굴, 헤어스타일, 흙 묻은 회색 셔츠 및 얼굴의 상처), 지폐가 드러난 돈 봉투, 뻗어 나온 상대방의 왼팔.",
        "hard_violations": [],
        "physics": "현우의 왼손이 돈 봉투를 덮고 있는 상대방의 왼쪽 손목을 위에서 아래로 체중을 실어 강하게 짓누르고 있으며, 동작과 지지면이 물리적으로 자연스러움."
       },
       {
        "label": "A",
        "direction": "인물들의 시선은 프레임 밖이며, 두 팔이 테이블 중앙의 한 지점을 향해 뻗어 있음.",
        "built_space": "식당 내부. 나무 테이블과 의자, 배경의 주방 입구 등 위치 레퍼런스의 공간 형태가 반영됨.",
        "entities": "두 사람의 팔, 테이블 위에 놓인 두꺼운 종이 패키지 형태의 봉투(지폐 여부 불분명). 손만 등장하여 신원 확인 불가.",
        "hard_violations": [
         "physically impossible anatomy: 맞잡은 두 손의 손가락들이 서로 융합되고 관절 구조가 뭉개져 손의 형태가 해부학적으로 불가능함"
        ],
        "physics": "왼쪽에서 온 손이 오른쪽에서 온 손의 바닥과 손가락을 부드럽게 감싸 쥔 상태로 테이블 위에 기대어 있으며, 상대방을 다급하게 제압하는 힘이나 동작이 전혀 읽히지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "봉투로 향하는 손을 누르는 동작은 명확하지만, 읽히는 메뉴 글자가 금지 조건을 위반하고 얼굴·상체까지 포함해 손 클로즈업의 집중도를 낮춘다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "맞은편에서 나온 두 손의 압박 접촉과 탁자 가장자리를 손 중심으로 담았지만, 제압당한 손끝이 봉투 반대쪽을 향해 집어 들려던 순간은 덜 명확하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 아래쪽의 겹친 손을 바라보며 오른쪽에서 들어온 남자의 손등과 손목 경계를 아래로 누른다. 아래 손의 손가락은 왼쪽 돈 봉투 위로 뻗어 있어 봉투를 가져가려는 목표가 명확하다.",
        "built_space": "전경의 나무 탁자 한 개는 상판과 가까운 가장자리가 보이고, 뒤쪽에는 다른 탁자 일부와 갈색 금속 의자 등받이 네 개가 보인다. 왼쪽에는 휴지함 하나와 수저통 일부가 있다. 낡은 투톤 벽은 참고와 유사하지만 오른쪽의 큰 검은 개구부는 참고에서 확인되지 않아 정확한 장소 일치는 약하다. 현우는 탁자 건너편에서 몸을 숙이고 상대 팔은 오른쪽에서 들어온다.",
        "entities": "현우의 앳된 동아시아계 남성 외모, 헝클어진 검은 머리와 회색 셔츠는 참고에 대체로 맞으며 이마 상처도 보인다. 상대 남자는 손과 팔만 보이고 추가 인물은 없다. 갈색 봉투 하나와 안쪽 지폐 묶음이 보이지만 지폐의 국가·액면은 식별되지 않는다. 벽 메뉴의 '계란말이', '공기밥'과 가격 숫자는 읽힌다.",
        "hard_violations": [
         "벽 메뉴에 읽을 수 있는 한글과 가격 숫자가 남아 있어 가독성 있는 문자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "상대 손과 팔은 탁자에 닿아 지지되고, 현우의 손은 상대 손등에 실제로 접촉한다. 뻗은 팔에서 손바닥으로 이어지는 하향 압박은 가능한 동작이다. 봉투와 지폐는 상판에 놓여 있으며 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "왼쪽에서 나온 현우의 손은 오른쪽 남자의 손등을 아래로 누르고 감싼다. 다만 아래 손의 손끝은 화면 아래 왼쪽으로 향하고 봉투는 오른쪽에 있어, 손이 봉투를 집으러 뻗던 방향은 뚜렷하지 않다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "전경 탁자 한 개의 상판과 가까운 가장자리가 선명하고, 뒤에는 벽 쪽 탁자 하나와 갈색 의자 등받이 하나가 보인다. 왼쪽 냉장고, 안쪽 미닫이문, 중앙 기둥과 달력, 병 상자 두 단, 낡은 투톤 벽이 참고 공간과 상당히 일치한다. 뒤 탁자에는 수저통 하나와 양념 용기들이 보인다. 벽의 큰 풍경 액자는 참고의 메뉴판과 다르다. 양쪽 팔이 탁자 양편에서 들어오는 배치는 자연스럽지만, 참고의 구석 탁자 자체인지까지는 확정하기 어렵다.",
        "entities": "두 사람의 손과 팔만 보여 손 중심 프레이밍을 지킨다. 현우로 읽히는 왼쪽 팔은 마른 체형과 해진 회색 셔츠가 참고에 부합한다. 얼굴이 없어 정확한 나이와 한국계 미국인 정체성은 확인할 수 없다. 오른쪽 상대도 회색 계열 긴소매를 입었다. 갈색 봉투 하나가 손 옆에 놓여 있지만 닫혀 있어 현금 내용물은 보이지 않는다. 배경 글자는 흐려 읽히지 않으며 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "아래 손의 손가락과 오른쪽 팔은 상판에 닿아 있고, 위 손의 손바닥과 굽힌 손가락은 아래 손등을 누른다. 팔의 연결과 접촉은 물리적으로 가능하며, 봉투도 상판에 놓여 있다. 다만 손등을 급히 꽉 내리누르는 힘보다는 손목 가까이를 붙잡는 느낌이 조금 강하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "봉투로 향하는 손을 누르는 동작은 명확하지만, 읽히는 메뉴 글자가 금지 조건을 위반하고 얼굴·상체까지 포함해 손 클로즈업의 집중도를 낮춘다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "맞은편에서 나온 두 손의 압박 접촉과 탁자 가장자리를 손 중심으로 담았지만, 제압당한 손끝이 봉투 반대쪽을 향해 집어 들려던 순간은 덜 명확하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 아래쪽의 겹친 손을 바라보며 오른쪽에서 들어온 남자의 손등과 손목 경계를 아래로 누른다. 아래 손의 손가락은 왼쪽 돈 봉투 위로 뻗어 있어 봉투를 가져가려는 목표가 명확하다.",
        "built_space": "전경의 나무 탁자 한 개는 상판과 가까운 가장자리가 보이고, 뒤쪽에는 다른 탁자 일부와 갈색 금속 의자 등받이 네 개가 보인다. 왼쪽에는 휴지함 하나와 수저통 일부가 있다. 낡은 투톤 벽은 참고와 유사하지만 오른쪽의 큰 검은 개구부는 참고에서 확인되지 않아 정확한 장소 일치는 약하다. 현우는 탁자 건너편에서 몸을 숙이고 상대 팔은 오른쪽에서 들어온다.",
        "entities": "현우의 앳된 동아시아계 남성 외모, 헝클어진 검은 머리와 회색 셔츠는 참고에 대체로 맞으며 이마 상처도 보인다. 상대 남자는 손과 팔만 보이고 추가 인물은 없다. 갈색 봉투 하나와 안쪽 지폐 묶음이 보이지만 지폐의 국가·액면은 식별되지 않는다. 벽 메뉴의 '계란말이', '공기밥'과 가격 숫자는 읽힌다.",
        "hard_violations": [
         "벽 메뉴에 읽을 수 있는 한글과 가격 숫자가 남아 있어 가독성 있는 문자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "상대 손과 팔은 탁자에 닿아 지지되고, 현우의 손은 상대 손등에 실제로 접촉한다. 뻗은 팔에서 손바닥으로 이어지는 하향 압박은 가능한 동작이다. 봉투와 지폐는 상판에 놓여 있으며 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "왼쪽에서 나온 현우의 손은 오른쪽 남자의 손등을 아래로 누르고 감싼다. 다만 아래 손의 손끝은 화면 아래 왼쪽으로 향하고 봉투는 오른쪽에 있어, 손이 봉투를 집으러 뻗던 방향은 뚜렷하지 않다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "전경 탁자 한 개의 상판과 가까운 가장자리가 선명하고, 뒤에는 벽 쪽 탁자 하나와 갈색 의자 등받이 하나가 보인다. 왼쪽 냉장고, 안쪽 미닫이문, 중앙 기둥과 달력, 병 상자 두 단, 낡은 투톤 벽이 참고 공간과 상당히 일치한다. 뒤 탁자에는 수저통 하나와 양념 용기들이 보인다. 벽의 큰 풍경 액자는 참고의 메뉴판과 다르다. 양쪽 팔이 탁자 양편에서 들어오는 배치는 자연스럽지만, 참고의 구석 탁자 자체인지까지는 확정하기 어렵다.",
        "entities": "두 사람의 손과 팔만 보여 손 중심 프레이밍을 지킨다. 현우로 읽히는 왼쪽 팔은 마른 체형과 해진 회색 셔츠가 참고에 부합한다. 얼굴이 없어 정확한 나이와 한국계 미국인 정체성은 확인할 수 없다. 오른쪽 상대도 회색 계열 긴소매를 입었다. 갈색 봉투 하나가 손 옆에 놓여 있지만 닫혀 있어 현금 내용물은 보이지 않는다. 배경 글자는 흐려 읽히지 않으며 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "아래 손의 손가락과 오른쪽 팔은 상판에 닿아 있고, 위 손의 손바닥과 굽힌 손가락은 아래 손등을 누른다. 팔의 연결과 접촉은 물리적으로 가능하며, 봉투도 상판에 놓여 있다. 다만 손등을 급히 꽉 내리누르는 힘보다는 손목 가까이를 붙잡는 느낌이 조금 강하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.179
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible anatomy: 맞잡은 두 손의 손가락들이 서로 융합되고 관절 구조가 뭉개져 손의 형태가 해부학적으로 불가능함"
    ],
    "B": [
     "[gpt-high] 벽 메뉴에 읽을 수 있는 한글과 가격 숫자가 남아 있어 가독성 있는 문자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1179,
   "A": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "요구된 손 클로즈업 숏보다 화각이 넓어져 인물의 상반신과 얼굴까지 노출되었으나, 돈 봉투 위로 뻗은 손등을 강하게 내리누르는 핵심 액션과 위치 레퍼런스를 정확하게 구현하였습니다.  ★위반: [gpt-high] 벽 메뉴에 읽을 수 있는 한글과 가격 숫자가 남아 있어 가독성 있는 문자를 전면 금지한 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "손 클로즈업이라는 프레임 스케일은 맞추었으나, 두 손의 구조가 해부학적으로 붕괴되었고 봉투를 제압하는 대신 손을 쥐고 있는 듯한 엉뚱한 액션을 연출하여 탈락입니다.  ★위반: [gemini-pro] physically impossible anatomy: 맞잡은 두 손의 손가락들이 서로 융합되고 관절 구조가 뭉개져 손의 형태가 해부학적으로 불가능함"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L162B01.png",
    "asset_id": "7a16f2f4-24b6-4a89-a184-9204ad75cbe4",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-b1b0-7920-aa2a-fc10eac3d5b5",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3__bgfirst_bg.png",
   "bg_asset_id": "33a03a51-7f51-4fab-9373-3f1601df1a45",
   "bg_record_key": "S10sh3::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S10sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:15:45.039836+00:00",
  "fingerprint": "496d5938efb779088b1c0dc7a78164a309d3965f90de09801fa9b619f31832c9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S10sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S10sh3_sel.png",
  "source_sha256": "a186e38714c08cf0a6504e762eb11dc169ed37a02aec572b24a63b72aed2841a",
  "file": "S10sh3_cine.png",
  "staged_sha256": "ba52780a5b0eb7d54ab371301fc329bc5c1e65f18a1185243f42a20f0457c078",
  "latency_ms": 10645
 },
 "S10sh6::signage": {
  "fp": "7794cd3d370fa59e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S10sh6": {
  "input_fingerprint": "6cdd8509bc71fa24",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 사이렌 소리가 울리는 가운데 현우를 향해 조롱하듯 여유롭게 웃고 있는 남자의 얼굴.\n\nLOCATION (lock): At the secluded dining table inside the shabby alley restaurant, with dim lighting during the nighttime conversation. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Secluded restaurant table (Remains between the two seated men) — Only a narrow edge is retained below the face and foreground shoulder; used as Maintains continuity with the interrupted transaction while allowing the mocking expression to lead.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the restaurant's ambient illumination and restrained contrast, with no flashing or colored light inferred from the siren.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The nighttime meeting remains at the table in the shabby restaurant's secluded area. 현우: Hyunwoo remains seated, tense with anger; his injured leg and earlier head injury persist.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 사이렌 소리가 울리는 가운데 현우를 향해 조롱하듯 여유롭게 웃고 있는 남자의 얼굴.\n\nLOCATION (lock): At the secluded dining table inside the shabby alley restaurant, with dim lighting during the nighttime conversation. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Secluded restaurant table (Remains between the two seated men) — Only a narrow edge is retained below the face and foreground shoulder; used as Maintains continuity with the interrupted transaction while allowing the mocking expression to lead.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the restaurant's ambient illumination and restrained contrast, with no flashing or colored light inferred from the siren.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The nighttime meeting remains at the table in the shabby restaurant's secluded area. 현우: Hyunwoo remains seated, tense with anger; his injured leg and earlier head injury persist.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 사이렌 소리가 울리는 가운데 현우를 향해 조롱하듯 여유롭게 웃고 있는 남자의 얼굴.\n\nLOCATION (lock): At the secluded dining table inside the shabby alley restaurant, with dim lighting during the nighttime conversation. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Secluded restaurant table (Remains between the two seated men) — Only a narrow edge is retained below the face and foreground shoulder; used as Maintains continuity with the interrupted transaction while allowing the mocking expression to lead.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the restaurant's ambient illumination and restrained contrast, with no flashing or colored light inferred from the siren.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The nighttime meeting remains at the table in the shabby restaurant's secluded area. 현우: Hyunwoo remains seated, tense with anger; his injured leg and earlier head injury persist.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "남자의 시선이 프레임 밖 아래쪽(테이블 건너편 현우의 위치)을 향하고 있음.",
    "built_space": "어두운 식당 내부 벽면과 포스터, 화면 하단에 테이블 모서리가 좁게 배치됨.",
    "entities": "올리브색 셔츠를 입은 남자(이전 샷과 일치)가 조롱하는 미소를 짓고 있으며, 지시대로 사이렌 불빛이 배제됨.",
    "hard_violations": [],
    "physics": "어깨와 목의 각도를 통해 테이블 앞에 자연스럽게 앉아 있는 자세가 확인됨."
   },
   {
    "label": "B",
    "direction": "남자의 시선이 전경에 걸쳐 있는 인물(현우)의 얼굴 측면을 향함.",
    "built_space": "식당 내부가 깊게 보이며, 왼쪽 벽의 메뉴판과 주방 출입구 구조가 드러남.",
    "entities": "올리브색 셔츠의 남자와 전경의 회색 셔츠를 입은 인물. 배경 출입구 쪽에 명백한 푸른색 사이렌 불빛이 존재함.",
    "hard_violations": [
     "[gpt-high] 왼쪽 벽 메뉴판에 읽을 수 있는 한글이 노출되어 글자 가독성을 전면 금지한 조건을 위반한다."
    ],
    "physics": "테이블 위에 손과 팔뚝을 올리고 안정적으로 앉아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 프레이밍을 잘 유지했고 금지된 사이렌 불빛을 배제했으나, 벽면에 일부 읽을 수 있는 텍스트가 남은 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "클로즈업 지시를 무시하고 샷을 넓게 잡았으며, 지문에서 명시적으로 금지한 푸른색 사이렌 불빛과 다량의 읽을 수 있는 메뉴판 텍스트가 포함되어 우선순위 지침을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자의 시선이 프레임 밖 아래쪽(테이블 건너편 현우의 위치)을 향하고 있음.",
        "built_space": "어두운 식당 내부 벽면과 포스터, 화면 하단에 테이블 모서리가 좁게 배치됨.",
        "entities": "올리브색 셔츠를 입은 남자(이전 샷과 일치)가 조롱하는 미소를 짓고 있으며, 지시대로 사이렌 불빛이 배제됨.",
        "hard_violations": [],
        "physics": "어깨와 목의 각도를 통해 테이블 앞에 자연스럽게 앉아 있는 자세가 확인됨."
       },
       {
        "label": "B",
        "direction": "남자의 시선이 전경에 걸쳐 있는 인물(현우)의 얼굴 측면을 향함.",
        "built_space": "식당 내부가 깊게 보이며, 왼쪽 벽의 메뉴판과 주방 출입구 구조가 드러남.",
        "entities": "올리브색 셔츠의 남자와 전경의 회색 셔츠를 입은 인물. 배경 출입구 쪽에 명백한 푸른색 사이렌 불빛이 존재함.",
        "hard_violations": [],
        "physics": "테이블 위에 손과 팔뚝을 올리고 안정적으로 앉아 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 프레이밍을 잘 유지했고 금지된 사이렌 불빛을 배제했으나, 벽면에 일부 읽을 수 있는 텍스트가 남은 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "클로즈업 지시를 무시하고 샷을 넓게 잡았으며, 지문에서 명시적으로 금지한 푸른색 사이렌 불빛과 다량의 읽을 수 있는 메뉴판 텍스트가 포함되어 우선순위 지침을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "남자의 시선이 프레임 밖 아래쪽(테이블 건너편 현우의 위치)을 향하고 있음.",
        "built_space": "어두운 식당 내부 벽면과 포스터, 화면 하단에 테이블 모서리가 좁게 배치됨.",
        "entities": "올리브색 셔츠를 입은 남자(이전 샷과 일치)가 조롱하는 미소를 짓고 있으며, 지시대로 사이렌 불빛이 배제됨.",
        "hard_violations": [],
        "physics": "어깨와 목의 각도를 통해 테이블 앞에 자연스럽게 앉아 있는 자세가 확인됨."
       },
       {
        "label": "B",
        "direction": "남자의 시선이 전경에 걸쳐 있는 인물(현우)의 얼굴 측면을 향함.",
        "built_space": "식당 내부가 깊게 보이며, 왼쪽 벽의 메뉴판과 주방 출입구 구조가 드러남.",
        "entities": "올리브색 셔츠의 남자와 전경의 회색 셔츠를 입은 인물. 배경 출입구 쪽에 명백한 푸른색 사이렌 불빛이 존재함.",
        "hard_violations": [],
        "physics": "테이블 위에 손과 팔뚝을 올리고 안정적으로 앉아 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우를 향한 여유로운 웃음은 맞지만, 얼굴 클로즈업 대신 상반신과 식당을 넓게 보여주며 읽히는 메뉴 글자와 푸른 배경광도 지시를 어긴다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "조롱하는 남자의 얼굴 클로즈업과 절제된 식당 조명이 가장 충실하지만, 지정된 전경 어깨와 두 사람 사이의 좁은 탁자 가장자리는 빠져 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자는 화면 오른쪽 전경에 앉은 현우의 얼굴 쪽을 바라보며 입꼬리를 올리고 있다. 현우도 남자 쪽으로 머리를 돌리고 있으나 눈은 보이지 않는다. 무기나 방향을 확인할 휴대 도구는 없다.",
        "built_space": "두 사람이 앞쪽 식탁을 사이에 두고 마주 앉아 있다. 주 식탁 외에 왼쪽 뒤로 식탁 두 개와 의자 등받이 세 개가 구분된다. 뒤에는 주방 통로, 반투명 문, 냉장고 한 대, 선반과 병 상자가 있고 천장 조명 두 개가 보인다. 낡은 투톤 벽과 목재 식탁은 참고 장소와 유사하지만, 얼굴보다 실내와 상반신을 크게 노출해 지정된 클로즈업에서 벗어난다. 반사는 문제되지 않는다.",
        "entities": "웃는 인물은 검은 짧은 머리의 중년 동아시아계 남성으로, 참고 장면의 상대 남성에게 어울리는 낡은 카키색 셔츠를 입었다. 오른쪽 현우는 검은 머리와 회색 계열 옷을 입은 뒷모습만 보여 얼굴·나이·부상 상태를 확인할 수 없다. 식탁 위 갈색 봉투는 이전 장면과 이어진다. 왼쪽 벽 메뉴에는 읽을 수 있는 한글이 남아 있고, 뒤쪽 문에는 지시와 달리 강한 푸른빛이 나타난다.",
        "hard_violations": [
         "왼쪽 벽 메뉴판에 읽을 수 있는 한글이 노출되어 글자 가독성을 전면 금지한 조건을 위반한다."
        ],
        "physics": "남자의 몸은 의자에 앉은 자세이며, 앞쪽 팔과 손은 식탁에 닿아 지지된다. 봉투도 식탁 위에 놓여 있다. 현우의 하체와 좌판 접촉은 전경 크롭으로 가려져 있지만 상체 배치에 부유나 불가능한 자세는 없다."
       },
       {
        "label": "B",
        "direction": "남자의 눈과 얼굴은 렌즈 정면보다 화면 왼쪽 가까운 곳을 향한다. 화면 밖 맞은편 현우를 보고 비웃는 시선으로 자연스럽게 읽히지만, 현우가 보이지 않아 실제 눈맞춤 위치를 직접 확인할 수는 없다. 비대칭으로 올라간 입꼬리와 좁아진 눈이 여유로운 조롱을 전달한다.",
        "built_space": "얼굴과 어깨가 화면 대부분을 차지한다. 왼쪽 뒤에 목재 식탁 하나, 금속 수저통 하나, 양념 용기군과 휴지함 하나가 보이며, 낡은 투톤 벽에는 풍경 달력 하나와 흐릿한 벽 게시물이 있다. 참고 식당의 재질과 비품 구성이 이어지고 불가능한 반사나 중복된 고정 설비는 보이지 않는다. 다만 왼쪽 아래 탁자는 배경 식탁이며, 요청된 두 사람 사이 식탁의 좁은 가장자리와 현우의 전경 어깨는 없다.",
        "entities": "주인공은 짧은 검은 머리와 수염 자국이 있는 중년 동아시아계 남성이고, 낡은 회갈색·카키색 셔츠가 이전 장면의 상대 남성 복장과 부합한다. 눈과 피부는 정상적인 인간 해부와 실물 질감으로 표현됐다. 현우, 그의 부상과 봉투는 클로즈업 밖이므로 확인할 수 없다. 배경 게시물은 흐려 명확히 읽히지 않으며 색광이나 그래픽 덧씌움도 없다.",
        "hard_violations": [],
        "physics": "남자는 상체를 약간 앞으로 기울인 자세이고 머리는 목과 어깨에 자연스럽게 연결된다. 좌판·손·발은 클로즈업 밖이라 접촉점을 확인할 수 없지만, 공중에 떠 있는 신체나 물체는 없다. 배경 수저통과 양념 용기, 휴지함은 식탁 위에 놓여 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우를 향한 여유로운 웃음은 맞지만, 얼굴 클로즈업 대신 상반신과 식당을 넓게 보여주며 읽히는 메뉴 글자와 푸른 배경광도 지시를 어긴다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "조롱하는 남자의 얼굴 클로즈업과 절제된 식당 조명이 가장 충실하지만, 지정된 전경 어깨와 두 사람 사이의 좁은 탁자 가장자리는 빠져 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "남자는 화면 오른쪽 전경에 앉은 현우의 얼굴 쪽을 바라보며 입꼬리를 올리고 있다. 현우도 남자 쪽으로 머리를 돌리고 있으나 눈은 보이지 않는다. 무기나 방향을 확인할 휴대 도구는 없다.",
        "built_space": "두 사람이 앞쪽 식탁을 사이에 두고 마주 앉아 있다. 주 식탁 외에 왼쪽 뒤로 식탁 두 개와 의자 등받이 세 개가 구분된다. 뒤에는 주방 통로, 반투명 문, 냉장고 한 대, 선반과 병 상자가 있고 천장 조명 두 개가 보인다. 낡은 투톤 벽과 목재 식탁은 참고 장소와 유사하지만, 얼굴보다 실내와 상반신을 크게 노출해 지정된 클로즈업에서 벗어난다. 반사는 문제되지 않는다.",
        "entities": "웃는 인물은 검은 짧은 머리의 중년 동아시아계 남성으로, 참고 장면의 상대 남성에게 어울리는 낡은 카키색 셔츠를 입었다. 오른쪽 현우는 검은 머리와 회색 계열 옷을 입은 뒷모습만 보여 얼굴·나이·부상 상태를 확인할 수 없다. 식탁 위 갈색 봉투는 이전 장면과 이어진다. 왼쪽 벽 메뉴에는 읽을 수 있는 한글이 남아 있고, 뒤쪽 문에는 지시와 달리 강한 푸른빛이 나타난다.",
        "hard_violations": [
         "왼쪽 벽 메뉴판에 읽을 수 있는 한글이 노출되어 글자 가독성을 전면 금지한 조건을 위반한다."
        ],
        "physics": "남자의 몸은 의자에 앉은 자세이며, 앞쪽 팔과 손은 식탁에 닿아 지지된다. 봉투도 식탁 위에 놓여 있다. 현우의 하체와 좌판 접촉은 전경 크롭으로 가려져 있지만 상체 배치에 부유나 불가능한 자세는 없다."
       },
       {
        "label": "A",
        "direction": "남자의 눈과 얼굴은 렌즈 정면보다 화면 왼쪽 가까운 곳을 향한다. 화면 밖 맞은편 현우를 보고 비웃는 시선으로 자연스럽게 읽히지만, 현우가 보이지 않아 실제 눈맞춤 위치를 직접 확인할 수는 없다. 비대칭으로 올라간 입꼬리와 좁아진 눈이 여유로운 조롱을 전달한다.",
        "built_space": "얼굴과 어깨가 화면 대부분을 차지한다. 왼쪽 뒤에 목재 식탁 하나, 금속 수저통 하나, 양념 용기군과 휴지함 하나가 보이며, 낡은 투톤 벽에는 풍경 달력 하나와 흐릿한 벽 게시물이 있다. 참고 식당의 재질과 비품 구성이 이어지고 불가능한 반사나 중복된 고정 설비는 보이지 않는다. 다만 왼쪽 아래 탁자는 배경 식탁이며, 요청된 두 사람 사이 식탁의 좁은 가장자리와 현우의 전경 어깨는 없다.",
        "entities": "주인공은 짧은 검은 머리와 수염 자국이 있는 중년 동아시아계 남성이고, 낡은 회갈색·카키색 셔츠가 이전 장면의 상대 남성 복장과 부합한다. 눈과 피부는 정상적인 인간 해부와 실물 질감으로 표현됐다. 현우, 그의 부상과 봉투는 클로즈업 밖이므로 확인할 수 없다. 배경 게시물은 흐려 명확히 읽히지 않으며 색광이나 그래픽 덧씌움도 없다.",
        "hard_violations": [],
        "physics": "남자는 상체를 약간 앞으로 기울인 자세이고 머리는 목과 어깨에 자연스럽게 연결된다. 좌판·손·발은 클로즈업 밖이라 접촉점을 확인할 수 없지만, 공중에 떠 있는 신체나 물체는 없다. 배경 수저통과 양념 용기, 휴지함은 식탁 위에 놓여 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.804
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.554
   },
   "violations": {
    "B": [
     "[gpt-high] 왼쪽 벽 메뉴판에 읽을 수 있는 한글이 노출되어 글자 가독성을 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 554
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 클로즈업 프레이밍을 잘 유지했고 금지된 사이렌 불빛을 배제했으나, 벽면에 일부 읽을 수 있는 텍스트가 남은 점이 감점 요인입니다."
   },
   {
    "label": "B",
    "score": 554,
    "verdict_ko": "클로즈업 지시를 무시하고 샷을 넓게 잡았으며, 지문에서 명시적으로 금지한 푸른색 사이렌 불빛과 다량의 읽을 수 있는 메뉴판 텍스트가 포함되어 우선순위 지침을 크게 위반했습니다.  ★위반: [gpt-high] 왼쪽 벽 메뉴판에 읽을 수 있는 한글이 노출되어 글자 가독성을 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3_sel.png",
    "asset_id": "0f19705b-748d-46e3-8c25-56c8e55e2d71",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-b4e4-73aa-a39e-611b55f48cc3",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S10sh3"
  }
 },
 "S10sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:16:47.926134+00:00",
  "fingerprint": "6e1015afb8608c34b138e5dd011d28f78ef35cc08f92c885f89a988d9df64465",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S10sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S10sh6_sel.png",
  "source_sha256": "7e72343c1f6e6933ea3556e88c337b1ffa66aae580ac603b55bb9ec11f0081d2",
  "file": "S10sh6_cine.png",
  "staged_sha256": "b5d181fd297956fe42dea591b205ae3b5f863c59403ac7ce0503ec9c5b6fbf98",
  "latency_ms": 10592
 },
 "S10sh7::signage": {
  "fp": "b5111a583a2f90ae",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S10sh7": {
  "input_fingerprint": "05ce99d6f07ca24d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 희미한 식당 조명 아래, 두 주먹을 꽉 쥔 채 무력하게 고개를 푹 숙이고 얼어붙은 현우의 전신.\n\nLOCATION (lock): In the secluded seating area of a shabby alley restaurant, under its faint interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Restaurant table (Beside 현우 at the secluded seating position) — Its side and part of its top are visible to the left of his body; used as Provides spatial context without concealing his fists or feet; 현우's seat (Occupied) — Seen obliquely from its open side; used as Supports the full-body silhouette and makes his collapsed seated posture legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Faint restaurant illumination preserves subdued facial and hand detail within restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The meeting remains in the secluded part of the shabby restaurant at night. 현우: Hyunwoo remains at the table, visibly frustrated and powerless, with his leg wound and earlier head injury unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 희미한 식당 조명 아래, 두 주먹을 꽉 쥔 채 무력하게 고개를 푹 숙이고 얼어붙은 현우의 전신.\n\nLOCATION (lock): In the secluded seating area of a shabby alley restaurant, under its faint interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Restaurant table (Beside 현우 at the secluded seating position) — Its side and part of its top are visible to the left of his body; used as Provides spatial context without concealing his fists or feet; 현우's seat (Occupied) — Seen obliquely from its open side; used as Supports the full-body silhouette and makes his collapsed seated posture legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Faint restaurant illumination preserves subdued facial and hand detail within restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The meeting remains in the secluded part of the shabby restaurant at night. 현우: Hyunwoo remains at the table, visibly frustrated and powerless, with his leg wound and earlier head injury unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 희미한 식당 조명 아래, 두 주먹을 꽉 쥔 채 무력하게 고개를 푹 숙이고 얼어붙은 현우의 전신.\n\nLOCATION (lock): In the secluded seating area of a shabby alley restaurant, under its faint interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Restaurant table (Beside 현우 at the secluded seating position) — Its side and part of its top are visible to the left of his body; used as Provides spatial context without concealing his fists or feet; 현우's seat (Occupied) — Seen obliquely from its open side; used as Supports the full-body silhouette and makes his collapsed seated posture legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Faint restaurant illumination preserves subdued facial and hand detail within restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The meeting remains in the secluded part of the shabby restaurant at night. 현우: Hyunwoo remains at the table, visibly frustrated and powerless, with his leg wound and earlier head injury unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "고개는 무릎을 향해 푹 숙이고 있으며 두 주먹은 양 허벅지 위에 놓여 있음.",
    "built_space": "식당 내부의 테이블이 왼쪽에 배치됨. 인물이 앉은 의자는 우측에 여분의 방석과 다리가 융합된 불가능한 구조임.",
    "entities": "현우의 인상착의와 얼굴 상처가 참조와 일치하며 왼쪽 무릎에 붕대가 있음.",
    "hard_violations": [
     "[gemini-pro] 기형적으로 융합된 물리적으로 불가능한 의자 구조",
     "[gemini-pro] 명시적 금지 사항을 위반하여 선명하게 읽을 수 있는 텍스트 노출",
     "[gpt-high] 배경 메뉴판의 한글과 ‘4,000’, ‘3,500’, ‘2,500’, ‘1,500’ 등의 가격이 선명하게 읽혀 글자 금지 조건을 위반한다."
    ],
    "physics": "발은 바닥에 닿아 있으나 지지대인 의자가 물리적으로 불가능하게 융합된 형태임."
   },
   {
    "label": "B",
    "direction": "고개를 아래로 깊게 숙이고 있음. 왼쪽 팔은 늘어뜨려 주먹을 쥐고 오른손은 허벅지 위에 둠.",
    "built_space": "낡은 식당 내부로 테이블 측면이 왼쪽에 보임. 벽면 텍스트는 얕은 심도로 읽을 수 없게 처리됨.",
    "entities": "현우의 외모와 의상 및 얼굴 상처가 참조 이미지와 일치함.",
    "hard_violations": [
     "[gpt-high] 오른쪽 위 벽 메뉴판에 읽을 수 있는 가격 숫자가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "나무 의자에 안정적으로 앉아 있으며 두 발이 바닥을 딛고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "의자가 기형적으로 융합된 물리적 오류와 금지된 텍스트가 선명하게 노출되는 치명적 위반이 있습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 프레이밍과 조명 분위기를 잘 구현했으며 얕은 심도를 활용해 텍스트 노출 금지 지시를 훌륭하게 따랐습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "고개는 무릎을 향해 푹 숙이고 있으며 두 주먹은 양 허벅지 위에 놓여 있음.",
        "built_space": "식당 내부의 테이블이 왼쪽에 배치됨. 인물이 앉은 의자는 우측에 여분의 방석과 다리가 융합된 불가능한 구조임.",
        "entities": "현우의 인상착의와 얼굴 상처가 참조와 일치하며 왼쪽 무릎에 붕대가 있음.",
        "hard_violations": [
         "기형적으로 융합된 물리적으로 불가능한 의자 구조",
         "명시적 금지 사항을 위반하여 선명하게 읽을 수 있는 텍스트 노출"
        ],
        "physics": "발은 바닥에 닿아 있으나 지지대인 의자가 물리적으로 불가능하게 융합된 형태임."
       },
       {
        "label": "B",
        "direction": "고개를 아래로 깊게 숙이고 있음. 왼쪽 팔은 늘어뜨려 주먹을 쥐고 오른손은 허벅지 위에 둠.",
        "built_space": "낡은 식당 내부로 테이블 측면이 왼쪽에 보임. 벽면 텍스트는 얕은 심도로 읽을 수 없게 처리됨.",
        "entities": "현우의 외모와 의상 및 얼굴 상처가 참조 이미지와 일치함.",
        "hard_violations": [],
        "physics": "나무 의자에 안정적으로 앉아 있으며 두 발이 바닥을 딛고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "의자가 기형적으로 융합된 물리적 오류와 금지된 텍스트가 선명하게 노출되는 치명적 위반이 있습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 프레이밍과 조명 분위기를 잘 구현했으며 얕은 심도를 활용해 텍스트 노출 금지 지시를 훌륭하게 따랐습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "고개는 무릎을 향해 푹 숙이고 있으며 두 주먹은 양 허벅지 위에 놓여 있음.",
        "built_space": "식당 내부의 테이블이 왼쪽에 배치됨. 인물이 앉은 의자는 우측에 여분의 방석과 다리가 융합된 불가능한 구조임.",
        "entities": "현우의 인상착의와 얼굴 상처가 참조와 일치하며 왼쪽 무릎에 붕대가 있음.",
        "hard_violations": [
         "기형적으로 융합된 물리적으로 불가능한 의자 구조",
         "명시적 금지 사항을 위반하여 선명하게 읽을 수 있는 텍스트 노출"
        ],
        "physics": "발은 바닥에 닿아 있으나 지지대인 의자가 물리적으로 불가능하게 융합된 형태임."
       },
       {
        "label": "B",
        "direction": "고개를 아래로 깊게 숙이고 있음. 왼쪽 팔은 늘어뜨려 주먹을 쥐고 오른손은 허벅지 위에 둠.",
        "built_space": "낡은 식당 내부로 테이블 측면이 왼쪽에 보임. 벽면 텍스트는 얕은 심도로 읽을 수 없게 처리됨.",
        "entities": "현우의 외모와 의상 및 얼굴 상처가 참조 이미지와 일치함.",
        "hard_violations": [],
        "physics": "나무 의자에 안정적으로 앉아 있으며 두 발이 바닥을 딛고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "고개를 숙인 착석은 맞지만 발이 잘리고 한쪽 주먹이 가려져 전신·두 주먹 구도를 놓쳤으며, 벽 메뉴판의 가격 숫자도 남아 있다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "전신과 두 주먹, 숙인 고개, 왼쪽 식탁 및 부상을 가장 충실히 보여주지만, 선명하게 읽히는 메뉴와 가격 때문에 최종 사용 조건은 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽을 향해 앉아 고개를 아래로 숙이고 있으며, 얼굴은 무릎 앞 바닥 쪽을 향한다. 보이는 손은 의자 옆에서 아래로 내려가 주먹을 쥐고 있다. 반대쪽 손은 몸에 가려져 두 주먹의 상태를 확인할 수 없다.",
        "built_space": "왼쪽 전경에 식탁 한 개의 상판과 옆면이 보이고 뒤쪽에도 식탁이 있다. 현우는 오른쪽 벽 앞 목제 등받이 의자 한 개에 앉아 있으며 카메라는 의자의 열린 측면을 본다. 천장등 한 개, 뒤쪽 출입구와 흰색 수납·가전 형태의 물체, 벽 메뉴판들이 보인다. 낡은 투톤 벽과 목재 식탁은 참고 장소의 재질과 유사하지만, 참고에 없던 깊은 통로가 공간에서 크게 드러난다. 식탁이 가까운 주먹을 가리지는 않지만 화면 하단이 신발을 잘라 전신 조건을 충족하지 못한다.",
        "entities": "젊은 동아시아계 남성 한 명이며 검은 머리, 회색 셔츠, 갈색 계열 카고바지와 벨트는 현우 참고와 대체로 맞는다. 숙인 옆얼굴 때문에 정확한 얼굴 일치는 제한적으로만 확인된다. 머리와 다리의 기존 부상은 뚜렷하게 확인되지 않는다. 식탁 위에는 금속 수저통과 목제 상자가 보인다. 오른쪽 위 메뉴판에는 가격 숫자가 식별 가능한 상태로 남아 있다.",
        "hard_violations": [
         "오른쪽 위 벽 메뉴판에 읽을 수 있는 가격 숫자가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "엉덩이는 의자 좌판에 놓이고 등 뒤에 등받이가 있어 착석 지지는 자연스럽다. 상체를 앞으로 접고 목을 숙인 자세도 가능하다. 보이는 주먹은 팔에 정상적으로 연결되어 있다. 발은 화면 경계에 잘려 바닥 접촉 전체를 확인할 수 없지만, 몸이 공중에 떠 있다는 증거는 없다."
       },
       {
        "label": "B",
        "direction": "현우는 상체와 고개를 앞으로 깊게 숙여 두 손 사이와 바닥 쪽을 향한다. 양쪽 주먹은 무릎 위쪽 앞 공간에서 서로 마주보는 방향으로 꽉 쥐어져 있고, 카메라를 바라보지 않는다. 무력하게 고개를 숙인 순간이 명확하다.",
        "built_space": "왼쪽 식탁 한 개의 상판 일부와 옆면이 보이며 두 주먹과 두 발을 가리지 않는다. 현우는 화면 오른쪽의 금속 프레임 등받이 의자에 앉아 있고, 왼쪽 뒤에는 빈 의자 좌판과 다리 일부가 보인다. 참고와 유사한 얼룩진 투톤 벽, 세로 메뉴판들, 병 그림 포스터, 식탁 위 수저통과 휴지통이 배치되어 있다. 의자의 열린 앞·측면에서 머리부터 신발까지 온전히 담아 요청한 전신 구도를 구현한다.",
        "entities": "검은 머리와 앳된 얼굴의 동아시아계 남성 한 명만 보인다. 회색 셔츠, 벨트, 카고바지, 낡은 갈색 신발은 현우 참고와 대체로 일치한다. 관자놀이 부근 상처와 피가 묻은 무릎 부위 붕대로 머리·다리 부상이 표현된다. 다만 이전 참고에는 현우의 부상 세부가 없어 정확히 같은 상처인지는 확인할 수 없다. 식탁과 점유된 좌석, 두 주먹이 모두 식별된다. 벽 메뉴판의 한글과 가격은 명백하게 읽힌다.",
        "hard_violations": [
         "배경 메뉴판의 한글과 ‘4,000’, ‘3,500’, ‘2,500’, ‘1,500’ 등의 가격이 선명하게 읽혀 글자 금지 조건을 위반한다."
        ],
        "physics": "엉덩이는 의자 좌판에 지지되고 두 신발은 바닥에 닿아 있다. 등받이는 등 뒤에 놓여 의자 사용 방향도 맞는다. 앞으로 기울인 몸은 좌판과 벌린 두 발로 지지되며, 주먹은 굽힌 팔로 유지된다. 부유하거나 지지 없이 놓인 신체·물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "고개를 숙인 착석은 맞지만 발이 잘리고 한쪽 주먹이 가려져 전신·두 주먹 구도를 놓쳤으며, 벽 메뉴판의 가격 숫자도 남아 있다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "전신과 두 주먹, 숙인 고개, 왼쪽 식탁 및 부상을 가장 충실히 보여주지만, 선명하게 읽히는 메뉴와 가격 때문에 최종 사용 조건은 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽을 향해 앉아 고개를 아래로 숙이고 있으며, 얼굴은 무릎 앞 바닥 쪽을 향한다. 보이는 손은 의자 옆에서 아래로 내려가 주먹을 쥐고 있다. 반대쪽 손은 몸에 가려져 두 주먹의 상태를 확인할 수 없다.",
        "built_space": "왼쪽 전경에 식탁 한 개의 상판과 옆면이 보이고 뒤쪽에도 식탁이 있다. 현우는 오른쪽 벽 앞 목제 등받이 의자 한 개에 앉아 있으며 카메라는 의자의 열린 측면을 본다. 천장등 한 개, 뒤쪽 출입구와 흰색 수납·가전 형태의 물체, 벽 메뉴판들이 보인다. 낡은 투톤 벽과 목재 식탁은 참고 장소의 재질과 유사하지만, 참고에 없던 깊은 통로가 공간에서 크게 드러난다. 식탁이 가까운 주먹을 가리지는 않지만 화면 하단이 신발을 잘라 전신 조건을 충족하지 못한다.",
        "entities": "젊은 동아시아계 남성 한 명이며 검은 머리, 회색 셔츠, 갈색 계열 카고바지와 벨트는 현우 참고와 대체로 맞는다. 숙인 옆얼굴 때문에 정확한 얼굴 일치는 제한적으로만 확인된다. 머리와 다리의 기존 부상은 뚜렷하게 확인되지 않는다. 식탁 위에는 금속 수저통과 목제 상자가 보인다. 오른쪽 위 메뉴판에는 가격 숫자가 식별 가능한 상태로 남아 있다.",
        "hard_violations": [
         "오른쪽 위 벽 메뉴판에 읽을 수 있는 가격 숫자가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "엉덩이는 의자 좌판에 놓이고 등 뒤에 등받이가 있어 착석 지지는 자연스럽다. 상체를 앞으로 접고 목을 숙인 자세도 가능하다. 보이는 주먹은 팔에 정상적으로 연결되어 있다. 발은 화면 경계에 잘려 바닥 접촉 전체를 확인할 수 없지만, 몸이 공중에 떠 있다는 증거는 없다."
       },
       {
        "label": "A",
        "direction": "현우는 상체와 고개를 앞으로 깊게 숙여 두 손 사이와 바닥 쪽을 향한다. 양쪽 주먹은 무릎 위쪽 앞 공간에서 서로 마주보는 방향으로 꽉 쥐어져 있고, 카메라를 바라보지 않는다. 무력하게 고개를 숙인 순간이 명확하다.",
        "built_space": "왼쪽 식탁 한 개의 상판 일부와 옆면이 보이며 두 주먹과 두 발을 가리지 않는다. 현우는 화면 오른쪽의 금속 프레임 등받이 의자에 앉아 있고, 왼쪽 뒤에는 빈 의자 좌판과 다리 일부가 보인다. 참고와 유사한 얼룩진 투톤 벽, 세로 메뉴판들, 병 그림 포스터, 식탁 위 수저통과 휴지통이 배치되어 있다. 의자의 열린 앞·측면에서 머리부터 신발까지 온전히 담아 요청한 전신 구도를 구현한다.",
        "entities": "검은 머리와 앳된 얼굴의 동아시아계 남성 한 명만 보인다. 회색 셔츠, 벨트, 카고바지, 낡은 갈색 신발은 현우 참고와 대체로 일치한다. 관자놀이 부근 상처와 피가 묻은 무릎 부위 붕대로 머리·다리 부상이 표현된다. 다만 이전 참고에는 현우의 부상 세부가 없어 정확히 같은 상처인지는 확인할 수 없다. 식탁과 점유된 좌석, 두 주먹이 모두 식별된다. 벽 메뉴판의 한글과 가격은 명백하게 읽힌다.",
        "hard_violations": [
         "배경 메뉴판의 한글과 ‘4,000’, ‘3,500’, ‘2,500’, ‘1,500’ 등의 가격이 선명하게 읽혀 글자 금지 조건을 위반한다."
        ],
        "physics": "엉덩이는 의자 좌판에 지지되고 두 신발은 바닥에 닿아 있다. 등받이는 등 뒤에 놓여 의자 사용 방향도 맞는다. 앞으로 기울인 몸은 좌판과 벌린 두 발로 지지되며, 주먹은 굽힌 팔로 유지된다. 부유하거나 지지 없이 놓인 신체·물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 기형적으로 융합된 물리적으로 불가능한 의자 구조",
     "[gemini-pro] 명시적 금지 사항을 위반하여 선명하게 읽을 수 있는 텍스트 노출",
     "[gpt-high] 배경 메뉴판의 한글과 ‘4,000’, ‘3,500’, ‘2,500’, ‘1,500’ 등의 가격이 선명하게 읽혀 글자 금지 조건을 위반한다."
    ],
    "B": [
     "[gpt-high] 오른쪽 위 벽 메뉴판에 읽을 수 있는 가격 숫자가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "의자가 기형적으로 융합된 물리적 오류와 금지된 텍스트가 선명하게 노출되는 치명적 위반이 있습니다.  ★위반: [gemini-pro] 기형적으로 융합된 물리적으로 불가능한 의자 구조 / [gemini-pro] 명시적 금지 사항을 위반하여 선명하게 읽을 수 있는 텍스트 노출 / [gpt-high] 배경 메뉴판의 한글과 ‘4,000’, ‘3,500’, ‘2,500’, ‘1,500’ 등의 가격이 선명하게 읽혀 글자 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "요구된 프레이밍과 조명 분위기를 잘 구현했으며 얕은 심도를 활용해 텍스트 노출 금지 지시를 훌륭하게 따랐습니다.  ★위반: [gpt-high] 오른쪽 위 벽 메뉴판에 읽을 수 있는 가격 숫자가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh6_sel.png",
    "asset_id": "a0733f24-ecf2-4137-acfd-737003a1607c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-b698-7f2e-a49a-9cf67a398e5d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S10sh6"
  }
 },
 "S10sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:18:17.613864+00:00",
  "fingerprint": "dbc5bcaf93454d901a79998a28cf55661be768793ca841a132448cefdba32bd2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S10sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S10sh7_sel.png",
  "source_sha256": "5eb67eaf03dffbbc63bd7d3187751d639b66b183b49b9749407714249ae9d833",
  "file": "S10sh7_cine.png",
  "staged_sha256": "e8f70a368be2d425867abe01dc300e6c8ffa1cce23cf3bc184fdf7c4a874e060",
  "latency_ms": 11155
 },
 "S11sh1::signage": {
  "fp": "3880848d190c97db",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S11sh1::ab_conti": {
  "input_fingerprint": "20f398e0fddffe58",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 화면 우측의 어두운 골목길을 향하고 있으며, 몸은 철문 틈새를 통과하는 방향으로 향해 있습니다.",
    "built_space": "화면 좌측에 마름모 장식이 있는 육중한 철창문이 닫히는 중이며, 우측에는 컨테이너 구조물들이 늘어선 골목길과 가로등, 그리고 우측 끝에 구형 조명과 확성기가 달린 기둥이 위치해 있습니다.",
    "entities": "현우(젊은 남성, 검은 머리, 회색 셔츠, 카고 바지)가 등장하며 얼굴과 팔다리에 상처와 핏자국이 묘사되어 레퍼런스와 일치합니다.",
    "hard_violations": [],
    "physics": "현우의 두 발은 바닥을 딛고 있으며, 오른손은 닫히는 철문의 끝부분을 짚고 있어 몸을 비틀어 빠져나가는 역동적인 자세를 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "주어진 레이아웃 스케치와 레퍼런스를 충실히 반영하여, 야간 배경 속 닫히는 철문 틈새로 몸을 구겨 넣는 현우의 전신과 부상 상태를 성공적으로 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측의 어두운 골목길을 향하고 있으며, 몸은 철문 틈새를 통과하는 방향으로 향해 있습니다.",
        "built_space": "화면 좌측에 마름모 장식이 있는 육중한 철창문이 닫히는 중이며, 우측에는 컨테이너 구조물들이 늘어선 골목길과 가로등, 그리고 우측 끝에 구형 조명과 확성기가 달린 기둥이 위치해 있습니다.",
        "entities": "현우(젊은 남성, 검은 머리, 회색 셔츠, 카고 바지)가 등장하며 얼굴과 팔다리에 상처와 핏자국이 묘사되어 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "현우의 두 발은 바닥을 딛고 있으며, 오른손은 닫히는 철문의 끝부분을 짚고 있어 몸을 비틀어 빠져나가는 역동적인 자세를 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "주어진 레이아웃 스케치와 레퍼런스를 충실히 반영하여, 야간 배경 속 닫히는 철문 틈새로 몸을 구겨 넣는 현우의 전신과 부상 상태를 성공적으로 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측의 어두운 골목길을 향하고 있으며, 몸은 철문 틈새를 통과하는 방향으로 향해 있습니다.",
        "built_space": "화면 좌측에 마름모 장식이 있는 육중한 철창문이 닫히는 중이며, 우측에는 컨테이너 구조물들이 늘어선 골목길과 가로등, 그리고 우측 끝에 구형 조명과 확성기가 달린 기둥이 위치해 있습니다.",
        "entities": "현우(젊은 남성, 검은 머리, 회색 셔츠, 카고 바지)가 등장하며 얼굴과 팔다리에 상처와 핏자국이 묘사되어 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "현우의 두 발은 바닥을 딛고 있으며, 오른손은 닫히는 철문의 끝부분을 짚고 있어 몸을 비틀어 빠져나가는 역동적인 자세를 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간 와이드숏과 문틈을 비집는 다급한 동작은 잘 구현했지만, 철문의 높이·장식·분할이 장소 사진과 다르고 정착촌 안으로 들어오는 이동 방향도 명확하지 않다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 골목을 향하고, 상체와 앞발도 오른쪽으로 나아간다. 뒤로 뻗은 손은 철문 세로 가장자리를 잡는다. 좁은 틈을 통과하는 행동은 분명하지만, 컨테이너 주거지가 오른쪽 골목에 보이므로 카메라가 정착촌 안에서 바깥을 본다는 조건과 현우의 안쪽 진입 경로를 동시에 명확히 확인하기는 어렵다.",
        "built_space": "왼쪽의 넓은 철문 면과 중앙의 좁은 철문 면 사이에 현우가 끼어 있으며, 하단에는 문 바퀴와 레일·문턱이 보인다. 오른쪽에는 구형 등 하나와 확성기 하나가 달린 콘크리트 기둥이 있고, 중앙 뒤쪽에도 작은 구형 등이 있는 기둥이 보인다. 오른쪽 배경에는 컨테이너 벽과 꺼진 가로등들이 이어진다. 배치 스케치의 큰 구도는 따르지만, 장소 사진의 낮고 여러 마름모 패널로 분할된 출입문 대신 사람보다 훨씬 높은 철창과 커다란 단일 마름모 장식을 보여 정확한 장소 구조 재현은 부족하다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 마른 체형은 인물 참고와 대체로 맞으며 국적은 외관만으로 확인할 수 없다. 회색 단추 셔츠, 올리브색 작업 바지와 어두운 부츠도 일치한다. 소매는 참고보다 걷혀 있고 팔과 바지에 상처 및 혈흔이 보이지만, 개에게 물린 상처인지와 기존 머리 부상의 상태는 명확하지 않다. 철문과 어두운 골목은 존재하며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 부츠가 문턱 안쪽 바닥에 닿아 체중을 받고, 뒤쪽 다리는 문틈 너머로 뻗어 있다. 뒤쪽 발의 접지는 문 가장자리와 어둠에 일부 가려져 있으나 앞발과 문을 잡은 손이 몸을 지지하므로 공중에 떠 있는 자세는 아니다. 굽힌 무릎과 기울여 비튼 상체는 급하게 좁은 틈을 통과하는 동작으로 가능하다. 철문은 하단 바퀴·레일과 기둥 구조로 지지되며, 정지 화면만으로 닫히는 운동 방향 자체는 확정하기 어렵다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간 와이드숏과 문틈을 비집는 다급한 동작은 잘 구현했지만, 철문의 높이·장식·분할이 장소 사진과 다르고 정착촌 안으로 들어오는 이동 방향도 명확하지 않다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 골목을 향하고, 상체와 앞발도 오른쪽으로 나아간다. 뒤로 뻗은 손은 철문 세로 가장자리를 잡는다. 좁은 틈을 통과하는 행동은 분명하지만, 컨테이너 주거지가 오른쪽 골목에 보이므로 카메라가 정착촌 안에서 바깥을 본다는 조건과 현우의 안쪽 진입 경로를 동시에 명확히 확인하기는 어렵다.",
        "built_space": "왼쪽의 넓은 철문 면과 중앙의 좁은 철문 면 사이에 현우가 끼어 있으며, 하단에는 문 바퀴와 레일·문턱이 보인다. 오른쪽에는 구형 등 하나와 확성기 하나가 달린 콘크리트 기둥이 있고, 중앙 뒤쪽에도 작은 구형 등이 있는 기둥이 보인다. 오른쪽 배경에는 컨테이너 벽과 꺼진 가로등들이 이어진다. 배치 스케치의 큰 구도는 따르지만, 장소 사진의 낮고 여러 마름모 패널로 분할된 출입문 대신 사람보다 훨씬 높은 철창과 커다란 단일 마름모 장식을 보여 정확한 장소 구조 재현은 부족하다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 마른 체형은 인물 참고와 대체로 맞으며 국적은 외관만으로 확인할 수 없다. 회색 단추 셔츠, 올리브색 작업 바지와 어두운 부츠도 일치한다. 소매는 참고보다 걷혀 있고 팔과 바지에 상처 및 혈흔이 보이지만, 개에게 물린 상처인지와 기존 머리 부상의 상태는 명확하지 않다. 철문과 어두운 골목은 존재하며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 부츠가 문턱 안쪽 바닥에 닿아 체중을 받고, 뒤쪽 다리는 문틈 너머로 뻗어 있다. 뒤쪽 발의 접지는 문 가장자리와 어둠에 일부 가려져 있으나 앞발과 문을 잡은 손이 몸을 지지하므로 공중에 떠 있는 자세는 아니다. 굽힌 무릎과 기울여 비튼 상체는 급하게 좁은 틈을 통과하는 동작으로 가능하다. 철문은 하단 바퀴·레일과 기둥 구조로 지지되며, 정지 화면만으로 닫히는 운동 방향 자체는 확정하기 어렵다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0
   },
   "adjusted": {
    "A": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000
  },
  "selected": "A",
  "ranking": [
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "주어진 레이아웃 스케치와 레퍼런스를 충실히 반영하여, 야간 배경 속 닫히는 철문 틈새로 몸을 구겨 넣는 현우의 전신과 부상 상태를 성공적으로 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "LAYOUT SKETCH — a bare thin-line layout guide, a REFERENCE ONLY: take from it ONLY the camera framing, figure placement, pose and size/depth order. It carries ZERO visual style — every texture, material, light and all realism come from the text and the photographic reference. Never let any line-drawing quality leak into the output.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S11sh1.png",
    "asset_id": "37dd0d43-2fdf-429b-82e4-045c520df701",
    "role": "conti_light"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06aaeb5d-51ab-7a16-a642-773277a8b1a2"
 },
 "S11sh4::signage": {
  "fp": "15ac91a737b02265",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::0e33b57ac265c530": {
  "subjects": [],
  "subject_text": "인천 난민촌 입구와 철창문 앞\n철창 출입문과 ‘난민 거주지역’ 표지판이 있는 진입로. 안쪽으로 낡은 철제 컨테이너들이 조밀하게 늘어서고 방송용 스피커가 설치돼 있다.",
  "identity": "canonical",
  "scope_id": "L163",
  "scope_role": "location_exterior",
  "scope_sha": "823ab6b5b4be054b"
 },
 "S11sh4::bgfirst_bg": {
  "input_fingerprint": "03c257130aebaf1b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4__bgfirst_bg.png",
  "asset_id": "ec7f53f2-f10f-4cb4-9c93-d5c6f78750b1",
  "input_asset_ids": [
   "478420ea-6c5c-40ae-a832-b5c0b14b89aa",
   "331ee13b-1a46-4cba-9a37-07f52e1d1493"
  ]
 },
 "S11sh4": {
  "input_fingerprint": "0bef826050e5a206",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The home container is marked 7-31, and its door is open for entry. The surrounding camp streets remain dark after curfew lights-out. 현우: Hyunwoo is entering his home with his upper garment still on; his face and leg remain injured, and the leg wound has not yet been treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The home container is marked 7-31, and its door is open for entry. The surrounding camp streets remain dark after curfew lights-out. 현우: Hyunwoo is entering his home with his upper garment still on; his face and leg remain injured, and the leg wound has not yet been treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The home container is marked 7-31, and its door is open for entry. The surrounding camp streets remain dark after curfew lights-out. 현우: Hyunwoo is entering his home with his upper garment still on; his face and leg remain injured, and the leg wound has not yet been treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4__bgfirst_bg.png",
     "asset_id": "ec7f53f2-f10f-4cb4-9c93-d5c6f78750b1",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S11sh4.png",
     "asset_id": "478420ea-6c5c-40ae-a832-b5c0b14b89aa",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
     "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우가 어두운 컨테이너 내부를 향해 똑바로 걸어 들어가는 방향입니다.",
    "built_space": "레퍼런스와 일치하는 포장된 도로와 늘어선 주거용 컨테이너들입니다. 우측 컨테이너 벽면에 '7-31'이 적혀 있고 측면 규격 출입문이 열려 있습니다.",
    "entities": "뒷모습의 현우입니다. 지시된 회색 셔츠와 올리브색 바지를 입고 있으며, 왼쪽 종아리 부분의 찢어진 옷과 상처가 보입니다.",
    "hard_violations": [
     "[gpt-high] 외벽의 7-31이 판독 가능하여 읽을 수 있는 글자 금지 조건을 위반합니다."
    ],
    "physics": "오른발로 바깥 지면을 딛고 왼발을 들어 컨테이너 문지방 안쪽으로 내딛고 있습니다."
   },
   {
    "label": "B",
    "direction": "몸은 컨테이너 안을 향하고 있으나, 고개를 오른쪽 뒤로 돌려 카메라 쪽을 응시하고 있습니다.",
    "built_space": "비포장 진흙 바닥에 화물용 양문형 컨테이너가 놓여 있습니다. '7-31'이 열린 문짝에 크게 그려져 있어 레퍼런스의 환경(포장도로, 일반 출입문)과 전혀 다릅니다.",
    "entities": "고개를 돌린 현우입니다. 얼굴의 상처와 지정된 의상은 일치하게 나타납니다.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스에 명시된 장소(포장도로 및 개조된 주거용 컨테이너)를 무시하고, 존재하지 않는 비포장 진흙길과 화물용 양문형 컨테이너를 임의로 생성함.",
     "[gpt-high] 참고 장소의 단일 주거용 출입문을 잠금봉이 달린 화물용 양문 구조로 대체했습니다.",
     "[gpt-high] 문짝에 적힌 7-31이 선명하게 판독되어 읽을 수 있는 글자 금지 조건을 위반합니다."
    ],
    "physics": "왼발로 진흙 바닥을 지탱하고 오른발을 컨테이너 입구 턱 위로 올리고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스 이미지의 장소(포장도로, 규격화된 출입문)를 정확히 구현했으며, 프롬프트가 지시한 뒷모습과 다리 상처까지 충실하게 표현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "레퍼런스와 전혀 다른 진흙길과 화물용 양문형 컨테이너를 생성하여 장소 일치 조건을 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우가 어두운 컨테이너 내부를 향해 똑바로 걸어 들어가는 방향입니다.",
        "built_space": "레퍼런스와 일치하는 포장된 도로와 늘어선 주거용 컨테이너들입니다. 우측 컨테이너 벽면에 '7-31'이 적혀 있고 측면 규격 출입문이 열려 있습니다.",
        "entities": "뒷모습의 현우입니다. 지시된 회색 셔츠와 올리브색 바지를 입고 있으며, 왼쪽 종아리 부분의 찢어진 옷과 상처가 보입니다.",
        "hard_violations": [],
        "physics": "오른발로 바깥 지면을 딛고 왼발을 들어 컨테이너 문지방 안쪽으로 내딛고 있습니다."
       },
       {
        "label": "B",
        "direction": "몸은 컨테이너 안을 향하고 있으나, 고개를 오른쪽 뒤로 돌려 카메라 쪽을 응시하고 있습니다.",
        "built_space": "비포장 진흙 바닥에 화물용 양문형 컨테이너가 놓여 있습니다. '7-31'이 열린 문짝에 크게 그려져 있어 레퍼런스의 환경(포장도로, 일반 출입문)과 전혀 다릅니다.",
        "entities": "고개를 돌린 현우입니다. 얼굴의 상처와 지정된 의상은 일치하게 나타납니다.",
        "hard_violations": [
         "레퍼런스에 명시된 장소(포장도로 및 개조된 주거용 컨테이너)를 무시하고, 존재하지 않는 비포장 진흙길과 화물용 양문형 컨테이너를 임의로 생성함."
        ],
        "physics": "왼발로 진흙 바닥을 지탱하고 오른발을 컨테이너 입구 턱 위로 올리고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스 이미지의 장소(포장도로, 규격화된 출입문)를 정확히 구현했으며, 프롬프트가 지시한 뒷모습과 다리 상처까지 충실하게 표현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "레퍼런스와 전혀 다른 진흙길과 화물용 양문형 컨테이너를 생성하여 장소 일치 조건을 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우가 어두운 컨테이너 내부를 향해 똑바로 걸어 들어가는 방향입니다.",
        "built_space": "레퍼런스와 일치하는 포장된 도로와 늘어선 주거용 컨테이너들입니다. 우측 컨테이너 벽면에 '7-31'이 적혀 있고 측면 규격 출입문이 열려 있습니다.",
        "entities": "뒷모습의 현우입니다. 지시된 회색 셔츠와 올리브색 바지를 입고 있으며, 왼쪽 종아리 부분의 찢어진 옷과 상처가 보입니다.",
        "hard_violations": [],
        "physics": "오른발로 바깥 지면을 딛고 왼발을 들어 컨테이너 문지방 안쪽으로 내딛고 있습니다."
       },
       {
        "label": "B",
        "direction": "몸은 컨테이너 안을 향하고 있으나, 고개를 오른쪽 뒤로 돌려 카메라 쪽을 응시하고 있습니다.",
        "built_space": "비포장 진흙 바닥에 화물용 양문형 컨테이너가 놓여 있습니다. '7-31'이 열린 문짝에 크게 그려져 있어 레퍼런스의 환경(포장도로, 일반 출입문)과 전혀 다릅니다.",
        "entities": "고개를 돌린 현우입니다. 얼굴의 상처와 지정된 의상은 일치하게 나타납니다.",
        "hard_violations": [
         "레퍼런스에 명시된 장소(포장도로 및 개조된 주거용 컨테이너)를 무시하고, 존재하지 않는 비포장 진흙길과 화물용 양문형 컨테이너를 임의로 생성함."
        ],
        "physics": "왼발로 진흙 바닥을 지탱하고 오른발을 컨테이너 입구 턱 위로 올리고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "문턱 밖에서 뒤돌아보는 동작이며, 주거용 출입문을 화물용 양문으로 바꾸고 번호를 크게 노출해 핵심 진입 순간과 장소 고정을 놓쳤습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "골목에서 본 와이드 뒷모습과 문턱을 넘는 발, 단일 주거용 문은 더 충실하지만, 켜진 가로등과 판독 가능한 번호는 명시 조건을 위반합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "몸통은 열린 입구 쪽이지만 고개는 오른쪽 어깨 너머 골목으로 돌아가 있습니다. 시선의 구체적인 대상은 보이지 않으며, 집 안을 향해 들어가는 집중된 동작보다는 밖을 돌아보는 순간입니다.",
        "built_space": "화물 컨테이너 끝면에 왼쪽 닫힌 문짝 하나와 오른쪽 열린 문짝 하나, 왼쪽 문짝의 수직 잠금봉 두 개가 보입니다. 참고 장소의 작은 단일 주거용 출입문과 다른 구조입니다. 인물은 높은 금속 문턱 바로 밖에 있고, 오른쪽 문짝은 비스듬히 열려 몸을 가리지 않습니다. 주변은 참고의 포장 통로가 아니라 진흙과 잔해가 있는 골목으로 표현됐습니다.",
        "entities": "인물은 한 명이며 검은 헝클어진 머리, 젊은 동아시아계 남성 외형, 회색 계열 셔츠와 카고바지, 부츠는 참고와 대체로 맞습니다. 보이는 얼굴에는 상처가 있으나 다리 상처는 확인되지 않습니다. 셔츠는 참고와 달리 바지 밖으로 나와 있습니다. 입구 옆의 7-31은 지나치게 크고 명확히 읽혀 최종 글자 비가독성 지시와 충돌합니다. 실내는 어둡습니다.",
        "hard_violations": [
         "참고 장소의 단일 주거용 출입문을 잠금봉이 달린 화물용 양문 구조로 대체했습니다.",
         "문짝에 적힌 7-31이 선명하게 판독되어 읽을 수 있는 글자 금지 조건을 위반합니다."
        ],
        "physics": "왼발은 문턱 밖 지면에 닿아 체중을 지지하고 오른발은 뒤꿈치를 든 보행 자세로 보입니다. 공중에 뜬 신체나 지지 없는 물체는 없습니다. 다만 두 부츠 모두 문턱 바깥에 있어 한 발을 이미 실내로 내디딘 순간은 구현되지 않았습니다."
       },
       {
        "label": "B",
        "direction": "등은 카메라를 향하고 몸통과 고개는 열린 출입구 안쪽을 향합니다. 눈 자체는 보이지 않지만 머리 방향은 실내 목적지와 맞고, 들어 올린 앞발도 문턱 안쪽으로 진행합니다.",
        "built_space": "주된 출입구는 하나이고 오른쪽에 손잡이가 달린 열린 문짝 하나가 있습니다. 입구 왼쪽에는 작은 창 하나와 상하로 배치된 설비 상자 두 개, 오른쪽 가장자리에는 창 하나가 보입니다. 금속 외벽과 주거용 문, 포장 통로, 반복되는 컨테이너와 외부 설비는 참고 장소에 더 가깝습니다. 카메라는 골목에서 뒤쪽 발보다 높은 위치로 문턱을 내려다보며, 문짝이 등과 발의 동작을 가리지 않습니다. 다만 참고의 해당 집 옆 좁은 통로보다 넓은 도로가 강조됐습니다.",
        "entities": "한 명의 젊고 마른 남성이 등장하며 검은 머리, 회색 긴팔 셔츠, 벨트, 카고바지와 부츠가 참고 복장에 대체로 부합합니다. 얼굴이 돌아서 있어 정확한 얼굴 일치와 얼굴 상처는 판단할 수 없습니다. 뒤쪽 다리에는 찢어지고 어두워진 상처 부위가 보이며 붕대는 없습니다. 입구 왼쪽 7-31은 읽을 수 있습니다. 밤이지만 골목의 여러 가로등이 켜져 있어 소등된 정착촌 조건과 다릅니다. 문 안에 밝게 켜진 실내조명은 보이지 않습니다.",
        "hard_violations": [
         "외벽의 7-31이 판독 가능하여 읽을 수 있는 글자 금지 조건을 위반합니다."
        ],
        "physics": "뒤쪽 부츠가 외부 포장면에 닿아 몸을 지지하고, 반대쪽 다리는 무릎을 굽혀 앞발을 문턱 안쪽으로 올리고 있습니다. 발이 향하는 착지면은 실내 바닥이며 실제 진입 보행으로 가능한 자세입니다. 열린 문은 문틀의 경첩으로 지지되고, 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "문턱 밖에서 뒤돌아보는 동작이며, 주거용 출입문을 화물용 양문으로 바꾸고 번호를 크게 노출해 핵심 진입 순간과 장소 고정을 놓쳤습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "골목에서 본 와이드 뒷모습과 문턱을 넘는 발, 단일 주거용 문은 더 충실하지만, 켜진 가로등과 판독 가능한 번호는 명시 조건을 위반합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "몸통은 열린 입구 쪽이지만 고개는 오른쪽 어깨 너머 골목으로 돌아가 있습니다. 시선의 구체적인 대상은 보이지 않으며, 집 안을 향해 들어가는 집중된 동작보다는 밖을 돌아보는 순간입니다.",
        "built_space": "화물 컨테이너 끝면에 왼쪽 닫힌 문짝 하나와 오른쪽 열린 문짝 하나, 왼쪽 문짝의 수직 잠금봉 두 개가 보입니다. 참고 장소의 작은 단일 주거용 출입문과 다른 구조입니다. 인물은 높은 금속 문턱 바로 밖에 있고, 오른쪽 문짝은 비스듬히 열려 몸을 가리지 않습니다. 주변은 참고의 포장 통로가 아니라 진흙과 잔해가 있는 골목으로 표현됐습니다.",
        "entities": "인물은 한 명이며 검은 헝클어진 머리, 젊은 동아시아계 남성 외형, 회색 계열 셔츠와 카고바지, 부츠는 참고와 대체로 맞습니다. 보이는 얼굴에는 상처가 있으나 다리 상처는 확인되지 않습니다. 셔츠는 참고와 달리 바지 밖으로 나와 있습니다. 입구 옆의 7-31은 지나치게 크고 명확히 읽혀 최종 글자 비가독성 지시와 충돌합니다. 실내는 어둡습니다.",
        "hard_violations": [
         "참고 장소의 단일 주거용 출입문을 잠금봉이 달린 화물용 양문 구조로 대체했습니다.",
         "문짝에 적힌 7-31이 선명하게 판독되어 읽을 수 있는 글자 금지 조건을 위반합니다."
        ],
        "physics": "왼발은 문턱 밖 지면에 닿아 체중을 지지하고 오른발은 뒤꿈치를 든 보행 자세로 보입니다. 공중에 뜬 신체나 지지 없는 물체는 없습니다. 다만 두 부츠 모두 문턱 바깥에 있어 한 발을 이미 실내로 내디딘 순간은 구현되지 않았습니다."
       },
       {
        "label": "A",
        "direction": "등은 카메라를 향하고 몸통과 고개는 열린 출입구 안쪽을 향합니다. 눈 자체는 보이지 않지만 머리 방향은 실내 목적지와 맞고, 들어 올린 앞발도 문턱 안쪽으로 진행합니다.",
        "built_space": "주된 출입구는 하나이고 오른쪽에 손잡이가 달린 열린 문짝 하나가 있습니다. 입구 왼쪽에는 작은 창 하나와 상하로 배치된 설비 상자 두 개, 오른쪽 가장자리에는 창 하나가 보입니다. 금속 외벽과 주거용 문, 포장 통로, 반복되는 컨테이너와 외부 설비는 참고 장소에 더 가깝습니다. 카메라는 골목에서 뒤쪽 발보다 높은 위치로 문턱을 내려다보며, 문짝이 등과 발의 동작을 가리지 않습니다. 다만 참고의 해당 집 옆 좁은 통로보다 넓은 도로가 강조됐습니다.",
        "entities": "한 명의 젊고 마른 남성이 등장하며 검은 머리, 회색 긴팔 셔츠, 벨트, 카고바지와 부츠가 참고 복장에 대체로 부합합니다. 얼굴이 돌아서 있어 정확한 얼굴 일치와 얼굴 상처는 판단할 수 없습니다. 뒤쪽 다리에는 찢어지고 어두워진 상처 부위가 보이며 붕대는 없습니다. 입구 왼쪽 7-31은 읽을 수 있습니다. 밤이지만 골목의 여러 가로등이 켜져 있어 소등된 정착촌 조건과 다릅니다. 문 안에 밝게 켜진 실내조명은 보이지 않습니다.",
        "hard_violations": [
         "외벽의 7-31이 판독 가능하여 읽을 수 있는 글자 금지 조건을 위반합니다."
        ],
        "physics": "뒤쪽 부츠가 외부 포장면에 닿아 몸을 지지하고, 반대쪽 다리는 무릎을 굽혀 앞발을 문턱 안쪽으로 올리고 있습니다. 발이 향하는 착지면은 실내 바닥이며 실제 진입 보행으로 가능한 자세입니다. 열린 문은 문틀의 경첩으로 지지되고, 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 레퍼런스에 명시된 장소(포장도로 및 개조된 주거용 컨테이너)를 무시하고, 존재하지 않는 비포장 진흙길과 화물용 양문형 컨테이너를 임의로 생성함.",
     "[gpt-high] 참고 장소의 단일 주거용 출입문을 잠금봉이 달린 화물용 양문 구조로 대체했습니다.",
     "[gpt-high] 문짝에 적힌 7-31이 선명하게 판독되어 읽을 수 있는 글자 금지 조건을 위반합니다."
    ],
    "A": [
     "[gpt-high] 외벽의 7-31이 판독 가능하여 읽을 수 있는 글자 금지 조건을 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스 이미지의 장소(포장도로, 규격화된 출입문)를 정확히 구현했으며, 프롬프트가 지시한 뒷모습과 다리 상처까지 충실하게 표현했습니다.  ★위반: [gpt-high] 외벽의 7-31이 판독 가능하여 읽을 수 있는 글자 금지 조건을 위반합니다."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "레퍼런스와 전혀 다른 진흙길과 화물용 양문형 컨테이너를 생성하여 장소 일치 조건을 심각하게 위반했습니다.  ★위반: [gemini-pro] 레퍼런스에 명시된 장소(포장도로 및 개조된 주거용 컨테이너)를 무시하고, 존재하지 않는 비포장 진흙길과 화물용 양문형 컨테이너를 임의로 생성함. / [gpt-high] 참고 장소의 단일 주거용 출입문을 잠금봉이 달린 화물용 양문 구조로 대체했습니다. / [gpt-high] 문짝에 적힌 7-31이 선명하게 판독되어 읽을 수 있는 글자 금지 조건을 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-bb6c-7e9a-8052-2970236034f6",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4__bgfirst_bg.png",
   "bg_asset_id": "ec7f53f2-f10f-4cb4-9c93-d5c6f78750b1",
   "bg_record_key": "S11sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "seed_bg"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S11sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:21:27.101199+00:00",
  "fingerprint": "2165e9a4713156f78e945147defa81bc310c0280278ff428c3ea38d95a394d37",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S11sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S11sh4_sel.png",
  "source_sha256": "4f25e17ccb729bbeb6ac2b11f521fd9a098e2b7ae80fc0e663f4aa211a6436ac",
  "file": "S11sh4_cine.png",
  "staged_sha256": "95dc22d0f2755b145d4e1839ac60c91e9831d33f7c9bfa3a17fddba061be061b",
  "latency_ms": 10476
 },
 "S12sh7::signage": {
  "fp": "b0d7431472acf0b6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::f424efd76bc0d46e": {
  "subjects": [],
  "subject_text": "현우의 컨테이너 내부\n식탁과 생활용품을 둔 비좁은 컨테이너 주거 공간. 안쪽에 작은 방으로 연결되는 문이 있고, 밤에는 식탁의 촛불이 실내를 밝힌다.",
  "identity": "canonical",
  "scope_id": "L166",
  "scope_role": "location_interior",
  "scope_sha": "ad03e7fae5aaa48f"
 },
 "S12sh7::bgfirst_bg": {
  "input_fingerprint": "bb68b70ba912526a",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7__bgfirst_bg.png",
  "asset_id": "520d1ce5-4724-4ae7-a50a-649a26087ab2",
  "input_asset_ids": [
   "70be5339-3395-4c59-bd14-9f3f92278784",
   "a4110c6d-4f09-489a-91bf-d6c8e7837b22"
  ]
 },
 "S12sh7": {
  "input_fingerprint": "98ba30e40a13d693",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A candle continues to light the dining table, where the meal was served; the mending clothes remain in the room and the medicine container is now open. 현우: Hyunwoo has removed his upper garment and is seated at the table with visible facial injuries and a bleeding dog-bite wound on his leg. 미연: Miyeon is at the dining table with the opened medicine container, preparing to examine the leg injury.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 미연 right now, so 미연's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 미연: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A candle continues to light the dining table, where the meal was served; the mending clothes remain in the room and the medicine container is now open. 현우: Hyunwoo has removed his upper garment and is seated at the table with visible facial injuries and a bleeding dog-bite wound on his leg. 미연: Miyeon is at the dining table with the opened medicine container, preparing to examine the leg injury.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 미연 right now, so 미연's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 미연: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A candle continues to light the dining table, where the meal was served; the mending clothes remain in the room and the medicine container is now open. 현우: Hyunwoo has removed his upper garment and is seated at the table with visible facial injuries and a bleeding dog-bite wound on his leg. 미연: Miyeon is at the dining table with the opened medicine container, preparing to examine the leg injury.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 미연 right now, so 미연's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 미연: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7__bgfirst_bg.png",
     "asset_id": "520d1ce5-4724-4ae7-a50a-649a26087ab2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S12sh7.png",
     "asset_id": "70be5339-3395-4c59-bd14-9f3f92278784",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L166B01.png",
     "asset_id": "a4110c6d-4f09-489a-91bf-d6c8e7837b22",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 미연이 내민 약병을 정확히 향하고 있으며 상체를 뒤로 물리고 있음.",
    "built_space": "식탁, 벤치, 배경의 선반과 냉장고가 참조 이미지와 일치함. 두 인물 모두 의자에 자연스럽게 착석해 있음.",
    "entities": "현우의 얼굴 상처와 다리 출혈이 정확히 반영됨. 식사와 촛불이 존재하나 약병 뚜껑이 닫혀 있음.",
    "hard_violations": [],
    "physics": "모든 사물과 인물이 바닥과 가구에 안정적으로 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 열린 약병을 향하고 있음.",
    "built_space": "식탁은 존재하나 배경의 주요 구조물(선반, 냉장고 등)이 누락되어 참조된 장소의 디테일이 부족함.",
    "entities": "다리의 피 나는 상처 대신 붕대가 감겨 있어 지시와 어긋남. 약병은 열려 있음.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 인체 구조 (약병을 잡은 미연의 손가락과 손목 관절이 심하게 왜곡됨)"
    ],
    "physics": "가구에 올바르게 앉아 있으나, 미연의 뻗은 팔과 손의 뼈대가 해부학적으로 성립하지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "상체를 뒤로 물린 자세, 배경, 다리의 상처 등 대부분의 요소를 훌륭히 구현했으나 약병이 닫혀 있어 일부 지시를 놓침."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "다리의 출혈 상처가 붕대로 가려져 지시를 위반했으며, 약병을 든 손의 인체 구조가 심각하게 왜곡됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 미연이 내민 약병을 정확히 향하고 있으며 상체를 뒤로 물리고 있음.",
        "built_space": "식탁, 벤치, 배경의 선반과 냉장고가 참조 이미지와 일치함. 두 인물 모두 의자에 자연스럽게 착석해 있음.",
        "entities": "현우의 얼굴 상처와 다리 출혈이 정확히 반영됨. 식사와 촛불이 존재하나 약병 뚜껑이 닫혀 있음.",
        "hard_violations": [],
        "physics": "모든 사물과 인물이 바닥과 가구에 안정적으로 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 열린 약병을 향하고 있음.",
        "built_space": "식탁은 존재하나 배경의 주요 구조물(선반, 냉장고 등)이 누락되어 참조된 장소의 디테일이 부족함.",
        "entities": "다리의 피 나는 상처 대신 붕대가 감겨 있어 지시와 어긋남. 약병은 열려 있음.",
        "hard_violations": [
         "물리적으로 불가능한 인체 구조 (약병을 잡은 미연의 손가락과 손목 관절이 심하게 왜곡됨)"
        ],
        "physics": "가구에 올바르게 앉아 있으나, 미연의 뻗은 팔과 손의 뼈대가 해부학적으로 성립하지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "상체를 뒤로 물린 자세, 배경, 다리의 상처 등 대부분의 요소를 훌륭히 구현했으나 약병이 닫혀 있어 일부 지시를 놓침."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "다리의 출혈 상처가 붕대로 가려져 지시를 위반했으며, 약병을 든 손의 인체 구조가 심각하게 왜곡됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 미연이 내민 약병을 정확히 향하고 있으며 상체를 뒤로 물리고 있음.",
        "built_space": "식탁, 벤치, 배경의 선반과 냉장고가 참조 이미지와 일치함. 두 인물 모두 의자에 자연스럽게 착석해 있음.",
        "entities": "현우의 얼굴 상처와 다리 출혈이 정확히 반영됨. 식사와 촛불이 존재하나 약병 뚜껑이 닫혀 있음.",
        "hard_violations": [],
        "physics": "모든 사물과 인물이 바닥과 가구에 안정적으로 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 열린 약병을 향하고 있음.",
        "built_space": "식탁은 존재하나 배경의 주요 구조물(선반, 냉장고 등)이 누락되어 참조된 장소의 디테일이 부족함.",
        "entities": "다리의 피 나는 상처 대신 붕대가 감겨 있어 지시와 어긋남. 약병은 열려 있음.",
        "hard_violations": [
         "물리적으로 불가능한 인체 구조 (약병을 잡은 미연의 손가락과 손목 관절이 심하게 왜곡됨)"
        ],
        "physics": "가구에 올바르게 앉아 있으나, 미연의 뻗은 팔과 손의 뼈대가 해부학적으로 성립하지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 약병이 아니라 미연의 얼굴을 노려보고 상체도 거의 곧게 세워, 핵심인 약병을 향한 시선과 뒤로 물러나는 거부 동작을 놓쳤다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "뒤로 물린 상체, 약병 쪽으로 낮춘 차가운 시선, 가슴 아래 내민 열린 약병과 몸 사이의 여백을 구현하고 장소와 복장도 더 충실히 맞췄다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈은 화면 오른쪽 위에 있는 미연의 얼굴을 향하며, 그보다 아래의 약병을 향하지 않는다. 미연은 현우 쪽을 내려다보며 오른쪽에서 왼쪽으로 약병을 내민다. 병은 현우의 가슴 아래가 아니라 가슴 높이에 놓여 있다.",
        "built_space": "하단을 비스듬히 가로지르는 식탁 하나와 뒤쪽 나무 등받이 벤치 하나가 보인다. 벤치 위에는 쿠션과 접힌 옷이 있고, 벽에는 사진 일부와 오른쪽 배관이 보인다. 컨테이너 식당의 재료와 분위기는 대체로 맞지만, 좁은 배경 때문에 참조 장소의 냉장고와 선반 배치는 확인되지 않는다. 현우는 식탁 뒤 좌석에 앉아 있고 미연은 오른쪽 가장자리에서 몸을 기울인다.",
        "entities": "현우는 앳된 동아시아계 남성으로 검은 머리, 벗은 상체, 볼과 코 주변 상처가 보인다. 참조보다 머리가 단정하다. 미연은 검은 머리의 중년 동아시아계 여성으로 보이지만, 보이는 소매는 참조의 체크 셔츠가 아닌 무지 회색이다. 손에는 열린 작은 약병이 있고 라벨 인쇄는 명확히 판독되지 않는다. 식탁에는 촛불 하나와 먹던 음식, 식기와 컵이 남아 있다. 다리에는 붕대가 보여 출혈 중인 개 물림 상처 자체는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "현우는 좌석에 앉아 있으나 상체가 거의 수직이어서 뒤로 물린 동작이 약하다. 들어 올린 무릎의 직접 받침은 가려졌지만 앉아서 다리를 굽힌 자세로 가능하다. 미연의 손가락이 약병을 감싸 받치며 팔은 화면 오른쪽 몸통에 연결된다. 음식과 촛불은 식탁 위에 놓여 있고, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 턱을 낮추고 화면 왼쪽 아래의 내민 약병 쪽을 노려본다. 미연의 보이는 옆얼굴은 현우를 향한다. 미연의 팔은 왼쪽에서 오른쪽으로 뻗고, 열린 약병은 현우의 가슴 하단보다 낮은 위치에서 몸통과 떨어져 있다.",
        "built_space": "식탁 하나가 하단을 대각선으로 가로지르고 현우는 벽 쪽 나무 등받이 벤치 하나에 앉아 있다. 뒤에는 냉장고 하나, 검은 철제 선반 하나와 그 안의 전자레인지 하나, 벽 사진군, 걸린 식물, 오른쪽 전선관과 전기함 하나가 보인다. 참조의 컨테이너 벽체와 주방·식탁의 공간 관계를 잘 유지한다. 미연은 식탁 맞은편 왼쪽 전경에 있으며, 현우와 약병 사이의 빈 공간이 분명하다.",
        "entities": "현우는 헝클어진 검은 머리와 앳된 얼굴의 동아시아계 남성으로 참조 인상에 가깝고, 상의를 벗었으며 얼굴 상처와 피가 흐르는 무릎 부근 상처가 보인다. 상처의 원인이 개 물림인지는 영상만으로 확정하기 어렵다. 미연은 검은 머리의 중년 동아시아계 여성으로 보이고 체크 셔츠와 앞치마 끈이 참조에 부합한다. 손에는 뚜껑이 제거된 갈색 약병이 있으며 읽을 수 있는 글자는 없다. 식탁에는 촛불 하나, 밥과 다른 음식이 든 그릇, 젓가락과 컵이 있다. 현우 무릎 위에 옷감이 있으나 수선 중인 옷인지는 불분명하다.",
        "hard_violations": [],
        "physics": "현우의 골반은 벤치 좌판에 놓이고 상체는 등받이 쪽으로 물러나 있다. 한 손은 무릎 부근에, 다른 손은 좌판 가장자리에 놓여 있어 거리를 벌리는 자세가 자연스럽다. 미연은 손가락으로 약병 몸체를 확실히 잡고 팔을 뻗는다. 촛불은 받침 위에, 받침과 식기는 식탁 위에 놓여 있어 지지가 모두 성립한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우가 약병이 아니라 미연의 얼굴을 노려보고 상체도 거의 곧게 세워, 핵심인 약병을 향한 시선과 뒤로 물러나는 거부 동작을 놓쳤다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "뒤로 물린 상체, 약병 쪽으로 낮춘 차가운 시선, 가슴 아래 내민 열린 약병과 몸 사이의 여백을 구현하고 장소와 복장도 더 충실히 맞췄다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈은 화면 오른쪽 위에 있는 미연의 얼굴을 향하며, 그보다 아래의 약병을 향하지 않는다. 미연은 현우 쪽을 내려다보며 오른쪽에서 왼쪽으로 약병을 내민다. 병은 현우의 가슴 아래가 아니라 가슴 높이에 놓여 있다.",
        "built_space": "하단을 비스듬히 가로지르는 식탁 하나와 뒤쪽 나무 등받이 벤치 하나가 보인다. 벤치 위에는 쿠션과 접힌 옷이 있고, 벽에는 사진 일부와 오른쪽 배관이 보인다. 컨테이너 식당의 재료와 분위기는 대체로 맞지만, 좁은 배경 때문에 참조 장소의 냉장고와 선반 배치는 확인되지 않는다. 현우는 식탁 뒤 좌석에 앉아 있고 미연은 오른쪽 가장자리에서 몸을 기울인다.",
        "entities": "현우는 앳된 동아시아계 남성으로 검은 머리, 벗은 상체, 볼과 코 주변 상처가 보인다. 참조보다 머리가 단정하다. 미연은 검은 머리의 중년 동아시아계 여성으로 보이지만, 보이는 소매는 참조의 체크 셔츠가 아닌 무지 회색이다. 손에는 열린 작은 약병이 있고 라벨 인쇄는 명확히 판독되지 않는다. 식탁에는 촛불 하나와 먹던 음식, 식기와 컵이 남아 있다. 다리에는 붕대가 보여 출혈 중인 개 물림 상처 자체는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "현우는 좌석에 앉아 있으나 상체가 거의 수직이어서 뒤로 물린 동작이 약하다. 들어 올린 무릎의 직접 받침은 가려졌지만 앉아서 다리를 굽힌 자세로 가능하다. 미연의 손가락이 약병을 감싸 받치며 팔은 화면 오른쪽 몸통에 연결된다. 음식과 촛불은 식탁 위에 놓여 있고, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 턱을 낮추고 화면 왼쪽 아래의 내민 약병 쪽을 노려본다. 미연의 보이는 옆얼굴은 현우를 향한다. 미연의 팔은 왼쪽에서 오른쪽으로 뻗고, 열린 약병은 현우의 가슴 하단보다 낮은 위치에서 몸통과 떨어져 있다.",
        "built_space": "식탁 하나가 하단을 대각선으로 가로지르고 현우는 벽 쪽 나무 등받이 벤치 하나에 앉아 있다. 뒤에는 냉장고 하나, 검은 철제 선반 하나와 그 안의 전자레인지 하나, 벽 사진군, 걸린 식물, 오른쪽 전선관과 전기함 하나가 보인다. 참조의 컨테이너 벽체와 주방·식탁의 공간 관계를 잘 유지한다. 미연은 식탁 맞은편 왼쪽 전경에 있으며, 현우와 약병 사이의 빈 공간이 분명하다.",
        "entities": "현우는 헝클어진 검은 머리와 앳된 얼굴의 동아시아계 남성으로 참조 인상에 가깝고, 상의를 벗었으며 얼굴 상처와 피가 흐르는 무릎 부근 상처가 보인다. 상처의 원인이 개 물림인지는 영상만으로 확정하기 어렵다. 미연은 검은 머리의 중년 동아시아계 여성으로 보이고 체크 셔츠와 앞치마 끈이 참조에 부합한다. 손에는 뚜껑이 제거된 갈색 약병이 있으며 읽을 수 있는 글자는 없다. 식탁에는 촛불 하나, 밥과 다른 음식이 든 그릇, 젓가락과 컵이 있다. 현우 무릎 위에 옷감이 있으나 수선 중인 옷인지는 불분명하다.",
        "hard_violations": [],
        "physics": "현우의 골반은 벤치 좌판에 놓이고 상체는 등받이 쪽으로 물러나 있다. 한 손은 무릎 부근에, 다른 손은 좌판 가장자리에 놓여 있어 거리를 벌리는 자세가 자연스럽다. 미연은 손가락으로 약병 몸체를 확실히 잡고 팔을 뻗는다. 촛불은 받침 위에, 받침과 식기는 식탁 위에 놓여 있어 지지가 모두 성립한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 인체 구조 (약병을 잡은 미연의 손가락과 손목 관절이 심하게 왜곡됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "상체를 뒤로 물린 자세, 배경, 다리의 상처 등 대부분의 요소를 훌륭히 구현했으나 약병이 닫혀 있어 일부 지시를 놓침."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "다리의 출혈 상처가 붕대로 가려져 지시를 위반했으며, 약병을 든 손의 인체 구조가 심각하게 왜곡됨.  ★위반: [gemini-pro] 물리적으로 불가능한 인체 구조 (약병을 잡은 미연의 손가락과 손목 관절이 심하게 왜곡됨)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L166B01.png",
    "asset_id": "a4110c6d-4f09-489a-91bf-d6c8e7837b22",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-bea5-7f95-b3c8-7015fdfaba7e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7__bgfirst_bg.png",
   "bg_asset_id": "520d1ce5-4724-4ae7-a50a-649a26087ab2",
   "bg_record_key": "S12sh7::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S12sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:22:49.931356+00:00",
  "fingerprint": "695ecb1ccffa0a9c4976c1270e71c5467225343b8ea7c52d7e143a58ad554a99",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S12sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S12sh7_sel.png",
  "source_sha256": "adfe94015fb82e4a04ac867c543a09769ef07f98dd2dc2adb754d076aabe54d4",
  "file": "S12sh7_cine.png",
  "staged_sha256": "16319cd8a9c379eb0bfe129671e73731f7d5eb20546e5752c8021518b64a8e75",
  "latency_ms": 11876
 },
 "S12sh8::signage": {
  "fp": "be121b0c2b1fc2f0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S12sh8": {
  "input_fingerprint": "e9ff60747e39f8b8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손이 식탁에 강하게 맞닿은 충격으로 촛불의 불꽃이 크게 꺾인 찰나, 양팔을 뻗은 채 굳어 있는 현우의 상체.\n\nLOCATION (lock): At the dining table inside the container home, illuminated by the candle whose flame bends with the impact. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Both of 현우's hands have struck its top) — The tabletop is seen at a shallow angle beneath his extended arms; used as Connects the two hand contacts and gives the impact a stable spatial base; Tabletop candle (Flame sharply bent at the instant of impact) — Seen side-on in the lower-left field, separate from the hand silhouettes; used as Provides a small, readable physical consequence of the impact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established candlelight retains its restrained warmth, with a momentary unevenness as the flame bends.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container dining area's table, interior surfaces, and candlelit nighttime atmosphere. Exclude the medicine bottle as a hand-held foreground element; do not transfer it from the earlier interaction.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The table candle remains the room's established light source, with the meal setting and opened medicine container still present. There is no established electrical lighting after curfew. 현우: Hyunwoo still has his upper garment off, with injured facial skin and a bleeding leg wound; no dressing has been applied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손이 식탁에 강하게 맞닿은 충격으로 촛불의 불꽃이 크게 꺾인 찰나, 양팔을 뻗은 채 굳어 있는 현우의 상체.\n\nLOCATION (lock): At the dining table inside the container home, illuminated by the candle whose flame bends with the impact. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Both of 현우's hands have struck its top) — The tabletop is seen at a shallow angle beneath his extended arms; used as Connects the two hand contacts and gives the impact a stable spatial base; Tabletop candle (Flame sharply bent at the instant of impact) — Seen side-on in the lower-left field, separate from the hand silhouettes; used as Provides a small, readable physical consequence of the impact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established candlelight retains its restrained warmth, with a momentary unevenness as the flame bends.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container dining area's table, interior surfaces, and candlelit nighttime atmosphere. Exclude the medicine bottle as a hand-held foreground element; do not transfer it from the earlier interaction.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The table candle remains the room's established light source, with the meal setting and opened medicine container still present. There is no established electrical lighting after curfew. 현우: Hyunwoo still has his upper garment off, with injured facial skin and a bleeding leg wound; no dressing has been applied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손이 식탁에 강하게 맞닿은 충격으로 촛불의 불꽃이 크게 꺾인 찰나, 양팔을 뻗은 채 굳어 있는 현우의 상체.\n\nLOCATION (lock): At the dining table inside the container home, illuminated by the candle whose flame bends with the impact. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Both of 현우's hands have struck its top) — The tabletop is seen at a shallow angle beneath his extended arms; used as Connects the two hand contacts and gives the impact a stable spatial base; Tabletop candle (Flame sharply bent at the instant of impact) — Seen side-on in the lower-left field, separate from the hand silhouettes; used as Provides a small, readable physical consequence of the impact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established candlelight retains its restrained warmth, with a momentary unevenness as the flame bends.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container dining area's table, interior surfaces, and candlelit nighttime atmosphere. Exclude the medicine bottle as a hand-held foreground element; do not transfer it from the earlier interaction.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The table candle remains the room's established light source, with the meal setting and opened medicine container still present. There is no established electrical lighting after curfew. 현우: Hyunwoo still has his upper garment off, with injured facial skin and a bleeding leg wound; no dressing has been applied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 식탁 아래쪽을 향하고 있으며, 곧게 뻗은 두 팔은 식탁 표면에 닿아 있다. 식탁 위 촛불의 불꽃은 충격으로 인해 왼쪽으로 크게 꺾여 있다.",
    "built_space": "컨테이너 내부의 식탁과 뒤편의 선반, 가전 등 레퍼런스의 공간 배경과 기물들이 올바른 위치에 배치되어 있다. 현우는 벤치에 올바르게 앉아 식탁과 상호작용하고 있다.",
    "entities": "현우는 캐릭터 레퍼런스와 일치하는 외모에 상의를 탈의하고 상처를 입은 상태이나, 이전 컷의 담요 대신 레퍼런스 이미지의 카고 바지를 입고 있다. 약병은 식탁 위에 놓여 있다.",
    "hard_violations": [],
    "physics": "현우의 엉덩이와 다리는 의자에 지지되어 있고 두 손은 단단한 식탁 표면에 평평하게 닿아 있다. 모든 사물이 중력에 맞게 지지면 위에 놓여 있다."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 양팔 사이의 식탁 표면을 향하고 있으며, 두 팔은 식탁을 강하게 짚고 있다. 촛불의 불꽃은 왼쪽으로 예리하게 꺾여 있다.",
    "built_space": "이전 컷과 동일한 식탁, 뒤편의 가구 및 조명 분위기가 올바른 위치와 비례로 묘사되어 있다.",
    "entities": "얼굴에 상처를 입고 상의를 탈의한 현우가 이전 컷에 등장한 담요를 덮고 피가 나는 다리 상처까지 일치하게 묘사되어 있으나, 지시된 약병은 화면에 보이지 않는다.",
    "hard_violations": [
     "[gemini-pro] 해부학적 오류: 화면 우측에 놓인 현우의 왼쪽 손에 엄지를 포함해 총 6개의 손가락이 그려져 있음",
     "[gemini-pro] 물리적 불가능성(재질 오류): 단단한 나무 재질의 식탁 표면이 양손이 닿은 지점에서 액체처럼 동심원 형태의 물결 파동을 일으키고 있음"
    ],
    "physics": "몸의 체중이 양손을 통해 식탁에 실려 있으나, 손이 닿은 지점의 나무 식탁 표면이 물리 법칙을 무시하고 액체의 파동처럼 왜곡되는 물리적 불가능성을 보인다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 컷의 담요 대신 바지를 입은 점과 촛불이 손 실루엣과 겹쳐진 구도는 다소 아쉬우나, 해부학적 및 물리적 치명적 오류 없이 지시된 충격의 순간을 안정적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "다리 상처와 담요 등 이전 컷의 묘사 연속성은 훌륭하나, 6개의 손가락과 단단한 나무 식탁이 물결치는 듯한 왜곡 등 심각한 해부학적/물리적 오류(Hard Violations)로 인해 탈락합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 식탁 아래쪽을 향하고 있으며, 곧게 뻗은 두 팔은 식탁 표면에 닿아 있다. 식탁 위 촛불의 불꽃은 충격으로 인해 왼쪽으로 크게 꺾여 있다.",
        "built_space": "컨테이너 내부의 식탁과 뒤편의 선반, 가전 등 레퍼런스의 공간 배경과 기물들이 올바른 위치에 배치되어 있다. 현우는 벤치에 올바르게 앉아 식탁과 상호작용하고 있다.",
        "entities": "현우는 캐릭터 레퍼런스와 일치하는 외모에 상의를 탈의하고 상처를 입은 상태이나, 이전 컷의 담요 대신 레퍼런스 이미지의 카고 바지를 입고 있다. 약병은 식탁 위에 놓여 있다.",
        "hard_violations": [],
        "physics": "현우의 엉덩이와 다리는 의자에 지지되어 있고 두 손은 단단한 식탁 표면에 평평하게 닿아 있다. 모든 사물이 중력에 맞게 지지면 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 양팔 사이의 식탁 표면을 향하고 있으며, 두 팔은 식탁을 강하게 짚고 있다. 촛불의 불꽃은 왼쪽으로 예리하게 꺾여 있다.",
        "built_space": "이전 컷과 동일한 식탁, 뒤편의 가구 및 조명 분위기가 올바른 위치와 비례로 묘사되어 있다.",
        "entities": "얼굴에 상처를 입고 상의를 탈의한 현우가 이전 컷에 등장한 담요를 덮고 피가 나는 다리 상처까지 일치하게 묘사되어 있으나, 지시된 약병은 화면에 보이지 않는다.",
        "hard_violations": [
         "해부학적 오류: 화면 우측에 놓인 현우의 왼쪽 손에 엄지를 포함해 총 6개의 손가락이 그려져 있음",
         "물리적 불가능성(재질 오류): 단단한 나무 재질의 식탁 표면이 양손이 닿은 지점에서 액체처럼 동심원 형태의 물결 파동을 일으키고 있음"
        ],
        "physics": "몸의 체중이 양손을 통해 식탁에 실려 있으나, 손이 닿은 지점의 나무 식탁 표면이 물리 법칙을 무시하고 액체의 파동처럼 왜곡되는 물리적 불가능성을 보인다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 컷의 담요 대신 바지를 입은 점과 촛불이 손 실루엣과 겹쳐진 구도는 다소 아쉬우나, 해부학적 및 물리적 치명적 오류 없이 지시된 충격의 순간을 안정적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "다리 상처와 담요 등 이전 컷의 묘사 연속성은 훌륭하나, 6개의 손가락과 단단한 나무 식탁이 물결치는 듯한 왜곡 등 심각한 해부학적/물리적 오류(Hard Violations)로 인해 탈락합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 식탁 아래쪽을 향하고 있으며, 곧게 뻗은 두 팔은 식탁 표면에 닿아 있다. 식탁 위 촛불의 불꽃은 충격으로 인해 왼쪽으로 크게 꺾여 있다.",
        "built_space": "컨테이너 내부의 식탁과 뒤편의 선반, 가전 등 레퍼런스의 공간 배경과 기물들이 올바른 위치에 배치되어 있다. 현우는 벤치에 올바르게 앉아 식탁과 상호작용하고 있다.",
        "entities": "현우는 캐릭터 레퍼런스와 일치하는 외모에 상의를 탈의하고 상처를 입은 상태이나, 이전 컷의 담요 대신 레퍼런스 이미지의 카고 바지를 입고 있다. 약병은 식탁 위에 놓여 있다.",
        "hard_violations": [],
        "physics": "현우의 엉덩이와 다리는 의자에 지지되어 있고 두 손은 단단한 식탁 표면에 평평하게 닿아 있다. 모든 사물이 중력에 맞게 지지면 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 양팔 사이의 식탁 표면을 향하고 있으며, 두 팔은 식탁을 강하게 짚고 있다. 촛불의 불꽃은 왼쪽으로 예리하게 꺾여 있다.",
        "built_space": "이전 컷과 동일한 식탁, 뒤편의 가구 및 조명 분위기가 올바른 위치와 비례로 묘사되어 있다.",
        "entities": "얼굴에 상처를 입고 상의를 탈의한 현우가 이전 컷에 등장한 담요를 덮고 피가 나는 다리 상처까지 일치하게 묘사되어 있으나, 지시된 약병은 화면에 보이지 않는다.",
        "hard_violations": [
         "해부학적 오류: 화면 우측에 놓인 현우의 왼쪽 손에 엄지를 포함해 총 6개의 손가락이 그려져 있음",
         "물리적 불가능성(재질 오류): 단단한 나무 재질의 식탁 표면이 양손이 닿은 지점에서 액체처럼 동심원 형태의 물결 파동을 일으키고 있음"
        ],
        "physics": "몸의 체중이 양손을 통해 식탁에 실려 있으나, 손이 닿은 지점의 나무 식탁 표면이 물리 법칙을 무시하고 액체의 파동처럼 왜곡되는 물리적 불가능성을 보인다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "양손의 강한 접촉과 꺾인 불꽃은 명확하지만, 무릎과 실내까지 넓게 담아 상체 중심 미디엄 숏에서 멀어지고 펼친 상판에 열린 약병도 보이지 않는다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "상체 중심 구도와 뻗은 양팔의 정지 자세, 크게 꺾인 불꽃이 지시된 순간에 더 가깝지만 초가 한쪽 손과 겹쳐 손 윤곽과의 분리 조건은 미흡하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 렌즈가 아니라 화면 왼쪽 앞의 식탁 너머를 내려다본다. 두 팔은 서로 떨어진 상판 접촉점으로 뻗고 두 손바닥은 아래를 향한다. 불꽃은 심지에서 화면 왼쪽으로 크게 휘며, 불꽃과 양손의 윤곽은 구분된다.",
        "built_space": "나무 식탁 하나와 오른쪽 벽을 따른 긴 나무 좌석 하나가 보인다. 뒤쪽에는 냉장고 하나, 전자레인지 하나가 놓인 금속 선반 한 조, 오른쪽 벽의 전기함 하나와 수직 배관이 있다. 패널 벽과 목재의 재질은 이전 장면과 대체로 이어진다. 현우는 좌석 앞에서 식탁 쪽으로 기울어 있고, 상판이 두 손의 접촉점을 연결한다. 다만 무릎과 실내 깊숙한 공간까지 보여 지정된 상체 중심 구도보다 넓다.",
        "entities": "인물은 현우로 보이는 젊은 동아시아계 남성 한 명뿐이며, 검은 헝클어진 머리와 얼굴 형태는 참조에 대체로 부합한다. 한국계 미국인이라는 국적 배경은 외형만으로 확인할 수 없다. 상의를 벗었고 볼의 상처, 무릎의 출혈, 허리와 허벅지의 체크 천이 보인다. 초와 받침 각각 하나, 컵 하나, 식사용 그릇 두 개와 젓가락이 있다. 열린 약병은 보이지 않는다. 판독 가능한 글자나 다른 인물은 없다.",
        "hard_violations": [],
        "physics": "두 손바닥이 상판에 밀착되어 앞으로 기울인 상체를 받치고, 골반은 뒤쪽 좌석 높이에 놓여 있어 지지 관계가 자연스럽다. 초와 식기는 상판에 놓여 있고 초는 받침에 서 있다. 뻗은 팔과 눌린 손은 식탁을 내려친 직후의 자세로 가능하다. 불꽃의 옆 방향 굽힘도 순간적인 공기 움직임으로 설명할 수 있으며, 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 왼쪽의 식탁 너머를 향하고 렌즈를 보지 않는다. 양팔은 앞쪽 상판으로 곧게 뻗으며 두 손바닥이 아래로 닿는다. 불꽃은 화면 왼쪽으로 급하게 꺾여 충격의 결과가 잘 읽힌다. 다만 초 몸통과 받침이 화면 왼쪽 손 일부를 가려, 초와 손 윤곽을 분리하라는 조건에는 덜 맞는다.",
        "built_space": "식탁 하나와 오른쪽 벽을 따른 긴 나무 좌석 하나가 있으며 현우는 그 좌석에 앉아 식탁을 향한다. 냉장고 하나, 금속 선반 한 조와 그 안의 전자레인지 하나, 오른쪽 전기함 하나와 수직 배관이 보인다. 패널 벽, 벽 사진, 선반과 냉장고의 관계가 이전 장면의 공간을 대체로 유지한다. 상판은 두 손 아래에서 낮은 사선으로 펼쳐지고, 얼굴과 상체 및 양팔이 중심이 되어 A보다 미디엄 숏 지시에 가깝다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 등장한다. 앳된 얼굴, 헝클어진 검은 머리, 체격과 볼의 상처는 참조와 대체로 일치하며 상의도 벗고 있다. 국적 배경은 영상만으로 판별할 수 없다. 다리 상처는 프레임 밖이므로 확인 대상이 아니다. 초와 받침 각각 하나, 컵 하나, 식사용 그릇 세 개와 젓가락, 뚜껑이 열린 갈색 약병 하나가 상판에 있다. 약병은 손에 들려 있지 않고 라벨도 읽을 수 없다.",
        "hard_violations": [],
        "physics": "골반은 나무 좌석에 놓이고 두 손은 상판에 평평하게 닿아 몸을 지지한다. 양팔을 편 채 어깨와 몸통이 멈춘 자세는 양손으로 식탁을 친 직후로 성립한다. 초는 받침 위에, 약병과 식기는 상판 위에 안정적으로 놓여 있다. 불꽃의 강한 굽힘은 순간적인 공기 흐름으로 가능한 형태이며, 지지 없이 떠 있는 물체나 인체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "양손의 강한 접촉과 꺾인 불꽃은 명확하지만, 무릎과 실내까지 넓게 담아 상체 중심 미디엄 숏에서 멀어지고 펼친 상판에 열린 약병도 보이지 않는다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "상체 중심 구도와 뻗은 양팔의 정지 자세, 크게 꺾인 불꽃이 지시된 순간에 더 가깝지만 초가 한쪽 손과 겹쳐 손 윤곽과의 분리 조건은 미흡하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 렌즈가 아니라 화면 왼쪽 앞의 식탁 너머를 내려다본다. 두 팔은 서로 떨어진 상판 접촉점으로 뻗고 두 손바닥은 아래를 향한다. 불꽃은 심지에서 화면 왼쪽으로 크게 휘며, 불꽃과 양손의 윤곽은 구분된다.",
        "built_space": "나무 식탁 하나와 오른쪽 벽을 따른 긴 나무 좌석 하나가 보인다. 뒤쪽에는 냉장고 하나, 전자레인지 하나가 놓인 금속 선반 한 조, 오른쪽 벽의 전기함 하나와 수직 배관이 있다. 패널 벽과 목재의 재질은 이전 장면과 대체로 이어진다. 현우는 좌석 앞에서 식탁 쪽으로 기울어 있고, 상판이 두 손의 접촉점을 연결한다. 다만 무릎과 실내 깊숙한 공간까지 보여 지정된 상체 중심 구도보다 넓다.",
        "entities": "인물은 현우로 보이는 젊은 동아시아계 남성 한 명뿐이며, 검은 헝클어진 머리와 얼굴 형태는 참조에 대체로 부합한다. 한국계 미국인이라는 국적 배경은 외형만으로 확인할 수 없다. 상의를 벗었고 볼의 상처, 무릎의 출혈, 허리와 허벅지의 체크 천이 보인다. 초와 받침 각각 하나, 컵 하나, 식사용 그릇 두 개와 젓가락이 있다. 열린 약병은 보이지 않는다. 판독 가능한 글자나 다른 인물은 없다.",
        "hard_violations": [],
        "physics": "두 손바닥이 상판에 밀착되어 앞으로 기울인 상체를 받치고, 골반은 뒤쪽 좌석 높이에 놓여 있어 지지 관계가 자연스럽다. 초와 식기는 상판에 놓여 있고 초는 받침에 서 있다. 뻗은 팔과 눌린 손은 식탁을 내려친 직후의 자세로 가능하다. 불꽃의 옆 방향 굽힘도 순간적인 공기 움직임으로 설명할 수 있으며, 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 시선은 화면 왼쪽의 식탁 너머를 향하고 렌즈를 보지 않는다. 양팔은 앞쪽 상판으로 곧게 뻗으며 두 손바닥이 아래로 닿는다. 불꽃은 화면 왼쪽으로 급하게 꺾여 충격의 결과가 잘 읽힌다. 다만 초 몸통과 받침이 화면 왼쪽 손 일부를 가려, 초와 손 윤곽을 분리하라는 조건에는 덜 맞는다.",
        "built_space": "식탁 하나와 오른쪽 벽을 따른 긴 나무 좌석 하나가 있으며 현우는 그 좌석에 앉아 식탁을 향한다. 냉장고 하나, 금속 선반 한 조와 그 안의 전자레인지 하나, 오른쪽 전기함 하나와 수직 배관이 보인다. 패널 벽, 벽 사진, 선반과 냉장고의 관계가 이전 장면의 공간을 대체로 유지한다. 상판은 두 손 아래에서 낮은 사선으로 펼쳐지고, 얼굴과 상체 및 양팔이 중심이 되어 A보다 미디엄 숏 지시에 가깝다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 등장한다. 앳된 얼굴, 헝클어진 검은 머리, 체격과 볼의 상처는 참조와 대체로 일치하며 상의도 벗고 있다. 국적 배경은 영상만으로 판별할 수 없다. 다리 상처는 프레임 밖이므로 확인 대상이 아니다. 초와 받침 각각 하나, 컵 하나, 식사용 그릇 세 개와 젓가락, 뚜껑이 열린 갈색 약병 하나가 상판에 있다. 약병은 손에 들려 있지 않고 라벨도 읽을 수 없다.",
        "hard_violations": [],
        "physics": "골반은 나무 좌석에 놓이고 두 손은 상판에 평평하게 닿아 몸을 지지한다. 양팔을 편 채 어깨와 몸통이 멈춘 자세는 양손으로 식탁을 친 직후로 성립한다. 초는 받침 위에, 약병과 식기는 상판 위에 안정적으로 놓여 있다. 불꽃의 강한 굽힘은 순간적인 공기 흐름으로 가능한 형태이며, 지지 없이 떠 있는 물체나 인체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.161
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.911
   },
   "violations": {
    "B": [
     "[gemini-pro] 해부학적 오류: 화면 우측에 놓인 현우의 왼쪽 손에 엄지를 포함해 총 6개의 손가락이 그려져 있음",
     "[gemini-pro] 물리적 불가능성(재질 오류): 단단한 나무 재질의 식탁 표면이 양손이 닿은 지점에서 액체처럼 동심원 형태의 물결 파동을 일으키고 있음"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 911
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "이전 컷의 담요 대신 바지를 입은 점과 촛불이 손 실루엣과 겹쳐진 구도는 다소 아쉬우나, 해부학적 및 물리적 치명적 오류 없이 지시된 충격의 순간을 안정적으로 구현했습니다."
   },
   {
    "label": "B",
    "score": 911,
    "verdict_ko": "다리 상처와 담요 등 이전 컷의 묘사 연속성은 훌륭하나, 6개의 손가락과 단단한 나무 식탁이 물결치는 듯한 왜곡 등 심각한 해부학적/물리적 오류(Hard Violations)로 인해 탈락합니다.  ★위반: [gemini-pro] 해부학적 오류: 화면 우측에 놓인 현우의 왼쪽 손에 엄지를 포함해 총 6개의 손가락이 그려져 있음 / [gemini-pro] 물리적 불가능성(재질 오류): 단단한 나무 재질의 식탁 표면이 양손이 닿은 지점에서 액체처럼 동심원 형태의 물결 파동을 일으키고 있음"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7_sel.png",
    "asset_id": "74f66d35-f2a4-400b-bf04-62ab61759576",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-c1fb-704e-947d-ec31d84f68fd",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S12sh7"
  }
 },
 "S12sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:24:22.327076+00:00",
  "fingerprint": "e904725753b9a76c3dd4a532f6848e4e3b3d29b2bf950dc2e86b888ca0cd3459",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S12sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S12sh8_sel.png",
  "source_sha256": "4923656aa959d2ee732c85d1ac7f7683e5ba4c10f56ab43cac8dab7bf676740c",
  "file": "S12sh8_cine.png",
  "staged_sha256": "34281a5f257c025a3f5d4bacf7cda3b9175206e0078ee4a5eae589c66a8ebe33",
  "latency_ms": 10439
 },
 "S12sh12::signage": {
  "fp": "2695a1071ee832be",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S12sh12": {
  "input_fingerprint": "5bb8f91a0696e65d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 식탁 앞에 홀로 앉아 굳은 표정으로 닫힌 문 쪽을 응시하는 미연의 상체.\n\nLOCATION (lock): At the candlelit dining table inside the container home, facing the closed front door. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Unoccupied dining-table space (No other person remains at the table with 미연) — The tabletop extends from beside her toward the lower-right field; used as Makes the absent relationship visible through empty space rather than an additional figure.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established intimate candlelight and subdued surrounding contrast without adding a new source after 현우's departure.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same dining table, candle, and surrounding container interior under nighttime candlelight. Exclude the young man, the girl, and the momentarily bent candle flame from the reference.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container's exterior door is now shut after being slammed. The table candle remains lit, with the meal setting and opened medicine container not yet cleared. 미연: Miyeon remains at the dining table after the argument.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 식탁 앞에 홀로 앉아 굳은 표정으로 닫힌 문 쪽을 응시하는 미연의 상체.\n\nLOCATION (lock): At the candlelit dining table inside the container home, facing the closed front door. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Unoccupied dining-table space (No other person remains at the table with 미연) — The tabletop extends from beside her toward the lower-right field; used as Makes the absent relationship visible through empty space rather than an additional figure.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established intimate candlelight and subdued surrounding contrast without adding a new source after 현우's departure.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same dining table, candle, and surrounding container interior under nighttime candlelight. Exclude the young man, the girl, and the momentarily bent candle flame from the reference.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container's exterior door is now shut after being slammed. The table candle remains lit, with the meal setting and opened medicine container not yet cleared. 미연: Miyeon remains at the dining table after the argument.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 식탁 앞에 홀로 앉아 굳은 표정으로 닫힌 문 쪽을 응시하는 미연의 상체.\n\nLOCATION (lock): At the candlelit dining table inside the container home, facing the closed front door. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Unoccupied dining-table space (No other person remains at the table with 미연) — The tabletop extends from beside her toward the lower-right field; used as Makes the absent relationship visible through empty space rather than an additional figure.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established intimate candlelight and subdued surrounding contrast without adding a new source after 현우's departure.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same dining table, candle, and surrounding container interior under nighttime candlelight. Exclude the young man, the girl, and the momentarily bent candle flame from the reference.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container's exterior door is now shut after being slammed. The table candle remains lit, with the meal setting and opened medicine container not yet cleared. 미연: Miyeon remains at the dining table after the argument.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "미연이 닫힌 문이 아닌 카메라 정면을 응시하고 있음.",
    "built_space": "컨테이너 내부 구조와 가구 배치는 이전 샷과 유사하나, 식탁과 문의 배치가 인물의 시선 방향과 조화를 이루지 못함.",
    "entities": "미연의 외모와 의상은 캐릭터 레퍼런스와 일치하나, 식탁 우측에 이전 샷에 없던 유리컵이 나타남.",
    "hard_violations": [
     "[gemini-pro] 없는 사물 추가 (식탁 위 유리컵)",
     "[gpt-high] 이전 장면의 식탁에 없고 프롬프트에서도 추가를 지시하지 않은 물 든 유리잔을 식탁 오른쪽에 새로 배치했다."
    ],
    "physics": "식탁에 팔을 올린 채 안정적으로 앉아 있으며 물리적인 오류는 없음."
   },
   {
    "label": "B",
    "direction": "미연의 시선이 프롬프트 지시대로 화면 좌측에 위치한 닫힌 문을 정확히 향하고 있음.",
    "built_space": "좌측의 문과 우측으로 확장되는 빈 식탁 및 벤치 공간 등 이전 샷의 공간적 디테일을 그대로 유지함.",
    "entities": "인물의 외모와 의상, 그리고 식탁 위 촛불, 약통, 그릇 등의 소품들이 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "의자에 앉아 손을 무릎 쪽에 두고 자연스러운 무게 중심과 자세를 보여줌."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "닫힌 문을 응시하는 시선과 우측 빈 식탁 공간을 정확히 구현하여 프롬프트의 의도를 완벽히 살림."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "시선이 닫힌 문이 아닌 카메라를 향하고 있으며, 이전 샷에 없던 유리컵이 식탁에 추가됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연이 닫힌 문이 아닌 카메라 정면을 응시하고 있음.",
        "built_space": "컨테이너 내부 구조와 가구 배치는 이전 샷과 유사하나, 식탁과 문의 배치가 인물의 시선 방향과 조화를 이루지 못함.",
        "entities": "미연의 외모와 의상은 캐릭터 레퍼런스와 일치하나, 식탁 우측에 이전 샷에 없던 유리컵이 나타남.",
        "hard_violations": [
         "없는 사물 추가 (식탁 위 유리컵)"
        ],
        "physics": "식탁에 팔을 올린 채 안정적으로 앉아 있으며 물리적인 오류는 없음."
       },
       {
        "label": "B",
        "direction": "미연의 시선이 프롬프트 지시대로 화면 좌측에 위치한 닫힌 문을 정확히 향하고 있음.",
        "built_space": "좌측의 문과 우측으로 확장되는 빈 식탁 및 벤치 공간 등 이전 샷의 공간적 디테일을 그대로 유지함.",
        "entities": "인물의 외모와 의상, 그리고 식탁 위 촛불, 약통, 그릇 등의 소품들이 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "의자에 앉아 손을 무릎 쪽에 두고 자연스러운 무게 중심과 자세를 보여줌."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "닫힌 문을 응시하는 시선과 우측 빈 식탁 공간을 정확히 구현하여 프롬프트의 의도를 완벽히 살림."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "시선이 닫힌 문이 아닌 카메라를 향하고 있으며, 이전 샷에 없던 유리컵이 식탁에 추가됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연이 닫힌 문이 아닌 카메라 정면을 응시하고 있음.",
        "built_space": "컨테이너 내부 구조와 가구 배치는 이전 샷과 유사하나, 식탁과 문의 배치가 인물의 시선 방향과 조화를 이루지 못함.",
        "entities": "미연의 외모와 의상은 캐릭터 레퍼런스와 일치하나, 식탁 우측에 이전 샷에 없던 유리컵이 나타남.",
        "hard_violations": [
         "없는 사물 추가 (식탁 위 유리컵)"
        ],
        "physics": "식탁에 팔을 올린 채 안정적으로 앉아 있으며 물리적인 오류는 없음."
       },
       {
        "label": "B",
        "direction": "미연의 시선이 프롬프트 지시대로 화면 좌측에 위치한 닫힌 문을 정확히 향하고 있음.",
        "built_space": "좌측의 문과 우측으로 확장되는 빈 식탁 및 벤치 공간 등 이전 샷의 공간적 디테일을 그대로 유지함.",
        "entities": "인물의 외모와 의상, 그리고 식탁 위 촛불, 약통, 그릇 등의 소품들이 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "의자에 앉아 손을 무릎 쪽에 두고 자연스러운 무게 중심과 자세를 보여줌."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "상체 중심 미디엄 숏과 오른쪽 아래로 펼쳐지는 빈 식탁은 충실하지만, 미연의 시선이 등 뒤의 닫힌 문이 아닌 화면 오른쪽을 향한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "미연이 닫힌 문 대신 카메라 쪽을 바라보고, 이전 식탁에 없던 물잔까지 추가되어 장면의 행동과 소품 연속성을 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 얼굴과 눈을 화면 오른쪽의 프레임 밖으로 향한다. 손잡이가 보이는 닫힌 문은 미연의 왼쪽 뒤에 있어 시선이 그 문에 닿지 않는다. 굳게 다문 입과 긴장된 표정은 요구에 맞는다.",
        "built_space": "나무 식탁 하나가 미연 옆에서 오른쪽 아래로 뻗고, 미연 쪽 의자 등받이 하나와 맞은편 빈 등받이 하나가 보인다. 왼쪽에 닫힌 문 하나, 뒤쪽에 냉장고 하나와 벽걸이 에어컨 하나, 오른쪽에 전자레인지가 놓인 금속 선반 하나가 있다. 패널 벽, 배관, 사진과 식물은 기존 컨테이너 실내를 대체로 유지한다. 다만 문이 미연 뒤에 배치되어 문을 응시하는 동선과 맞지 않는다.",
        "entities": "중년 한국인 여성으로 읽히는 인물 한 명만 있으며 검은 머리, 얼굴 윤곽, 체크 셔츠와 갈색 앞치마는 미연 참조와 대체로 맞는다. 참조의 모자는 없다. 식탁에는 켜진 초 하나와 받침, 열린 갈색 약병 하나, 도자기 컵 하나, 밥그릇과 다른 그릇 및 식기가 남아 있다. 이전 인물은 없고 읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연은 등받이가 뒤에 있는 의자에 앉아 있으며 몸의 무게가 좌석으로 내려가는 자세다. 팔은 무릎 쪽으로 내려가고 손 부위는 일부 가려져 있다. 그릇과 약병, 컵, 촛대는 모두 식탁 위에 놓여 있고 초는 받침에 지지된다. 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "미연의 눈은 거의 카메라를 직접 향한다. 손잡이가 달린 닫힌 문은 왼쪽 뒤에 있으므로 눈길의 목표가 문이 아니다. 손은 식탁 위에서 서로 맞잡고 있으며 별도의 지시 동작은 없다.",
        "built_space": "나무 식탁 하나가 오른쪽 아래를 차지하고 왼쪽 아래에 의자 등받이 하나가 보인다. 왼쪽 뒤에 창이 있는 닫힌 문 하나, 후방에 에어컨 하나와 냉장고 하나, 오른쪽에 전자레인지가 있는 금속 선반 하나가 있다. 기존의 좁고 긴 패널 실내와 재료는 대체로 유지하지만, 문이 인물 뒤에 있어 문을 바라보는 장면이 되지 않는다.",
        "entities": "중년 한국인 여성 한 명의 검은 머리와 체크 셔츠, 갈색 앞치마는 참조에 대체로 부합하며 모자는 없다. 초 하나, 받침, 열린 약병, 도자기 컵, 식사 그릇과 식기가 보인다. 식탁 오른쪽에는 이전 장면에 없던 물이 든 투명 유리잔 하나가 추가되어 있다. 다른 사람이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [
         "이전 장면의 식탁에 없고 프롬프트에서도 추가를 지시하지 않은 물 든 유리잔을 식탁 오른쪽에 새로 배치했다."
        ],
        "physics": "미연의 하체는 식탁 옆 좌석에 놓이고 상체는 앞으로 조금 기울어 있다. 양팔이 식탁에 닿아 맞잡은 손을 지지하므로 자세는 물리적으로 가능하다. 약병과 식기, 유리잔, 촛대는 모두 식탁 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "상체 중심 미디엄 숏과 오른쪽 아래로 펼쳐지는 빈 식탁은 충실하지만, 미연의 시선이 등 뒤의 닫힌 문이 아닌 화면 오른쪽을 향한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미연이 닫힌 문 대신 카메라 쪽을 바라보고, 이전 식탁에 없던 물잔까지 추가되어 장면의 행동과 소품 연속성을 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 얼굴과 눈을 화면 오른쪽의 프레임 밖으로 향한다. 손잡이가 보이는 닫힌 문은 미연의 왼쪽 뒤에 있어 시선이 그 문에 닿지 않는다. 굳게 다문 입과 긴장된 표정은 요구에 맞는다.",
        "built_space": "나무 식탁 하나가 미연 옆에서 오른쪽 아래로 뻗고, 미연 쪽 의자 등받이 하나와 맞은편 빈 등받이 하나가 보인다. 왼쪽에 닫힌 문 하나, 뒤쪽에 냉장고 하나와 벽걸이 에어컨 하나, 오른쪽에 전자레인지가 놓인 금속 선반 하나가 있다. 패널 벽, 배관, 사진과 식물은 기존 컨테이너 실내를 대체로 유지한다. 다만 문이 미연 뒤에 배치되어 문을 응시하는 동선과 맞지 않는다.",
        "entities": "중년 한국인 여성으로 읽히는 인물 한 명만 있으며 검은 머리, 얼굴 윤곽, 체크 셔츠와 갈색 앞치마는 미연 참조와 대체로 맞는다. 참조의 모자는 없다. 식탁에는 켜진 초 하나와 받침, 열린 갈색 약병 하나, 도자기 컵 하나, 밥그릇과 다른 그릇 및 식기가 남아 있다. 이전 인물은 없고 읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연은 등받이가 뒤에 있는 의자에 앉아 있으며 몸의 무게가 좌석으로 내려가는 자세다. 팔은 무릎 쪽으로 내려가고 손 부위는 일부 가려져 있다. 그릇과 약병, 컵, 촛대는 모두 식탁 위에 놓여 있고 초는 받침에 지지된다. 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "미연의 눈은 거의 카메라를 직접 향한다. 손잡이가 달린 닫힌 문은 왼쪽 뒤에 있으므로 눈길의 목표가 문이 아니다. 손은 식탁 위에서 서로 맞잡고 있으며 별도의 지시 동작은 없다.",
        "built_space": "나무 식탁 하나가 오른쪽 아래를 차지하고 왼쪽 아래에 의자 등받이 하나가 보인다. 왼쪽 뒤에 창이 있는 닫힌 문 하나, 후방에 에어컨 하나와 냉장고 하나, 오른쪽에 전자레인지가 있는 금속 선반 하나가 있다. 기존의 좁고 긴 패널 실내와 재료는 대체로 유지하지만, 문이 인물 뒤에 있어 문을 바라보는 장면이 되지 않는다.",
        "entities": "중년 한국인 여성 한 명의 검은 머리와 체크 셔츠, 갈색 앞치마는 참조에 대체로 부합하며 모자는 없다. 초 하나, 받침, 열린 약병, 도자기 컵, 식사 그릇과 식기가 보인다. 식탁 오른쪽에는 이전 장면에 없던 물이 든 투명 유리잔 하나가 추가되어 있다. 다른 사람이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [
         "이전 장면의 식탁에 없고 프롬프트에서도 추가를 지시하지 않은 물 든 유리잔을 식탁 오른쪽에 새로 배치했다."
        ],
        "physics": "미연의 하체는 식탁 옆 좌석에 놓이고 상체는 앞으로 조금 기울어 있다. 양팔이 식탁에 닿아 맞잡은 손을 지지하므로 자세는 물리적으로 가능하다. 약병과 식기, 유리잔, 촛대는 모두 식탁 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.975,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.725,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 없는 사물 추가 (식탁 위 유리컵)",
     "[gpt-high] 이전 장면의 식탁에 없고 프롬프트에서도 추가를 지시하지 않은 물 든 유리잔을 식탁 오른쪽에 새로 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 725
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "닫힌 문을 응시하는 시선과 우측 빈 식탁 공간을 정확히 구현하여 프롬프트의 의도를 완벽히 살림."
   },
   {
    "label": "A",
    "score": 725,
    "verdict_ko": "시선이 닫힌 문이 아닌 카메라를 향하고 있으며, 이전 샷에 없던 유리컵이 식탁에 추가됨.  ★위반: [gemini-pro] 없는 사물 추가 (식탁 위 유리컵) / [gpt-high] 이전 장면의 식탁에 없고 프롬프트에서도 추가를 지시하지 않은 물 든 유리잔을 식탁 오른쪽에 새로 배치했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh8_sel.png",
    "asset_id": "a3b93bed-3edf-4f19-99e1-875d84c449e6",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-c3ba-7f31-a793-0a19fa3ea365",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S12sh8"
  }
 },
 "S12sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:25:25.226852+00:00",
  "fingerprint": "f70e3aa86cbde83dfbfca6a6632d1bd79560cbcea9a37f840f3c66ce1e4458a5",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S12sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S12sh12_sel.png",
  "source_sha256": "277f2241042132f31948e508a99d1c5f735bfc9d9bc6bc64344012e6688df1eb",
  "file": "S12sh12_cine.png",
  "staged_sha256": "feff6c6fb0a3ea19a36848efbb216fdda7b35e0c6809b7c2c72d2b8ee014fa79",
  "latency_ms": 11579
 },
 "S13sh10::signage": {
  "fp": "d23f70edea146926",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::roadside_hideout": {
  "input_fingerprint": "84a69df1b2a9c646",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "roadside_hideout",
    "tags": [
     "S13sh10"
    ]
   },
   "context_sig": "6e556fbde0803929"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 인공제방 바로 옆 노상에 쿠마의 아지트가 있다.\n- 쿠마, 현우를 아랑곳하지 않고 인공제방 쪽으로 성큼성큼 앞서 걷는다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 인공제방 바로 옆 노상에 쿠마의 아지트가 있다.\n- 쿠마, 현우를 아랑곳하지 않고 인공제방 쪽으로 성큼성큼 앞서 걷는다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_roadside_hideout_f5126d.png",
  "asset_id": "3da48160-9267-4344-8758-06815d5af5e3",
  "input_asset_ids": [
   "a2b7d59b-6427-4f49-8773-4ed140b22cee"
  ],
  "origin_tag": "S13sh10",
  "place_text": "In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.",
  "origin_inputs": {
   "place_text": "In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.",
   "time_of_day_en": "night",
   "conti_asset_id": "a2b7d59b-6427-4f49-8773-4ed140b22cee"
  }
 },
 "S13sh10::bgfirst_bg": {
  "input_fingerprint": "cc716612c2959c56",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh10__bgfirst_bg.png",
  "asset_id": "51a6f4c4-1510-44cd-986d-c3a6e8933313",
  "input_asset_ids": [
   "a2b7d59b-6427-4f49-8773-4ed140b22cee",
   "3da48160-9267-4344-8758-06815d5af5e3"
  ]
 },
 "S13sh10": {
  "input_fingerprint": "351d4059c1591ca3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A fire burns in a metal drum at the gang's roadside hideout, and chicken skewers are on the barbecue grill. The adjoining seawall remains cracked and leaking. 현우: Hyunwoo wears his upper garment and holds a seized gun in an aiming posture; his facial injuries and dog-bitten leg remain untreated. 쿠마: Kuma remains at the barbecue gathering with a chicken skewer in hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A fire burns in a metal drum at the gang's roadside hideout, and chicken skewers are on the barbecue grill. The adjoining seawall remains cracked and leaking. 현우: Hyunwoo wears his upper garment and holds a seized gun in an aiming posture; his facial injuries and dog-bitten leg remain untreated. 쿠마: Kuma remains at the barbecue gathering with a chicken skewer in hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A fire burns in a metal drum at the gang's roadside hideout, and chicken skewers are on the barbecue grill. The adjoining seawall remains cracked and leaking. 현우: Hyunwoo wears his upper garment and holds a seized gun in an aiming posture; his facial injuries and dog-bitten leg remain untreated. 쿠마: Kuma remains at the barbecue gathering with a chicken skewer in hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh10__bgfirst_bg.png",
     "asset_id": "51a6f4c4-1510-44cd-986d-c3a6e8933313",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S13sh10.png",
     "asset_id": "a2b7d59b-6427-4f49-8773-4ed140b22cee",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:806945>",
     "asset_id": "fe00e8c2-fc45-4464-b8b5-95ce8056561f",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_roadside_hideout_f5126d.png",
     "asset_id": "3da48160-9267-4344-8758-06815d5af5e3",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:806945>",
     "asset_id": "fe00e8c2-fc45-4464-b8b5-95ce8056561f",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우가 든 총의 총구는 상대방의 이마를 향하고 있음.",
    "built_space": "방파제 곁의 야외 공간으로, 불타는 드럼통과 바비큐 그릴이 배경에 존재함.",
    "entities": "현우는 참고 이미지와 일치하나, 표적이 된 인물은 쿠마가 아닌 현우 본인과 닮은 다른 인물임. 권총은 개머리판이 생략되어 소품 참고 이미지와 일치하지 않음.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물 2명이 배경에 추가됨 (extra bodies)",
     "[gemini-pro] 표적 인물이 쿠마가 아님 (invented people)",
     "[gpt-high] 샷에 허용되지 않은 배경 인물 두 명을 추가했다.",
     "[gpt-high] 참조 장소의 단일 그릴을 두 개로 중복했다.",
     "[gpt-high] 참조 장소의 단일 금속 드럼통을 두 개로 중복했다."
    ],
    "physics": "인물들은 바닥을 딛고 서 있으며, 무기와 꼬치구이는 각각 손으로 쥐어 지지되고 있음."
   },
   {
    "label": "B",
    "direction": "현우가 쥔 권총의 총구가 쿠마의 이마 정중앙에 정확하게 맞닿아 있음.",
    "built_space": "방파제 근처의 야외로, 배경에 화로와 그릴이 올바르게 배치됨.",
    "entities": "현우와 쿠마 모두 참고 이미지의 얼굴, 헤어스타일, 의상과 정확히 일치함. 권총 역시 개머리판이 장착된 형태까지 소품 레퍼런스와 동일함(단, 총기의 영문 각인이 읽히는 상태로 남음).",
    "hard_violations": [
     "[gpt-high] 권총 측면에 읽을 수 있는 문자·숫자 각인이 노출되어, 화면 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "현우의 오른손이 총을 안정적으로 쥐고 있으며, 화면 하단 끄트머리에서 쿠마의 손가락이 꼬치구이를 쥐고 지지함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 클로즈업 샷 구도를 무시하고 화면을 넓혔으며, 쿠마의 외형이 불일치하고 배경에 지시되지 않은 인물이 등장하는 치명적인 위반이 발생했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물의 옆얼굴을 총구가 이어주는 정확한 클로즈업 구도를 구현했으며, 인물들의 외모와 무기 소품의 디테일을 프롬프트에 맞게 훌륭히 반영했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우가 든 총의 총구는 상대방의 이마를 향하고 있음.",
        "built_space": "방파제 곁의 야외 공간으로, 불타는 드럼통과 바비큐 그릴이 배경에 존재함.",
        "entities": "현우는 참고 이미지와 일치하나, 표적이 된 인물은 쿠마가 아닌 현우 본인과 닮은 다른 인물임. 권총은 개머리판이 생략되어 소품 참고 이미지와 일치하지 않음.",
        "hard_violations": [
         "프롬프트에 없는 인물 2명이 배경에 추가됨 (extra bodies)",
         "표적 인물이 쿠마가 아님 (invented people)"
        ],
        "physics": "인물들은 바닥을 딛고 서 있으며, 무기와 꼬치구이는 각각 손으로 쥐어 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "현우가 쥔 권총의 총구가 쿠마의 이마 정중앙에 정확하게 맞닿아 있음.",
        "built_space": "방파제 근처의 야외로, 배경에 화로와 그릴이 올바르게 배치됨.",
        "entities": "현우와 쿠마 모두 참고 이미지의 얼굴, 헤어스타일, 의상과 정확히 일치함. 권총 역시 개머리판이 장착된 형태까지 소품 레퍼런스와 동일함(단, 총기의 영문 각인이 읽히는 상태로 남음).",
        "hard_violations": [],
        "physics": "현우의 오른손이 총을 안정적으로 쥐고 있으며, 화면 하단 끄트머리에서 쿠마의 손가락이 꼬치구이를 쥐고 지지함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 클로즈업 샷 구도를 무시하고 화면을 넓혔으며, 쿠마의 외형이 불일치하고 배경에 지시되지 않은 인물이 등장하는 치명적인 위반이 발생했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물의 옆얼굴을 총구가 이어주는 정확한 클로즈업 구도를 구현했으며, 인물들의 외모와 무기 소품의 디테일을 프롬프트에 맞게 훌륭히 반영했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우가 든 총의 총구는 상대방의 이마를 향하고 있음.",
        "built_space": "방파제 곁의 야외 공간으로, 불타는 드럼통과 바비큐 그릴이 배경에 존재함.",
        "entities": "현우는 참고 이미지와 일치하나, 표적이 된 인물은 쿠마가 아닌 현우 본인과 닮은 다른 인물임. 권총은 개머리판이 생략되어 소품 참고 이미지와 일치하지 않음.",
        "hard_violations": [
         "프롬프트에 없는 인물 2명이 배경에 추가됨 (extra bodies)",
         "표적 인물이 쿠마가 아님 (invented people)"
        ],
        "physics": "인물들은 바닥을 딛고 서 있으며, 무기와 꼬치구이는 각각 손으로 쥐어 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "현우가 쥔 권총의 총구가 쿠마의 이마 정중앙에 정확하게 맞닿아 있음.",
        "built_space": "방파제 근처의 야외로, 배경에 화로와 그릴이 올바르게 배치됨.",
        "entities": "현우와 쿠마 모두 참고 이미지의 얼굴, 헤어스타일, 의상과 정확히 일치함. 권총 역시 개머리판이 장착된 형태까지 소품 레퍼런스와 동일함(단, 총기의 영문 각인이 읽히는 상태로 남음).",
        "hard_violations": [],
        "physics": "현우의 오른손이 총을 안정적으로 쥐고 있으며, 화면 하단 끄트머리에서 쿠마의 손가락이 꼬치구이를 쥐고 지지함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "두 옆얼굴의 클로즈업과 이마에 맞닿은 총구, 현우의 손 및 개머리판 달린 총은 충실하지만, 읽을 수 있는 총기 각인이 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "배경 인물 두 명과 중복된 드럼통·그릴이 실격 사유이며, 넓어진 구도와 쿠마의 외모·의상 및 총기 형태도 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 쿠마를 응시하고 쿠마도 현우 쪽을 본다. 권총은 현우의 손에서 오른쪽으로 향하며 약간 위로 기울어져 쿠마의 눈썹 위 이마에 직접 닿는다. 측면에서 총열 축과 접촉점이 모두 읽히며, 이마 중앙을 겨누는 행동에 부합한다.",
        "built_space": "두 사람의 얼굴과 현우의 손이 전경을 크게 차지한다. 뒤에는 불타는 금속 드럼통 한 개와 직사각형 그릴 한 개가 보이고, 오른쪽 방조제와 왼쪽 작업장·트럭, 멀리 바다와 항만 조명이 이어져 장소 참조와 일치한다. 구조물 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 현우와 쿠마 두 명뿐이다. 현우는 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 회색 셔츠와 치료되지 않은 얼굴 상처를 갖췄다. 쿠마는 짧은 검은 머리의 젊은 성인 남성이며 참조의 갈색 가죽 재킷을 입었다. 구체적인 혈통은 외관만으로 확정할 수 없지만 두 인물의 외형은 참조에 대체로 부합한다. 총은 참조처럼 개머리판이 달린 금속 권총이며, 그릴 위 닭꼬치와 쿠마가 든 꼬치도 보인다. 다만 총 측면에는 판독 가능한 문자와 숫자 각인이 있다.",
        "hard_violations": [
         "권총 측면에 읽을 수 있는 문자·숫자 각인이 노출되어, 화면 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "현우의 회색 소매에서 이어지는 손목과 손이 권총 손잡이를 감싸 총을 지탱한다. 총구와 이마의 접촉도 물리적으로 가능하다. 쿠마의 닭꼬치는 화면 하단에 보이는 손가락이 아래쪽 막대를 잡고 있어 떠 있지 않다. 두 사람의 하체는 클로즈업 밖이며, 보이는 상체에 부유나 불가능한 관절 자세는 없다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 아래의 쿠마를 내려다보고, 쿠마는 현우를 올려다본다. 현우가 뻗은 팔의 권총은 오른쪽 아래로 향하며 쿠마의 윗이마에 닿는다. 총열 방향과 접촉은 읽히지만, 서로 비슷한 높이에서 마주 보는 두 옆얼굴보다 위에서 아래로 제압하는 배치가 강조된다.",
        "built_space": "방조제, 바다, 작업장과 트럭은 참조 장소를 따른다. 그러나 왼쪽 뒤와 오른쪽 중앙에 드럼통이 각각 하나씩, 왼쪽과 하단 중앙에 그릴이 각각 하나씩 보여 단일 드럼통·그릴 배치가 중복된다. 두 주인공 사이 통로에는 추가 인물 두 명이 서 있다. 현우의 몸통과 뻗은 팔, 쿠마의 상반신 및 넓은 배경까지 포함해 요구한 얼굴 중심 클로즈업보다 넓다.",
        "entities": "현우는 검은 머리와 얼굴 상처, 회색 계열 셔츠를 갖춘 젊은 동아시아계 남성으로 보인다. 쿠마 역할의 인물은 참조보다 훨씬 앳되고 머리가 길며, 갈색 가죽 재킷 대신 어두운 티셔츠를 입어 참조 정체성과 의상이 맞지 않는다. 배경에는 허용되지 않은 성인 남성 두 명이 추가됐다. 권총에는 참조의 개머리판이 없고 총몸 형태도 다르다. 쿠마의 손에 닭꼬치가 있으며 전경 그릴에도 꼬치가 놓여 있다.",
        "hard_violations": [
         "샷에 허용되지 않은 배경 인물 두 명을 추가했다.",
         "참조 장소의 단일 그릴을 두 개로 중복했다.",
         "참조 장소의 단일 금속 드럼통을 두 개로 중복했다."
        ],
        "physics": "현우의 손과 곧게 뻗은 팔이 권총을 지탱하며, 쿠마의 손은 닭꼬치 막대를 잡고 있다. 쿠마는 방조제 앞 낮은 위치에 있고 하체와 좌면은 프레임 밖이므로 지지 방식을 확정할 수 없지만, 보이는 부분만으로 부유한다고 판단할 근거는 없다. 그릴과 드럼통은 바닥에 놓여 있고 꼬치는 석쇠 위에 받쳐져 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "두 옆얼굴의 클로즈업과 이마에 맞닿은 총구, 현우의 손 및 개머리판 달린 총은 충실하지만, 읽을 수 있는 총기 각인이 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "배경 인물 두 명과 중복된 드럼통·그릴이 실격 사유이며, 넓어진 구도와 쿠마의 외모·의상 및 총기 형태도 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 쿠마를 응시하고 쿠마도 현우 쪽을 본다. 권총은 현우의 손에서 오른쪽으로 향하며 약간 위로 기울어져 쿠마의 눈썹 위 이마에 직접 닿는다. 측면에서 총열 축과 접촉점이 모두 읽히며, 이마 중앙을 겨누는 행동에 부합한다.",
        "built_space": "두 사람의 얼굴과 현우의 손이 전경을 크게 차지한다. 뒤에는 불타는 금속 드럼통 한 개와 직사각형 그릴 한 개가 보이고, 오른쪽 방조제와 왼쪽 작업장·트럭, 멀리 바다와 항만 조명이 이어져 장소 참조와 일치한다. 구조물 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 현우와 쿠마 두 명뿐이다. 현우는 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 회색 셔츠와 치료되지 않은 얼굴 상처를 갖췄다. 쿠마는 짧은 검은 머리의 젊은 성인 남성이며 참조의 갈색 가죽 재킷을 입었다. 구체적인 혈통은 외관만으로 확정할 수 없지만 두 인물의 외형은 참조에 대체로 부합한다. 총은 참조처럼 개머리판이 달린 금속 권총이며, 그릴 위 닭꼬치와 쿠마가 든 꼬치도 보인다. 다만 총 측면에는 판독 가능한 문자와 숫자 각인이 있다.",
        "hard_violations": [
         "권총 측면에 읽을 수 있는 문자·숫자 각인이 노출되어, 화면 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "현우의 회색 소매에서 이어지는 손목과 손이 권총 손잡이를 감싸 총을 지탱한다. 총구와 이마의 접촉도 물리적으로 가능하다. 쿠마의 닭꼬치는 화면 하단에 보이는 손가락이 아래쪽 막대를 잡고 있어 떠 있지 않다. 두 사람의 하체는 클로즈업 밖이며, 보이는 상체에 부유나 불가능한 관절 자세는 없다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 아래의 쿠마를 내려다보고, 쿠마는 현우를 올려다본다. 현우가 뻗은 팔의 권총은 오른쪽 아래로 향하며 쿠마의 윗이마에 닿는다. 총열 방향과 접촉은 읽히지만, 서로 비슷한 높이에서 마주 보는 두 옆얼굴보다 위에서 아래로 제압하는 배치가 강조된다.",
        "built_space": "방조제, 바다, 작업장과 트럭은 참조 장소를 따른다. 그러나 왼쪽 뒤와 오른쪽 중앙에 드럼통이 각각 하나씩, 왼쪽과 하단 중앙에 그릴이 각각 하나씩 보여 단일 드럼통·그릴 배치가 중복된다. 두 주인공 사이 통로에는 추가 인물 두 명이 서 있다. 현우의 몸통과 뻗은 팔, 쿠마의 상반신 및 넓은 배경까지 포함해 요구한 얼굴 중심 클로즈업보다 넓다.",
        "entities": "현우는 검은 머리와 얼굴 상처, 회색 계열 셔츠를 갖춘 젊은 동아시아계 남성으로 보인다. 쿠마 역할의 인물은 참조보다 훨씬 앳되고 머리가 길며, 갈색 가죽 재킷 대신 어두운 티셔츠를 입어 참조 정체성과 의상이 맞지 않는다. 배경에는 허용되지 않은 성인 남성 두 명이 추가됐다. 권총에는 참조의 개머리판이 없고 총몸 형태도 다르다. 쿠마의 손에 닭꼬치가 있으며 전경 그릴에도 꼬치가 놓여 있다.",
        "hard_violations": [
         "샷에 허용되지 않은 배경 인물 두 명을 추가했다.",
         "참조 장소의 단일 그릴을 두 개로 중복했다.",
         "참조 장소의 단일 금속 드럼통을 두 개로 중복했다."
        ],
        "physics": "현우의 손과 곧게 뻗은 팔이 권총을 지탱하며, 쿠마의 손은 닭꼬치 막대를 잡고 있다. 쿠마는 방조제 앞 낮은 위치에 있고 하체와 좌면은 프레임 밖이므로 지지 방식을 확정할 수 없지만, 보이는 부분만으로 부유한다고 판단할 근거는 없다. 그릴과 드럼통은 바닥에 놓여 있고 꼬치는 석쇠 위에 받쳐져 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.679,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.429,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 없는 인물 2명이 배경에 추가됨 (extra bodies)",
     "[gemini-pro] 표적 인물이 쿠마가 아님 (invented people)",
     "[gpt-high] 샷에 허용되지 않은 배경 인물 두 명을 추가했다.",
     "[gpt-high] 참조 장소의 단일 그릴을 두 개로 중복했다.",
     "[gpt-high] 참조 장소의 단일 금속 드럼통을 두 개로 중복했다."
    ],
    "B": [
     "[gpt-high] 권총 측면에 읽을 수 있는 문자·숫자 각인이 노출되어, 화면 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 429,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 429,
    "verdict_ko": "지정된 클로즈업 샷 구도를 무시하고 화면을 넓혔으며, 쿠마의 외형이 불일치하고 배경에 지시되지 않은 인물이 등장하는 치명적인 위반이 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물 2명이 배경에 추가됨 (extra bodies) / [gemini-pro] 표적 인물이 쿠마가 아님 (invented people) / [gpt-high] 샷에 허용되지 않은 배경 인물 두 명을 추가했다. / [gpt-high] 참조 장소의 단일 그릴을 두 개로 중복했다. / [gpt-high] 참조 장소의 단일 금속 드럼통을 두 개로 중복했다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "두 인물의 옆얼굴을 총구가 이어주는 정확한 클로즈업 구도를 구현했으며, 인물들의 외모와 무기 소품의 디테일을 프롬프트에 맞게 훌륭히 반영했습니다.  ★위반: [gpt-high] 권총 측면에 읽을 수 있는 문자·숫자 각인이 노출되어, 화면 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_roadside_hideout_f5126d.png",
    "asset_id": "3da48160-9267-4344-8758-06815d5af5e3",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:806945>",
    "asset_id": "fe00e8c2-fc45-4464-b8b5-95ce8056561f",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-c580-7f34-8556-9ba5d61d8f43",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh10__bgfirst_bg.png",
   "bg_asset_id": "51a6f4c4-1510-44cd-986d-c3a6e8933313",
   "bg_record_key": "S13sh10::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "roadside_hideout",
   "groupbg_asset_id": "3da48160-9267-4344-8758-06815d5af5e3"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S13sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:27:11.616098+00:00",
  "fingerprint": "eac264c9a59166955f1e3f556a8b5f398512f02a286e8f5819c2c92bc40c72aa",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S13sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S13sh10_sel.png",
  "source_sha256": "8a479f997cda156b36770d4f5c759447a3b0e8f692758b971b9c20738107c9c9",
  "file": "S13sh10_cine.png",
  "staged_sha256": "876b7e5c8b6eba48e94d4a987a488311c4757c54bdcd4fb114b87072f7ea9a3d",
  "latency_ms": 9466
 },
 "S13sh13::signage": {
  "fp": "1bd17acfabc10faf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S13sh13": {
  "input_fingerprint": "9e94b789a25aef50",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 갈라진 콘크리트 틈새로 탁한 물이 줄줄 흘러내리는 거대한 인공제방의 젖은 벽면 클로즈업.\n\nLOCATION (lock): Directly against the exterior face of the refugee settlement's concrete seawall, where water seeps through deep cracks at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Wet, with murky water running through the concrete cracks) — Its exposed face is viewed obliquely, revealing the crack and the wall's recession; used as Provides the environmental field around the narrow band of active seepage; Water flowing from the cracks (Running downward over the wall); used as Forms the primary moving detail without relying on a reflection or an added atmospheric effect; Embankment supports (Water also runs over the supports) — Only a peripheral portion is visible along the receding wall edge; used as Maintains structural context at the boundary of the close view.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the wall within the scene's established nighttime darkness, using restrained tonal separation to reveal the flowing water without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has visible cracks with seawater seeping through, and water runs along its supports in the nighttime darkness. The nearby hideout's drum fire and barbecue remain established.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 갈라진 콘크리트 틈새로 탁한 물이 줄줄 흘러내리는 거대한 인공제방의 젖은 벽면 클로즈업.\n\nLOCATION (lock): Directly against the exterior face of the refugee settlement's concrete seawall, where water seeps through deep cracks at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Wet, with murky water running through the concrete cracks) — Its exposed face is viewed obliquely, revealing the crack and the wall's recession; used as Provides the environmental field around the narrow band of active seepage; Water flowing from the cracks (Running downward over the wall); used as Forms the primary moving detail without relying on a reflection or an added atmospheric effect; Embankment supports (Water also runs over the supports) — Only a peripheral portion is visible along the receding wall edge; used as Maintains structural context at the boundary of the close view.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the wall within the scene's established nighttime darkness, using restrained tonal separation to reveal the flowing water without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has visible cracks with seawater seeping through, and water runs along its supports in the nighttime darkness. The nearby hideout's drum fire and barbecue remain established.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 갈라진 콘크리트 틈새로 탁한 물이 줄줄 흘러내리는 거대한 인공제방의 젖은 벽면 클로즈업.\n\nLOCATION (lock): Directly against the exterior face of the refugee settlement's concrete seawall, where water seeps through deep cracks at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Wet, with murky water running through the concrete cracks) — Its exposed face is viewed obliquely, revealing the crack and the wall's recession; used as Provides the environmental field around the narrow band of active seepage; Water flowing from the cracks (Running downward over the wall); used as Forms the primary moving detail without relying on a reflection or an added atmospheric effect; Embankment supports (Water also runs over the supports) — Only a peripheral portion is visible along the receding wall edge; used as Maintains structural context at the boundary of the close view.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the wall within the scene's established nighttime darkness, using restrained tonal separation to reveal the flowing water without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has visible cracks with seawater seeping through, and water runs along its supports in the nighttime darkness. The nearby hideout's drum fire and barbecue remain established.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "해당 사항 없음.",
    "built_space": "비스듬한 각도의 패널형 콘크리트 벽면, 수직 균열 및 배관.",
    "entities": "갈라진 제방 벽(질감이 레퍼런스와 다름), 탁한 물, 배관. 요구된 제방 지지대 없음. 인물 없음.",
    "hard_violations": [],
    "physics": "물이 배관과 균열에서 아래로 떨어짐."
   },
   {
    "label": "B",
    "direction": "해당 사항 없음.",
    "built_space": "비스듬한 각도의 거친 콘크리트 벽면, 깊은 균열, 우측 가장자리를 따라 늘어선 A자형 금속 제방 지지대.",
    "entities": "갈라진 제방 벽(레퍼런스와 질감 일치), 탁한 물, 우측의 제방 지지대. 인물 없음.",
    "hard_violations": [],
    "physics": "물이 균열을 타고 중력에 맞춰 자연스럽게 아래로 흐름."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 거친 콘크리트 질감을 잘 살렸으며, 우측 가장자리에 프롬프트가 요구한 제방 지지대를 정확히 묘사함."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "벽면 질감이 레퍼런스와 달리 지나치게 매끄러우며, 화면 가장자리에 보여야 할 제방 지지대가 누락됨."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "해당 사항 없음.",
        "built_space": "비스듬한 각도의 거친 콘크리트 벽면, 깊은 균열, 우측 가장자리를 따라 늘어선 A자형 금속 제방 지지대.",
        "entities": "갈라진 제방 벽(레퍼런스와 질감 일치), 탁한 물, 우측의 제방 지지대. 인물 없음.",
        "hard_violations": [],
        "physics": "물이 균열을 타고 중력에 맞춰 자연스럽게 아래로 흐름."
       },
       {
        "label": "A",
        "direction": "해당 사항 없음.",
        "built_space": "비스듬한 각도의 패널형 콘크리트 벽면, 수직 균열 및 배관.",
        "entities": "갈라진 제방 벽(질감이 레퍼런스와 다름), 탁한 물, 배관. 요구된 제방 지지대 없음. 인물 없음.",
        "hard_violations": [],
        "physics": "물이 배관과 균열에서 아래로 떨어짐."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 거친 콘크리트 질감을 잘 살렸으며, 우측 가장자리에 프롬프트가 요구한 제방 지지대를 정확히 묘사함."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "벽면 질감이 레퍼런스와 달리 지나치게 매끄러우며, 화면 가장자리에 보여야 할 제방 지지대가 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "해당 사항 없음.",
        "built_space": "비스듬한 각도의 거친 콘크리트 벽면, 깊은 균열, 우측 가장자리를 따라 늘어선 A자형 금속 제방 지지대.",
        "entities": "갈라진 제방 벽(레퍼런스와 질감 일치), 탁한 물, 우측의 제방 지지대. 인물 없음.",
        "hard_violations": [],
        "physics": "물이 균열을 타고 중력에 맞춰 자연스럽게 아래로 흐름."
       },
       {
        "label": "A",
        "direction": "해당 사항 없음.",
        "built_space": "비스듬한 각도의 패널형 콘크리트 벽면, 수직 균열 및 배관.",
        "entities": "갈라진 제방 벽(질감이 레퍼런스와 다름), 탁한 물, 배관. 요구된 제방 지지대 없음. 인물 없음.",
        "hard_violations": [],
        "physics": "물이 배관과 균열에서 아래로 떨어짐."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "젖고 파손된 벽면을 비스듬한 클로즈업으로 담고, 균열에서 내려오는 탁한 물과 가장자리 지지 구조를 함께 보여 주어 핵심 구도와 장소 연속성에 더 충실하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "야간 벽면 클로즈업과 균열의 하향 물줄기는 맞지만, 주변부 지지대 대신 양쪽 관의 방류가 두드러져 요청된 누수 중심 구성과 구조적 맥락이 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 깊은 균열에서 나온 갈색 물이 돌출된 콘크리트 면을 타고 화면 아래로 흐른다. 벽은 오른쪽 뒤로 멀어지며, 물의 진행 방향은 벽을 따라 내려오는 누수라는 지시와 맞는다. 사람이나 시선, 겨냥하는 물체는 없다.",
        "built_space": "파손된 콘크리트 외벽이 화면 대부분을 차지한다. 중앙 오른쪽에 고정 띠가 달린 세로 관 하나가 뚜렷하고, 오른쪽 아래에는 삼각형 지지대 하나와 뒤로 이어지는 가로 관·지지 구조 일부가 보인다. 지지 구조가 가까운 벽의 가장자리에 제한되어 있어 요청된 주변부 배치에 가깝다. 참조의 크게 갈라진 벽과 관형 구조물의 재질도 이어진다.",
        "entities": "거친 회색 콘크리트, 깊게 벌어진 균열, 광택이 있는 젖은 표면, 탁한 갈색 물, 가장자리 지지 구조가 보인다. 사람·얼굴·읽을 수 있는 글자는 없다. 드럼통 불과 바비큐는 이 클로즈업에 보이지 않으며 이를 보여 주려고 구도를 넓히지 않았다. 야간의 어두운 색조이며 별도의 광원은 화면에 없다.",
        "hard_violations": [],
        "physics": "물은 균열이라는 배출 지점에서 시작해 콘크리트 표면에 닿은 채 중력 방향으로 흘러내린다. 벽의 돌출부에서 물줄기가 굽어지는 모습도 가능하다. 관은 벽의 고정 띠로 지지되고, 주변 지지대는 하단까지 이어진다. 지지대 표면은 젖어 있으나 그 위를 흐르는 개별 물줄기는 중앙 누수만큼 명확하지 않다. 근거 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "중앙 세로 균열에서 나온 물이 화면 아래쪽으로 흐른다. 좌우의 관 끝에서도 각각 물줄기가 아래로 떨어진다. 중앙 누수의 방향은 맞지만, 흐름의 출발점 두 곳이 콘크리트 균열이 아니라 관의 출구로 보인다. 사람이나 시선은 없다.",
        "built_space": "콘크리트 벽을 가까이 비스듬하게 보며 벽은 왼쪽 뒤로 멀어진다. 중앙에 큰 세로 균열 하나와 가지 균열들이 있고, 왼쪽과 오른쪽에 세로 관이 하나씩, 왼쪽 위에 상단 난간 일부가 보인다. 참조의 콘크리트와 관 재질은 유사하지만, 요청된 가장자리 지지대 부분은 식별되지 않는다.",
        "entities": "젖은 회색 콘크리트, 벌어진 균열, 중앙의 회갈색 물줄기와 관 두 개가 보인다. 물은 불투명하지만 A보다 탁한 갈색 느낌이 약하다. 사람·얼굴·읽을 수 있는 글자는 없고, 화면 밖 소품을 추가로 끌어들이지 않았다. 야간 색조는 유지되며 새로운 광원 자체는 보이지 않는다.",
        "hard_violations": [],
        "physics": "중앙 물줄기는 균열에서 나와 벽에 접촉하며 내려오고, 양쪽 물줄기는 관의 열린 끝에서 낙하하므로 출발점과 중력 방향이 설명된다. 관들은 벽을 따라 설치되어 고정 부위가 보인다. 물이나 고체가 근거 없이 공중에 정지한 모습은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "젖고 파손된 벽면을 비스듬한 클로즈업으로 담고, 균열에서 내려오는 탁한 물과 가장자리 지지 구조를 함께 보여 주어 핵심 구도와 장소 연속성에 더 충실하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간 벽면 클로즈업과 균열의 하향 물줄기는 맞지만, 주변부 지지대 대신 양쪽 관의 방류가 두드러져 요청된 누수 중심 구성과 구조적 맥락이 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "중앙의 깊은 균열에서 나온 갈색 물이 돌출된 콘크리트 면을 타고 화면 아래로 흐른다. 벽은 오른쪽 뒤로 멀어지며, 물의 진행 방향은 벽을 따라 내려오는 누수라는 지시와 맞는다. 사람이나 시선, 겨냥하는 물체는 없다.",
        "built_space": "파손된 콘크리트 외벽이 화면 대부분을 차지한다. 중앙 오른쪽에 고정 띠가 달린 세로 관 하나가 뚜렷하고, 오른쪽 아래에는 삼각형 지지대 하나와 뒤로 이어지는 가로 관·지지 구조 일부가 보인다. 지지 구조가 가까운 벽의 가장자리에 제한되어 있어 요청된 주변부 배치에 가깝다. 참조의 크게 갈라진 벽과 관형 구조물의 재질도 이어진다.",
        "entities": "거친 회색 콘크리트, 깊게 벌어진 균열, 광택이 있는 젖은 표면, 탁한 갈색 물, 가장자리 지지 구조가 보인다. 사람·얼굴·읽을 수 있는 글자는 없다. 드럼통 불과 바비큐는 이 클로즈업에 보이지 않으며 이를 보여 주려고 구도를 넓히지 않았다. 야간의 어두운 색조이며 별도의 광원은 화면에 없다.",
        "hard_violations": [],
        "physics": "물은 균열이라는 배출 지점에서 시작해 콘크리트 표면에 닿은 채 중력 방향으로 흘러내린다. 벽의 돌출부에서 물줄기가 굽어지는 모습도 가능하다. 관은 벽의 고정 띠로 지지되고, 주변 지지대는 하단까지 이어진다. 지지대 표면은 젖어 있으나 그 위를 흐르는 개별 물줄기는 중앙 누수만큼 명확하지 않다. 근거 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "중앙 세로 균열에서 나온 물이 화면 아래쪽으로 흐른다. 좌우의 관 끝에서도 각각 물줄기가 아래로 떨어진다. 중앙 누수의 방향은 맞지만, 흐름의 출발점 두 곳이 콘크리트 균열이 아니라 관의 출구로 보인다. 사람이나 시선은 없다.",
        "built_space": "콘크리트 벽을 가까이 비스듬하게 보며 벽은 왼쪽 뒤로 멀어진다. 중앙에 큰 세로 균열 하나와 가지 균열들이 있고, 왼쪽과 오른쪽에 세로 관이 하나씩, 왼쪽 위에 상단 난간 일부가 보인다. 참조의 콘크리트와 관 재질은 유사하지만, 요청된 가장자리 지지대 부분은 식별되지 않는다.",
        "entities": "젖은 회색 콘크리트, 벌어진 균열, 중앙의 회갈색 물줄기와 관 두 개가 보인다. 물은 불투명하지만 A보다 탁한 갈색 느낌이 약하다. 사람·얼굴·읽을 수 있는 글자는 없고, 화면 밖 소품을 추가로 끌어들이지 않았다. 야간 색조는 유지되며 새로운 광원 자체는 보이지 않는다.",
        "hard_violations": [],
        "physics": "중앙 물줄기는 균열에서 나와 벽에 접촉하며 내려오고, 양쪽 물줄기는 관의 열린 끝에서 낙하하므로 출발점과 중력 방향이 설명된다. 관들은 벽을 따라 설치되어 고정 부위가 보인다. 물이나 고체가 근거 없이 공중에 정지한 모습은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.492,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.492,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1492
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "레퍼런스의 거친 콘크리트 질감을 잘 살렸으며, 우측 가장자리에 프롬프트가 요구한 제방 지지대를 정확히 묘사함."
   },
   {
    "label": "A",
    "score": 1492,
    "verdict_ko": "벽면 질감이 레퍼런스와 달리 지나치게 매끄러우며, 화면 가장자리에 보여야 할 제방 지지대가 누락됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11_sel.png",
    "asset_id": "92376a28-6a24-41e0-8ef5-2c96ea1c4239",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-ca5e-73c9-b6f1-3db559be39b8",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S7sh11"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S13sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:28:00.945346+00:00",
  "fingerprint": "344728e8cf65fb4646477757bfad17b4e097a8f459ef1b5b27ab788f7a3f45b3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S13sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S13sh13_sel.png",
  "source_sha256": "1447bf8ba0fd5d69fb1df1bb65728d5a6dd633b8e5f7bc36e7c8491e5523b128",
  "file": "S13sh13_cine.png",
  "staged_sha256": "8c53a0a5e531757302f3500b91530f528235e2ea04867fe0cc49fa980b17aed7",
  "latency_ms": 9304
 },
 "S13sh20::signage": {
  "fp": "e96c4612775fbad4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S13sh20": {
  "input_fingerprint": "ee3b8ef4e8d56ae2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 쿠마의 손에 들린 지폐 뭉치를 재빠르게 움켜쥔 찰나의 현우 손 클로즈업.\n\nLOCATION (lock): At the foot of the leaking seawall beside the refugee settlement, where the nighttime negotiation concludes. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Banknotes (Still supported by 쿠마 as 현우 grips them) — Overlapping note faces are seen obliquely between the hands; no denomination needs to be legible; used as Marks the exact transfer of possession while remaining subordinate in size to the hands; Embankment wall (Cracked and wet in the surrounding location) — A soft, partial section remains beyond the cropped torsos; used as Retains the negotiation's location without competing with the hand action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime ambient visibility at the embankment without carrying the hideout's firelight into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the cracked concrete wall, persistent seawater seepage, and dark nighttime coloration as the environmental reference. Exclude the barbecue grill, beer, and burning barrel from the separate hideout.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains wet and cracked, with water running along the supports under nighttime darkness. The nearby drum fire and barbecue have not been extinguished. 현우: Hyunwoo is taking possession of the cash payment and still retains the seized gun; his upper garment is on and his facial and leg injuries persist. 쿠마: Kuma stands beside the seawall with one palm still wet, completing the cash handover from inside his clothing.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마, 현우 right now, so 쿠마, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 쿠마의 손에 들린 지폐 뭉치를 재빠르게 움켜쥔 찰나의 현우 손 클로즈업.\n\nLOCATION (lock): At the foot of the leaking seawall beside the refugee settlement, where the nighttime negotiation concludes. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Banknotes (Still supported by 쿠마 as 현우 grips them) — Overlapping note faces are seen obliquely between the hands; no denomination needs to be legible; used as Marks the exact transfer of possession while remaining subordinate in size to the hands; Embankment wall (Cracked and wet in the surrounding location) — A soft, partial section remains beyond the cropped torsos; used as Retains the negotiation's location without competing with the hand action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime ambient visibility at the embankment without carrying the hideout's firelight into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the cracked concrete wall, persistent seawater seepage, and dark nighttime coloration as the environmental reference. Exclude the barbecue grill, beer, and burning barrel from the separate hideout.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains wet and cracked, with water running along the supports under nighttime darkness. The nearby drum fire and barbecue have not been extinguished. 현우: Hyunwoo is taking possession of the cash payment and still retains the seized gun; his upper garment is on and his facial and leg injuries persist. 쿠마: Kuma stands beside the seawall with one palm still wet, completing the cash handover from inside his clothing.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마, 현우 right now, so 쿠마, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 쿠마의 손에 들린 지폐 뭉치를 재빠르게 움켜쥔 찰나의 현우 손 클로즈업.\n\nLOCATION (lock): At the foot of the leaking seawall beside the refugee settlement, where the nighttime negotiation concludes. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Banknotes (Still supported by 쿠마 as 현우 grips them) — Overlapping note faces are seen obliquely between the hands; no denomination needs to be legible; used as Marks the exact transfer of possession while remaining subordinate in size to the hands; Embankment wall (Cracked and wet in the surrounding location) — A soft, partial section remains beyond the cropped torsos; used as Retains the negotiation's location without competing with the hand action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime ambient visibility at the embankment without carrying the hideout's firelight into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the cracked concrete wall, persistent seawater seepage, and dark nighttime coloration as the environmental reference. Exclude the barbecue grill, beer, and burning barrel from the separate hideout.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains wet and cracked, with water running along the supports under nighttime darkness. The nearby drum fire and barbecue have not been extinguished. 현우: Hyunwoo is taking possession of the cash payment and still retains the seized gun; his upper garment is on and his facial and leg injuries persist. 쿠마: Kuma stands beside the seawall with one palm still wet, completing the cash handover from inside his clothing.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마, 현우 right now, so 쿠마, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "양측에서 손이 뻗어 지폐 다발을 잡고 있으며, 우측 인물의 시선이 돈을 향함.",
    "built_space": "갈라지고 물이 흐르는 콘크리트 벽이 배경에 적절히 배치됨.",
    "entities": "우측에 현우의 얼굴과 손, 좌측에 쿠마의 팔이 등장하나, 현우의 필수 소지품인 총이 없음.",
    "hard_violations": [
     "[gemini-pro] 손 클로즈업 및 상체 크롭 지시를 위반하고 현우의 얼굴을 넓게 프레이밍함",
     "[gemini-pro] 현우가 쥐고 있어야 할 총(seized gun) 누락",
     "[gpt-high] 펼쳐진 지폐에 액면 숫자와 인쇄 문자가 판독 가능하게 노출되어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "양손이 지폐를 물리적으로 자연스럽게 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "양측에서 뻗은 손이 화면 중앙에서 지폐 다발을 맞잡고 교환을 진행함.",
    "built_space": "젖어서 갈라진 콘크리트 방파제가 배경에 잘 반영됨.",
    "entities": "잘린 상체로 등장한 현우(좌측, 한 손에 총을 듦)와 쿠마(우측, 젖은 손)가 지시사항과 일치함.",
    "hard_violations": [
     "[gpt-high] 지폐 윗부분의 액면 숫자가 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
    ],
    "physics": "손이 지폐 다발과 총의 그립을 올바르게 쥐고 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "크롭된 상체와 손 클로즈업 프레이밍을 정확히 구현했으며, 현우의 총과 쿠마의 젖은 손 등 요구된 세부 묘사를 충실히 반영했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "손 클로즈업 프레이밍 지시를 어기고 인물의 얼굴을 노출했으며, 현우가 들고 있어야 할 총이 누락되어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "양측에서 손이 뻗어 지폐 다발을 잡고 있으며, 우측 인물의 시선이 돈을 향함.",
        "built_space": "갈라지고 물이 흐르는 콘크리트 벽이 배경에 적절히 배치됨.",
        "entities": "우측에 현우의 얼굴과 손, 좌측에 쿠마의 팔이 등장하나, 현우의 필수 소지품인 총이 없음.",
        "hard_violations": [
         "손 클로즈업 및 상체 크롭 지시를 위반하고 현우의 얼굴을 넓게 프레이밍함",
         "현우가 쥐고 있어야 할 총(seized gun) 누락"
        ],
        "physics": "양손이 지폐를 물리적으로 자연스럽게 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "양측에서 뻗은 손이 화면 중앙에서 지폐 다발을 맞잡고 교환을 진행함.",
        "built_space": "젖어서 갈라진 콘크리트 방파제가 배경에 잘 반영됨.",
        "entities": "잘린 상체로 등장한 현우(좌측, 한 손에 총을 듦)와 쿠마(우측, 젖은 손)가 지시사항과 일치함.",
        "hard_violations": [],
        "physics": "손이 지폐 다발과 총의 그립을 올바르게 쥐고 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "크롭된 상체와 손 클로즈업 프레이밍을 정확히 구현했으며, 현우의 총과 쿠마의 젖은 손 등 요구된 세부 묘사를 충실히 반영했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "손 클로즈업 프레이밍 지시를 어기고 인물의 얼굴을 노출했으며, 현우가 들고 있어야 할 총이 누락되어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "양측에서 손이 뻗어 지폐 다발을 잡고 있으며, 우측 인물의 시선이 돈을 향함.",
        "built_space": "갈라지고 물이 흐르는 콘크리트 벽이 배경에 적절히 배치됨.",
        "entities": "우측에 현우의 얼굴과 손, 좌측에 쿠마의 팔이 등장하나, 현우의 필수 소지품인 총이 없음.",
        "hard_violations": [
         "손 클로즈업 및 상체 크롭 지시를 위반하고 현우의 얼굴을 넓게 프레이밍함",
         "현우가 쥐고 있어야 할 총(seized gun) 누락"
        ],
        "physics": "양손이 지폐를 물리적으로 자연스럽게 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "양측에서 뻗은 손이 화면 중앙에서 지폐 다발을 맞잡고 교환을 진행함.",
        "built_space": "젖어서 갈라진 콘크리트 방파제가 배경에 잘 반영됨.",
        "entities": "잘린 상체로 등장한 현우(좌측, 한 손에 총을 듦)와 쿠마(우측, 젖은 손)가 지시사항과 일치함.",
        "hard_violations": [],
        "physics": "손이 지폐 다발과 총의 그립을 올바르게 쥐고 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 쿠마의 젖은 손에 받쳐진 돈을 움켜쥐는 관계와 손 중심 구도는 더 정확하지만, 지폐의 읽히는 액면 숫자는 명시적 금지 위반이다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "읽히는 지폐 숫자에 더해, 현우로 보이는 오른쪽 인물이 돈을 받치고 왼쪽 손이 가져가는 역할 역전과 불필요한 얼굴 노출로 핵심 순간의 충실도가 낮다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우의 손이 오른쪽 쿠마가 받친 지폐 뭉치를 향해 뻗어 가장자리를 움켜쥔다. 쿠마의 손바닥은 위를 향해 돈 아래에 남아 있다. 현우의 다른 손에 든 권총 총구는 화면 오른쪽 아래를 향하며 상대를 겨누지 않는다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "두 사람의 잘린 몸통 사이로 젖고 크게 갈라진 콘크리트 방벽 한 면이 보인다. 오른쪽 뒤에는 수직 배관 한 줄, 아래에는 금속 지지대 일부가 보이며 참고 장소와 부합한다. 고정물 중복이나 불가능한 반사는 없다. 다만 벽의 균열과 누수 질감이 상당히 선명해 요청한 부드러운 부분 배경보다 두드러진다.",
        "entities": "현우 쪽에는 참고와 맞는 남색 상의, 젊은 사람의 맨팔과 상처 난 손, 다른 손에 잡힌 권총이 보인다. 쿠마 쪽에는 젖은 손바닥이 보이지만 검은 지퍼 겉옷은 참고의 남색 티셔츠와 다르다. 얼굴이 없어 정확한 나이와 한국계 미국인·중국계 혼혈 정체성은 확인할 수 없다. 겹친 실물 지폐가 손 사이에 비스듬히 놓였으나 일부 액면 숫자를 읽을 수 있다. 통화 국가는 확정하기 어렵고 프롬프트도 특정하지 않았다. 추가 인물, 화로, 맥주, 그릴은 없다.",
        "hard_violations": [
         "지폐 윗부분의 액면 숫자가 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
        ],
        "physics": "지폐는 쿠마의 손바닥과 굽힌 손가락이 아래에서 받치고 현우의 손이 반대쪽을 움켜쥐어 양쪽 접촉이 유지된다. 권총도 별도의 손이 손잡이를 쥐고 있어 떠 있지 않다. 팔은 각자의 몸통 방향으로 이어지며, 돈을 빠르게 가져가기 시작하는 동작으로 가능한 자세다. 발과 하체는 구도 밖이므로 지면 접촉은 평가할 수 없다."
       },
       {
        "label": "B",
        "direction": "왼쪽에서 들어온 손이 지폐 윗부분을 움켜쥐고, 오른쪽 손은 두꺼운 뭉치를 아래에서 받친다. 오른쪽 인물은 고개와 시선을 돈 쪽으로 내리고 있다. 이 얼굴과 헝클어진 머리는 현우 쪽에 가까워, 현우가 움켜쥐는 대신 돈을 받쳐 주는 역할로 읽힌다. 무기는 보이지 않는다.",
        "built_space": "젖고 갈라진 콘크리트 방벽 한 면이 화면 대부분을 차지하고, 오른쪽 뒤로 수직 배관 한 줄과 하단 금속 지지대 일부가 이어진다. 참고의 균열과 누수, 어두운 색조는 유지된다. 중복 고정물이나 불가능한 반사는 없다. 벽이 선명하고 넓게 드러나 손 뒤의 흐린 부분 배경이라는 지시에서는 벗어난다.",
        "entities": "오른쪽에 현우를 닮은 젊은 남성의 얼굴 일부와 검은 헝클어진 머리가 보인다. 왼쪽 인물은 팔과 손뿐이라 쿠마의 신원은 확인하기 어렵고, 털과 피부 질감은 참고의 젊은 인상보다 성숙하게 보인다. 양쪽의 회녹색 긴소매는 두 참고의 남색 티셔츠와 다르다. 지폐는 손에 비해 크게 펼쳐져 있고 액면 숫자와 일부 인쇄 문자가 판독 가능하다. 얼굴·다리 부상과 총의 유지 여부는 보이는 범위만으로 확정할 수 없으며, 총이 프레임 밖이라는 이유 자체는 감점하지 않는다.",
        "hard_violations": [
         "펼쳐진 지폐에 액면 숫자와 인쇄 문자가 판독 가능하게 노출되어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "아래쪽 지폐 뭉치는 오른쪽 손바닥과 손가락에 받쳐져 있고, 위로 펼쳐진 부분은 왼쪽 손이 잡고 있다. 종이가 두 손의 힘으로 벌어진 상태는 물리적으로 가능하며 지지 없는 물체는 없다. 오른쪽 인물의 숙인 머리도 몸에서 이어지는 자연스러운 자세다. 다만 위쪽 일부를 떼어 가는 모습이 강해, 현우가 쿠마의 뭉치 전체를 재빨리 움켜쥔 순간은 덜 분명하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우가 쿠마의 젖은 손에 받쳐진 돈을 움켜쥐는 관계와 손 중심 구도는 더 정확하지만, 지폐의 읽히는 액면 숫자는 명시적 금지 위반이다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "읽히는 지폐 숫자에 더해, 현우로 보이는 오른쪽 인물이 돈을 받치고 왼쪽 손이 가져가는 역할 역전과 불필요한 얼굴 노출로 핵심 순간의 충실도가 낮다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우의 손이 오른쪽 쿠마가 받친 지폐 뭉치를 향해 뻗어 가장자리를 움켜쥔다. 쿠마의 손바닥은 위를 향해 돈 아래에 남아 있다. 현우의 다른 손에 든 권총 총구는 화면 오른쪽 아래를 향하며 상대를 겨누지 않는다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "두 사람의 잘린 몸통 사이로 젖고 크게 갈라진 콘크리트 방벽 한 면이 보인다. 오른쪽 뒤에는 수직 배관 한 줄, 아래에는 금속 지지대 일부가 보이며 참고 장소와 부합한다. 고정물 중복이나 불가능한 반사는 없다. 다만 벽의 균열과 누수 질감이 상당히 선명해 요청한 부드러운 부분 배경보다 두드러진다.",
        "entities": "현우 쪽에는 참고와 맞는 남색 상의, 젊은 사람의 맨팔과 상처 난 손, 다른 손에 잡힌 권총이 보인다. 쿠마 쪽에는 젖은 손바닥이 보이지만 검은 지퍼 겉옷은 참고의 남색 티셔츠와 다르다. 얼굴이 없어 정확한 나이와 한국계 미국인·중국계 혼혈 정체성은 확인할 수 없다. 겹친 실물 지폐가 손 사이에 비스듬히 놓였으나 일부 액면 숫자를 읽을 수 있다. 통화 국가는 확정하기 어렵고 프롬프트도 특정하지 않았다. 추가 인물, 화로, 맥주, 그릴은 없다.",
        "hard_violations": [
         "지폐 윗부분의 액면 숫자가 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
        ],
        "physics": "지폐는 쿠마의 손바닥과 굽힌 손가락이 아래에서 받치고 현우의 손이 반대쪽을 움켜쥐어 양쪽 접촉이 유지된다. 권총도 별도의 손이 손잡이를 쥐고 있어 떠 있지 않다. 팔은 각자의 몸통 방향으로 이어지며, 돈을 빠르게 가져가기 시작하는 동작으로 가능한 자세다. 발과 하체는 구도 밖이므로 지면 접촉은 평가할 수 없다."
       },
       {
        "label": "A",
        "direction": "왼쪽에서 들어온 손이 지폐 윗부분을 움켜쥐고, 오른쪽 손은 두꺼운 뭉치를 아래에서 받친다. 오른쪽 인물은 고개와 시선을 돈 쪽으로 내리고 있다. 이 얼굴과 헝클어진 머리는 현우 쪽에 가까워, 현우가 움켜쥐는 대신 돈을 받쳐 주는 역할로 읽힌다. 무기는 보이지 않는다.",
        "built_space": "젖고 갈라진 콘크리트 방벽 한 면이 화면 대부분을 차지하고, 오른쪽 뒤로 수직 배관 한 줄과 하단 금속 지지대 일부가 이어진다. 참고의 균열과 누수, 어두운 색조는 유지된다. 중복 고정물이나 불가능한 반사는 없다. 벽이 선명하고 넓게 드러나 손 뒤의 흐린 부분 배경이라는 지시에서는 벗어난다.",
        "entities": "오른쪽에 현우를 닮은 젊은 남성의 얼굴 일부와 검은 헝클어진 머리가 보인다. 왼쪽 인물은 팔과 손뿐이라 쿠마의 신원은 확인하기 어렵고, 털과 피부 질감은 참고의 젊은 인상보다 성숙하게 보인다. 양쪽의 회녹색 긴소매는 두 참고의 남색 티셔츠와 다르다. 지폐는 손에 비해 크게 펼쳐져 있고 액면 숫자와 일부 인쇄 문자가 판독 가능하다. 얼굴·다리 부상과 총의 유지 여부는 보이는 범위만으로 확정할 수 없으며, 총이 프레임 밖이라는 이유 자체는 감점하지 않는다.",
        "hard_violations": [
         "펼쳐진 지폐에 액면 숫자와 인쇄 문자가 판독 가능하게 노출되어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "아래쪽 지폐 뭉치는 오른쪽 손바닥과 손가락에 받쳐져 있고, 위로 펼쳐진 부분은 왼쪽 손이 잡고 있다. 종이가 두 손의 힘으로 벌어진 상태는 물리적으로 가능하며 지지 없는 물체는 없다. 오른쪽 인물의 숙인 머리도 몸에서 이어지는 자연스러운 자세다. 다만 위쪽 일부를 떼어 가는 모습이 강해, 현우가 쿠마의 뭉치 전체를 재빨리 움켜쥔 순간은 덜 분명하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 손 클로즈업 및 상체 크롭 지시를 위반하고 현우의 얼굴을 넓게 프레이밍함",
     "[gemini-pro] 현우가 쥐고 있어야 할 총(seized gun) 누락",
     "[gpt-high] 펼쳐진 지폐에 액면 숫자와 인쇄 문자가 판독 가능하게 노출되어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "B": [
     "[gpt-high] 지폐 윗부분의 액면 숫자가 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "크롭된 상체와 손 클로즈업 프레이밍을 정확히 구현했으며, 현우의 총과 쿠마의 젖은 손 등 요구된 세부 묘사를 충실히 반영했습니다.  ★위반: [gpt-high] 지폐 윗부분의 액면 숫자가 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "손 클로즈업 프레이밍 지시를 어기고 인물의 얼굴을 노출했으며, 현우가 들고 있어야 할 총이 누락되어 감점되었습니다.  ★위반: [gemini-pro] 손 클로즈업 및 상체 크롭 지시를 위반하고 현우의 얼굴을 넓게 프레이밍함 / [gemini-pro] 현우가 쥐고 있어야 할 총(seized gun) 누락 / [gpt-high] 펼쳐진 지폐에 액면 숫자와 인쇄 문자가 판독 가능하게 노출되어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh13_sel.png",
    "asset_id": "99bf797a-ae7b-4d7c-b294-36d0f6c10c38",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1218434>",
    "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-cc0f-7cb4-a7ac-33fcc2d8625a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S13sh13"
  }
 },
 "S13sh20::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:29:28.714075+00:00",
  "fingerprint": "760d92cb2d987f7854bc29e2533e7e7d7612ce0025e3d1513df35c4e4aaa07a8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S13sh20_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S13sh20_sel.png",
  "source_sha256": "33faaccfa58c967dd5118781a03649256044b82476a74d0e9899186d62d3b47c",
  "file": "S13sh20_cine.png",
  "staged_sha256": "070b50b938006ec0a7958f6e2e6974e7b1b0f901c20489d93ec50a3ca3df0e99",
  "latency_ms": 8480
 },
 "S14sh3::signage": {
  "fp": "3862b78c3466387e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S14sh3": {
  "input_fingerprint": "13f473b9f6c424a9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 깨진 화병 조각 위로 발끝을 슬며시 밀고 있는 mid-action 자세의 찰리 발밑 클로즈업.\n\nLOCATION (lock): On the floor of the container home's shared living area, beside broken vase fragments in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Broken vase fragments (Being pushed along the floor by 찰리's foot); used as A small scattered group ahead of the toe makes the concealment attempt readable without exaggerated fragment scale; Container interior floor (Supporting 찰리's feet and the broken fragments) — Viewed at a downward oblique angle from beside his feet; used as Provides uninterrupted spatial context for the direction and extent of the sliding action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained daytime ambient illumination appropriate to the container interior, preserving readable floor-level detail without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments are on the container floor, being pushed aside rather than removed. Charlie retains his dirty aged gorilla-shaped casing, blue-lit eyes, worn Ubik chest logo, netting, and previously donned old coat and hat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 깨진 화병 조각 위로 발끝을 슬며시 밀고 있는 mid-action 자세의 찰리 발밑 클로즈업.\n\nLOCATION (lock): On the floor of the container home's shared living area, beside broken vase fragments in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Broken vase fragments (Being pushed along the floor by 찰리's foot); used as A small scattered group ahead of the toe makes the concealment attempt readable without exaggerated fragment scale; Container interior floor (Supporting 찰리's feet and the broken fragments) — Viewed at a downward oblique angle from beside his feet; used as Provides uninterrupted spatial context for the direction and extent of the sliding action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained daytime ambient illumination appropriate to the container interior, preserving readable floor-level detail without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments are on the container floor, being pushed aside rather than removed. Charlie retains his dirty aged gorilla-shaped casing, blue-lit eyes, worn Ubik chest logo, netting, and previously donned old coat and hat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 깨진 화병 조각 위로 발끝을 슬며시 밀고 있는 mid-action 자세의 찰리 발밑 클로즈업.\n\nLOCATION (lock): On the floor of the container home's shared living area, beside broken vase fragments in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Broken vase fragments (Being pushed along the floor by 찰리's foot); used as A small scattered group ahead of the toe makes the concealment attempt readable without exaggerated fragment scale; Container interior floor (Supporting 찰리's feet and the broken fragments) — Viewed at a downward oblique angle from beside his feet; used as Provides uninterrupted spatial context for the direction and extent of the sliding action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained daytime ambient illumination appropriate to the container interior, preserving readable floor-level detail without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments are on the container floor, being pushed aside rather than removed. Charlie retains his dirty aged gorilla-shaped casing, blue-lit eyes, worn Ubik chest logo, netting, and previously donned old coat and hat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "오른발(화면 좌측)의 발끝이 깨진 화병 조각을 향해 있으며 그 위에 얹혀 있음.",
    "built_space": "나무 판자 바닥과 좌측의 회색 문, 우측 후경의 의자 등 레퍼런스의 컨테이너 내부 구조와 일치하는 공간이 보임.",
    "entities": "찰리의 기계식 발, 샌드 베이지 장갑판, 낡은 코트 자락이 보이며 참조 이미지와 일치함. 바닥에 깨진 화병 조각이 있음.",
    "hard_violations": [],
    "physics": "찰리의 두 발은 바닥을 딛고 있으며, 오른발은 화병 조각 위를 살짝 밟고 지탱하는 안정적인 자세임. 우측 끝에 길게 늘어진 팔의 끝부분이 자연스럽게 매달려 있음."
   },
   {
    "label": "B",
    "direction": "오른발(화면 좌측)의 발끝이 아래를 향해 깨진 화병 조각들을 밀어내는 방향을 가리킴.",
    "built_space": "배경에 의자들이 보이나, 바닥 재질이 레퍼런스의 나무 판자가 아닌 매끄러운 단일 표면(콘크리트 또는 장판)으로 되어 있음.",
    "entities": "찰리의 기계식 발과 코트 자락이 참조 이미지의 디자인과 일치하게 나타남. 발밑에 깨진 화병 조각들이 흩어져 있음.",
    "hard_violations": [],
    "physics": "왼발은 바닥에 평평하게 닿아 있고, 오른발은 발뒤꿈치를 들고 발끝으로 조각을 미는 물리적으로 가능한 동적인 자세를 취함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "레퍼런스 이미지의 나무 바닥과 좌측 문 등 공간의 특징을 잘 유지했으며, 지시된 클로즈업 앵글과 캐릭터의 디테일을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "발끝으로 조각을 미는 동작은 자연스러우나, 바닥의 재질이 레퍼런스에 제시된 나무 바닥과 전혀 일치하지 않아 로케이션 고정 조건을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른발(화면 좌측)의 발끝이 깨진 화병 조각을 향해 있으며 그 위에 얹혀 있음.",
        "built_space": "나무 판자 바닥과 좌측의 회색 문, 우측 후경의 의자 등 레퍼런스의 컨테이너 내부 구조와 일치하는 공간이 보임.",
        "entities": "찰리의 기계식 발, 샌드 베이지 장갑판, 낡은 코트 자락이 보이며 참조 이미지와 일치함. 바닥에 깨진 화병 조각이 있음.",
        "hard_violations": [],
        "physics": "찰리의 두 발은 바닥을 딛고 있으며, 오른발은 화병 조각 위를 살짝 밟고 지탱하는 안정적인 자세임. 우측 끝에 길게 늘어진 팔의 끝부분이 자연스럽게 매달려 있음."
       },
       {
        "label": "B",
        "direction": "오른발(화면 좌측)의 발끝이 아래를 향해 깨진 화병 조각들을 밀어내는 방향을 가리킴.",
        "built_space": "배경에 의자들이 보이나, 바닥 재질이 레퍼런스의 나무 판자가 아닌 매끄러운 단일 표면(콘크리트 또는 장판)으로 되어 있음.",
        "entities": "찰리의 기계식 발과 코트 자락이 참조 이미지의 디자인과 일치하게 나타남. 발밑에 깨진 화병 조각들이 흩어져 있음.",
        "hard_violations": [],
        "physics": "왼발은 바닥에 평평하게 닿아 있고, 오른발은 발뒤꿈치를 들고 발끝으로 조각을 미는 물리적으로 가능한 동적인 자세를 취함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "레퍼런스 이미지의 나무 바닥과 좌측 문 등 공간의 특징을 잘 유지했으며, 지시된 클로즈업 앵글과 캐릭터의 디테일을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "발끝으로 조각을 미는 동작은 자연스러우나, 바닥의 재질이 레퍼런스에 제시된 나무 바닥과 전혀 일치하지 않아 로케이션 고정 조건을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "오른발(화면 좌측)의 발끝이 깨진 화병 조각을 향해 있으며 그 위에 얹혀 있음.",
        "built_space": "나무 판자 바닥과 좌측의 회색 문, 우측 후경의 의자 등 레퍼런스의 컨테이너 내부 구조와 일치하는 공간이 보임.",
        "entities": "찰리의 기계식 발, 샌드 베이지 장갑판, 낡은 코트 자락이 보이며 참조 이미지와 일치함. 바닥에 깨진 화병 조각이 있음.",
        "hard_violations": [],
        "physics": "찰리의 두 발은 바닥을 딛고 있으며, 오른발은 화병 조각 위를 살짝 밟고 지탱하는 안정적인 자세임. 우측 끝에 길게 늘어진 팔의 끝부분이 자연스럽게 매달려 있음."
       },
       {
        "label": "B",
        "direction": "오른발(화면 좌측)의 발끝이 아래를 향해 깨진 화병 조각들을 밀어내는 방향을 가리킴.",
        "built_space": "배경에 의자들이 보이나, 바닥 재질이 레퍼런스의 나무 판자가 아닌 매끄러운 단일 표면(콘크리트 또는 장판)으로 되어 있음.",
        "entities": "찰리의 기계식 발과 코트 자락이 참조 이미지의 디자인과 일치하게 나타남. 발밑에 깨진 화병 조각들이 흩어져 있음.",
        "hard_violations": [],
        "physics": "왼발은 바닥에 평평하게 닿아 있고, 오른발은 발뒤꿈치를 들고 발끝으로 조각을 미는 물리적으로 가능한 동적인 자세를 취함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "발 옆에서 내려다보는 클로즈업과 파편에 직접 닿아 미는 발끝이 핵심 동작에 더 충실하지만, 손잡이 달린 화병 잔해는 다소 주전자처럼 보인다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "찰리의 외장과 작은 파편 무리는 맞지만, 낮은 정면 시점과 파편 위로 들린 발끝 때문에 슬며시 밀어내는 순간보다 밟기 직전처럼 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 중앙의 동작하는 발끝은 왼쪽 아래의 파편 무리를 향하며, 앞쪽의 휘어진 조각에 직접 닿아 있다. 파편을 바닥을 따라 왼쪽 앞으로 밀려는 방향이 읽힌다. 다른 발도 대체로 같은 방향을 향한다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "마모된 회색 바닥이 두 발과 파편 사이에 연속적으로 보이고, 왼쪽에는 세로 골이 있는 벽과 의자 한 개, 뒤쪽에는 일부 가구 다리가 보인다. 발 옆의 비스듬한 하향 시점으로 파편의 이동 공간을 보여준다. 참고 장소의 낡은 컨테이너 내장과는 어울리지만, 참고 사진에서 바닥 노출이 적어 정확한 바닥 재질 일치까지 확정하기는 어렵다. 중복 설비나 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 두 하퇴와 기계식 발, 낡은 코트 자락이 보인다. 긁힌 샌드 베이지 장갑판과 분절된 발가락은 캐릭터 참고와 부합한다. 얼굴·눈·모자·가슴 표식은 화면 밖이므로 평가 대상이 아니다. 옅은 청록색 도자기 파편은 발보다 작고 앞쪽에 모여 있으나, 손잡이와 긴 목이 남은 큰 잔해는 일반적인 화병보다 주전자에 가깝게 보인다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "오른쪽의 지지 발은 바닥에 놓여 있고, 동작하는 발도 뒤쪽과 옆쪽 접지부로 지지된다. 앞 발가락은 바닥에 놓인 곡면 파편과 접촉하므로 미는 힘을 전달할 수 있다. 세워진 듯한 곡면 조각도 바닥과 발끝 사이에서 지지되며, 나머지 잔해는 바닥에 놓여 있다. 근거 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "앞쪽 발끝은 화면 오른쪽 아래의 파편 무리를 향하지만, 발가락 밑면이 파편 위에 들려 있어 수평으로 밀어내는 접촉은 명확하지 않다. 반대 발은 오른쪽 뒤에서 같은 쪽을 향한다. 파편을 슬쩍 밀기보다는 위에서 밟으려는 방향으로 읽히며, 시선은 화면 밖이다.",
        "built_space": "널빤지 이음과 나뭇결이 뚜렷한 바닥, 왼쪽 벽의 문 한 개, 오른쪽 의자 한 개, 먼 뒤쪽의 작은 선반이 보인다. 발 앞쪽의 매우 낮은 시점이라 요청된 발 옆 하향 사선보다 정면에 가깝고, 다리와 코트가 상대적으로 많이 들어온다. 낡은 컨테이너 분위기는 유지하지만 참고에서 확인되지 않는 목재 바닥 무늬가 두드러진다. 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 장갑 다리 두 개와 기계식 발, 낡은 코트가 보이며, 베이지 외장과 발 관절은 참고 정체성에 부합한다. 얼굴과 가슴의 세부는 프레임 밖이다. 발 앞에는 옅은 청록색 화병의 목·입구 잔해와 작은 곡면 파편들이 모여 있고, 크기도 과장되지 않았다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤쪽 발은 바닥에 접지하고, 앞쪽 발은 뒤꿈치 쪽을 바닥에 둔 채 발끝을 들어 올린 모습으로 읽힌다. 따라서 몸이나 발이 무지지 상태로 떠 있는 것은 아니다. 도자기 잔해도 바닥에 지지된다. 다만 발끝과 파편 사이의 간격 때문에 바닥을 따라 파편을 밀고 있다는 힘의 전달은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "발 옆에서 내려다보는 클로즈업과 파편에 직접 닿아 미는 발끝이 핵심 동작에 더 충실하지만, 손잡이 달린 화병 잔해는 다소 주전자처럼 보인다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "찰리의 외장과 작은 파편 무리는 맞지만, 낮은 정면 시점과 파편 위로 들린 발끝 때문에 슬며시 밀어내는 순간보다 밟기 직전처럼 읽힌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 중앙의 동작하는 발끝은 왼쪽 아래의 파편 무리를 향하며, 앞쪽의 휘어진 조각에 직접 닿아 있다. 파편을 바닥을 따라 왼쪽 앞으로 밀려는 방향이 읽힌다. 다른 발도 대체로 같은 방향을 향한다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "마모된 회색 바닥이 두 발과 파편 사이에 연속적으로 보이고, 왼쪽에는 세로 골이 있는 벽과 의자 한 개, 뒤쪽에는 일부 가구 다리가 보인다. 발 옆의 비스듬한 하향 시점으로 파편의 이동 공간을 보여준다. 참고 장소의 낡은 컨테이너 내장과는 어울리지만, 참고 사진에서 바닥 노출이 적어 정확한 바닥 재질 일치까지 확정하기는 어렵다. 중복 설비나 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 두 하퇴와 기계식 발, 낡은 코트 자락이 보인다. 긁힌 샌드 베이지 장갑판과 분절된 발가락은 캐릭터 참고와 부합한다. 얼굴·눈·모자·가슴 표식은 화면 밖이므로 평가 대상이 아니다. 옅은 청록색 도자기 파편은 발보다 작고 앞쪽에 모여 있으나, 손잡이와 긴 목이 남은 큰 잔해는 일반적인 화병보다 주전자에 가깝게 보인다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "오른쪽의 지지 발은 바닥에 놓여 있고, 동작하는 발도 뒤쪽과 옆쪽 접지부로 지지된다. 앞 발가락은 바닥에 놓인 곡면 파편과 접촉하므로 미는 힘을 전달할 수 있다. 세워진 듯한 곡면 조각도 바닥과 발끝 사이에서 지지되며, 나머지 잔해는 바닥에 놓여 있다. 근거 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "앞쪽 발끝은 화면 오른쪽 아래의 파편 무리를 향하지만, 발가락 밑면이 파편 위에 들려 있어 수평으로 밀어내는 접촉은 명확하지 않다. 반대 발은 오른쪽 뒤에서 같은 쪽을 향한다. 파편을 슬쩍 밀기보다는 위에서 밟으려는 방향으로 읽히며, 시선은 화면 밖이다.",
        "built_space": "널빤지 이음과 나뭇결이 뚜렷한 바닥, 왼쪽 벽의 문 한 개, 오른쪽 의자 한 개, 먼 뒤쪽의 작은 선반이 보인다. 발 앞쪽의 매우 낮은 시점이라 요청된 발 옆 하향 사선보다 정면에 가깝고, 다리와 코트가 상대적으로 많이 들어온다. 낡은 컨테이너 분위기는 유지하지만 참고에서 확인되지 않는 목재 바닥 무늬가 두드러진다. 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 장갑 다리 두 개와 기계식 발, 낡은 코트가 보이며, 베이지 외장과 발 관절은 참고 정체성에 부합한다. 얼굴과 가슴의 세부는 프레임 밖이다. 발 앞에는 옅은 청록색 화병의 목·입구 잔해와 작은 곡면 파편들이 모여 있고, 크기도 과장되지 않았다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤쪽 발은 바닥에 접지하고, 앞쪽 발은 뒤꿈치 쪽을 바닥에 둔 채 발끝을 들어 올린 모습으로 읽힌다. 따라서 몸이나 발이 무지지 상태로 떠 있는 것은 아니다. 도자기 잔해도 바닥에 지지된다. 다만 발끝과 파편 사이의 간격 때문에 바닥을 따라 파편을 밀고 있다는 힘의 전달은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.667
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스 이미지의 나무 바닥과 좌측 문 등 공간의 특징을 잘 유지했으며, 지시된 클로즈업 앵글과 캐릭터의 디테일을 충실히 구현했습니다."
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "발끝으로 조각을 미는 동작은 자연스러우나, 바닥의 재질이 레퍼런스에 제시된 나무 바닥과 전혀 일치하지 않아 로케이션 고정 조건을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh12_sel.png",
    "asset_id": "5a3c7b67-f591-40f4-a60b-c97969b6d232",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-cdc2-78d4-88f3-dd92c680a5fd",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S12sh12"
  }
 },
 "S14sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:30:22.524448+00:00",
  "fingerprint": "f8545f669f7823ff3933a7ff7e3a2c6f7b0575e972674fa8154f428e076a536d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S14sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S14sh3_sel.png",
  "source_sha256": "8a5dd070063f3e16216b91e213afcef51b175ce3bc871f0712162e21ce21f9c9",
  "file": "S14sh3_cine.png",
  "staged_sha256": "bb2fa4cf8c875a0ff02fc13e02342535de870f2c7d52b1377de53c26c3700ab4",
  "latency_ms": 12425
 },
 "S14sh5::signage": {
  "fp": "dd22c80933257b71",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S14sh5": {
  "input_fingerprint": "ed7ff72a787e519b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양팔을 벌린 채 앞을 가로막고 서 있는 앰버의 굳은 전신.\n\nLOCATION (lock): In the shared living area inside the refugee family's container home, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container interior (Interior of 현우's home) — The floor and interior boundaries remain visible around Amber's full figure; used as Establish the space she is blocking without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination with controlled contrast, preserving the tension in her face without introducing a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments remain on the container floor, partly pushed aside rather than removed. Charlie is an old, dirt-covered gorilla-shaped robot with blue-lit eyes and a worn Ubik chest logo, still wearing the old coat and hat used as a disguise. 앰버: She still wears her mask and waist tool pouch and has a recurring cough.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양팔을 벌린 채 앞을 가로막고 서 있는 앰버의 굳은 전신.\n\nLOCATION (lock): In the shared living area inside the refugee family's container home, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container interior (Interior of 현우's home) — The floor and interior boundaries remain visible around Amber's full figure; used as Establish the space she is blocking without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination with controlled contrast, preserving the tension in her face without introducing a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments remain on the container floor, partly pushed aside rather than removed. Charlie is an old, dirt-covered gorilla-shaped robot with blue-lit eyes and a worn Ubik chest logo, still wearing the old coat and hat used as a disguise. 앰버: She still wears her mask and waist tool pouch and has a recurring cough.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양팔을 벌린 채 앞을 가로막고 서 있는 앰버의 굳은 전신.\n\nLOCATION (lock): In the shared living area inside the refugee family's container home, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container interior (Interior of 현우's home) — The floor and interior boundaries remain visible around Amber's full figure; used as Establish the space she is blocking without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination with controlled contrast, preserving the tension in her face without introducing a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments remain on the container floor, partly pushed aside rather than removed. Charlie is an old, dirt-covered gorilla-shaped robot with blue-lit eyes and a worn Ubik chest logo, still wearing the old coat and hat used as a disguise. 앰버: She still wears her mask and waist tool pouch and has a recurring cough.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버의 몸과 시선은 카메라 쪽 정면을 향하고, 두 팔과 펼친 손은 통로 양옆으로 뻗어 있다. 화면 밖 카메라 쪽에 있는 현우를 막는 배치로 읽히지만, 현우가 보이지 않아 실제 시선 도착점까지 확인할 수는 없다.",
    "built_space": "중앙 통로에 앰버가 서 있고 머리부터 부츠까지, 양옆 벽과 천장 및 발 주변 바닥이 보인다. 왼쪽 회색 문 1개, 뒤쪽 냉장고 1대와 수납 선반 1개, 오른쪽 벤치 1개와 좁은 수납장 1개 및 화면 가장자리의 높은 목제 장 1개가 보인다. 참고의 낡은 목재 바닥, 골진 금속 벽, 왼쪽 문과 오른쪽 벤치 배치는 대체로 유지된다. 전경에는 몸을 가리는 장애물이 없으며 반사상도 없다.",
    "entities": "금발의 어린 여자아이 1명만 등장하며 체격, 드러난 둥근 얼굴과 큰 눈은 앰버 참고에 가깝다. 한국계 백인 혼혈이라는 설정과 외형상 뚜렷한 충돌은 없다. 남색 반소매, 낡은 갈색 멜빵 작업복, 부츠, 허리 공구 주머니가 유지된다. 다만 마스크는 얼굴이 아니라 이마에 올라가 있다. 깨진 옅은 청록색 꽃병과 파편이 발 앞 오른쪽에 남아 있다. 현우나 찰리 및 다른 인물은 없으며, 이는 앰버만 보여 주는 이번 숏에 맞는다. 읽을 수 있는 글자는 보이지 않는다.",
    "hard_violations": [],
    "physics": "양쪽 부츠가 바닥에 닿아 몸을 지탱하며, 어깨에서 양팔을 벌린 자세는 실제로 유지할 수 있다. 공구와 주머니는 허리띠에 걸려 있고 마스크는 머리에 얹혀 있다. 꽃병 몸통과 파편은 바닥에 놓여 있어 지지 없는 물체는 없다."
   },
   {
    "label": "A",
    "direction": "앰버는 정면의 카메라 쪽을 응시하며 두 팔을 좌우로 거의 수평으로 뻗어 통로를 막는다. 현우는 화면에 없지만 몸과 시선이 동일한 전방을 향해, 화면 밖 현우를 가로막는 관계가 자연스럽게 읽힌다.",
    "built_space": "앰버의 전신 주위로 바닥, 좌우 금속 벽과 천장이 남는 와이드 구도다. 왼쪽 회색 문 1개, 뒤쪽 수납 선반 1개와 밝은 소형 냉장고 1대, 후면의 열린 문 1개, 오른쪽 작업대 1개와 벤치 1개, 가장자리 목제 장 1개가 보인다. 참고의 바닥 마모와 금속 벽, 문과 벤치의 기본 위치는 이어지지만 뒤쪽 가구 배치는 완전히 같다고 확인하기 어렵다. 앰버는 가구가 아닌 빈 중앙 통로를 막고 있으며 전경 가림이나 불가능한 반사는 없다.",
    "entities": "금발의 어린 여자아이 1명으로, 참고의 앰버와 머리색 및 체격이 가깝다. 마스크가 얼굴 대부분을 가려 얼굴 윤곽과 혼혈 정체성의 세부 일치는 확인이 제한되지만, 보이는 눈과 이마에 뚜렷한 불일치는 없다. 얼굴을 덮는 호흡 마스크, 남색 반소매, 낡은 갈색 멜빵 작업복, 부츠와 허리 공구 주머니를 착용한다. 꽃병 몸통과 파편은 바닥 오른쪽으로 모여 남아 있다. 다른 인물이나 찰리는 등장하지 않으며 이번 숏의 인물 범위에 맞는다. 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "두 발이 바닥에 붙어 벌어진 다리로 체중을 받치며, 양팔을 수평으로 든 정지 자세도 해부학적으로 가능하다. 마스크는 머리끈으로 고정되고 공구 주머니와 공구는 허리띠에 지지된다. 꽃병 파편과 주변 가구는 바닥이나 작업대 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신과 주변 바닥을 담은 와이드 구도와 통로를 막는 동작은 맞지만, 마스크를 이마에 올려 얼굴을 노출한 점이 착용 유지 지시와 어긋난다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "마스크와 허리 공구 주머니를 착용한 앰버가 양팔을 벌려 통로를 막는 굳은 전신을 보여 주어 핵심 구도와 행동, 착용 상태를 가장 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 몸과 시선은 카메라 쪽 정면을 향하고, 두 팔과 펼친 손은 통로 양옆으로 뻗어 있다. 화면 밖 카메라 쪽에 있는 현우를 막는 배치로 읽히지만, 현우가 보이지 않아 실제 시선 도착점까지 확인할 수는 없다.",
        "built_space": "중앙 통로에 앰버가 서 있고 머리부터 부츠까지, 양옆 벽과 천장 및 발 주변 바닥이 보인다. 왼쪽 회색 문 1개, 뒤쪽 냉장고 1대와 수납 선반 1개, 오른쪽 벤치 1개와 좁은 수납장 1개 및 화면 가장자리의 높은 목제 장 1개가 보인다. 참고의 낡은 목재 바닥, 골진 금속 벽, 왼쪽 문과 오른쪽 벤치 배치는 대체로 유지된다. 전경에는 몸을 가리는 장애물이 없으며 반사상도 없다.",
        "entities": "금발의 어린 여자아이 1명만 등장하며 체격, 드러난 둥근 얼굴과 큰 눈은 앰버 참고에 가깝다. 한국계 백인 혼혈이라는 설정과 외형상 뚜렷한 충돌은 없다. 남색 반소매, 낡은 갈색 멜빵 작업복, 부츠, 허리 공구 주머니가 유지된다. 다만 마스크는 얼굴이 아니라 이마에 올라가 있다. 깨진 옅은 청록색 꽃병과 파편이 발 앞 오른쪽에 남아 있다. 현우나 찰리 및 다른 인물은 없으며, 이는 앰버만 보여 주는 이번 숏에 맞는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "양쪽 부츠가 바닥에 닿아 몸을 지탱하며, 어깨에서 양팔을 벌린 자세는 실제로 유지할 수 있다. 공구와 주머니는 허리띠에 걸려 있고 마스크는 머리에 얹혀 있다. 꽃병 몸통과 파편은 바닥에 놓여 있어 지지 없는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "앰버는 정면의 카메라 쪽을 응시하며 두 팔을 좌우로 거의 수평으로 뻗어 통로를 막는다. 현우는 화면에 없지만 몸과 시선이 동일한 전방을 향해, 화면 밖 현우를 가로막는 관계가 자연스럽게 읽힌다.",
        "built_space": "앰버의 전신 주위로 바닥, 좌우 금속 벽과 천장이 남는 와이드 구도다. 왼쪽 회색 문 1개, 뒤쪽 수납 선반 1개와 밝은 소형 냉장고 1대, 후면의 열린 문 1개, 오른쪽 작업대 1개와 벤치 1개, 가장자리 목제 장 1개가 보인다. 참고의 바닥 마모와 금속 벽, 문과 벤치의 기본 위치는 이어지지만 뒤쪽 가구 배치는 완전히 같다고 확인하기 어렵다. 앰버는 가구가 아닌 빈 중앙 통로를 막고 있으며 전경 가림이나 불가능한 반사는 없다.",
        "entities": "금발의 어린 여자아이 1명으로, 참고의 앰버와 머리색 및 체격이 가깝다. 마스크가 얼굴 대부분을 가려 얼굴 윤곽과 혼혈 정체성의 세부 일치는 확인이 제한되지만, 보이는 눈과 이마에 뚜렷한 불일치는 없다. 얼굴을 덮는 호흡 마스크, 남색 반소매, 낡은 갈색 멜빵 작업복, 부츠와 허리 공구 주머니를 착용한다. 꽃병 몸통과 파편은 바닥 오른쪽으로 모여 남아 있다. 다른 인물이나 찰리는 등장하지 않으며 이번 숏의 인물 범위에 맞는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 발이 바닥에 붙어 벌어진 다리로 체중을 받치며, 양팔을 수평으로 든 정지 자세도 해부학적으로 가능하다. 마스크는 머리끈으로 고정되고 공구 주머니와 공구는 허리띠에 지지된다. 꽃병 파편과 주변 가구는 바닥이나 작업대 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "전신과 주변 바닥을 담은 와이드 구도와 통로를 막는 동작은 맞지만, 마스크를 이마에 올려 얼굴을 노출한 점이 착용 유지 지시와 어긋난다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "마스크와 허리 공구 주머니를 착용한 앰버가 양팔을 벌려 통로를 막는 굳은 전신을 보여 주어 핵심 구도와 행동, 착용 상태를 가장 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 몸과 시선은 카메라 쪽 정면을 향하고, 두 팔과 펼친 손은 통로 양옆으로 뻗어 있다. 화면 밖 카메라 쪽에 있는 현우를 막는 배치로 읽히지만, 현우가 보이지 않아 실제 시선 도착점까지 확인할 수는 없다.",
        "built_space": "중앙 통로에 앰버가 서 있고 머리부터 부츠까지, 양옆 벽과 천장 및 발 주변 바닥이 보인다. 왼쪽 회색 문 1개, 뒤쪽 냉장고 1대와 수납 선반 1개, 오른쪽 벤치 1개와 좁은 수납장 1개 및 화면 가장자리의 높은 목제 장 1개가 보인다. 참고의 낡은 목재 바닥, 골진 금속 벽, 왼쪽 문과 오른쪽 벤치 배치는 대체로 유지된다. 전경에는 몸을 가리는 장애물이 없으며 반사상도 없다.",
        "entities": "금발의 어린 여자아이 1명만 등장하며 체격, 드러난 둥근 얼굴과 큰 눈은 앰버 참고에 가깝다. 한국계 백인 혼혈이라는 설정과 외형상 뚜렷한 충돌은 없다. 남색 반소매, 낡은 갈색 멜빵 작업복, 부츠, 허리 공구 주머니가 유지된다. 다만 마스크는 얼굴이 아니라 이마에 올라가 있다. 깨진 옅은 청록색 꽃병과 파편이 발 앞 오른쪽에 남아 있다. 현우나 찰리 및 다른 인물은 없으며, 이는 앰버만 보여 주는 이번 숏에 맞는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "양쪽 부츠가 바닥에 닿아 몸을 지탱하며, 어깨에서 양팔을 벌린 자세는 실제로 유지할 수 있다. 공구와 주머니는 허리띠에 걸려 있고 마스크는 머리에 얹혀 있다. 꽃병 몸통과 파편은 바닥에 놓여 있어 지지 없는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "앰버는 정면의 카메라 쪽을 응시하며 두 팔을 좌우로 거의 수평으로 뻗어 통로를 막는다. 현우는 화면에 없지만 몸과 시선이 동일한 전방을 향해, 화면 밖 현우를 가로막는 관계가 자연스럽게 읽힌다.",
        "built_space": "앰버의 전신 주위로 바닥, 좌우 금속 벽과 천장이 남는 와이드 구도다. 왼쪽 회색 문 1개, 뒤쪽 수납 선반 1개와 밝은 소형 냉장고 1대, 후면의 열린 문 1개, 오른쪽 작업대 1개와 벤치 1개, 가장자리 목제 장 1개가 보인다. 참고의 바닥 마모와 금속 벽, 문과 벤치의 기본 위치는 이어지지만 뒤쪽 가구 배치는 완전히 같다고 확인하기 어렵다. 앰버는 가구가 아닌 빈 중앙 통로를 막고 있으며 전경 가림이나 불가능한 반사는 없다.",
        "entities": "금발의 어린 여자아이 1명으로, 참고의 앰버와 머리색 및 체격이 가깝다. 마스크가 얼굴 대부분을 가려 얼굴 윤곽과 혼혈 정체성의 세부 일치는 확인이 제한되지만, 보이는 눈과 이마에 뚜렷한 불일치는 없다. 얼굴을 덮는 호흡 마스크, 남색 반소매, 낡은 갈색 멜빵 작업복, 부츠와 허리 공구 주머니를 착용한다. 꽃병 몸통과 파편은 바닥 오른쪽으로 모여 남아 있다. 다른 인물이나 찰리는 등장하지 않으며 이번 숏의 인물 범위에 맞는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 발이 바닥에 붙어 벌어진 다리로 체중을 받치며, 양팔을 수평으로 든 정지 자세도 해부학적으로 가능하다. 마스크는 머리끈으로 고정되고 공구 주머니와 공구는 허리띠에 지지된다. 꽃병 파편과 주변 가구는 바닥이나 작업대 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 7,
   "A": 9
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 7,
    "verdict_ko": "전신과 주변 바닥을 담은 와이드 구도와 통로를 막는 동작은 맞지만, 마스크를 이마에 올려 얼굴을 노출한 점이 착용 유지 지시와 어긋난다."
   },
   {
    "label": "A",
    "score": 9,
    "verdict_ko": "마스크와 허리 공구 주머니를 착용한 앰버가 양팔을 벌려 통로를 막는 굳은 전신을 보여 주어 핵심 구도와 행동, 착용 상태를 가장 충실히 구현한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S14sh3_sel.png",
    "asset_id": "8c1e52ed-b248-48d8-971a-f33b5a485d75",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-cf69-7f38-9876-70a2f530ab26",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S14sh3"
  }
 },
 "S14sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:31:20.504434+00:00",
  "fingerprint": "2a60c1298f2e7d74f2ff98f29fc23cc1afb640f044a2d67cfcc3c7aa627aa675",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S14sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S14sh5_sel.png",
  "source_sha256": "060d888583e786c00e538482a123c37bea0c7f73313f91ea6d5ad506f767a64a",
  "file": "S14sh5_cine.png",
  "staged_sha256": "8311d52f546d7e577065deb9989201dd5d01b7abc6f32b96ce75b6aa32962f94",
  "latency_ms": 11339
 },
 "S14sh9::signage": {
  "fp": "5602a8de271d82eb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S14sh9": {
  "input_fingerprint": "a750077782eb8938",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 험상궂은 표정으로 찰리의 팔을 거칠게 움켜쥔 현우의 밀착된 상체.\n\nLOCATION (lock): Inside the container home's shared living area, at the spot where the robot has been handling household objects in daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior (Same interior as the blocking exchange) — Only a narrow portion of the interior remains behind the two upper bodies; used as Maintain location continuity without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the container's neutral ambient illumination and restrained contrast, keeping the face and gripping hand equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same container interior, daylight, and household fixtures, including the broken vase fragments that remain on the floor. Exclude an intact replacement vase and any shards suspended in the air.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The broken vase fragments remain on the floor where they were pushed aside. Charlie retains his dirty, worn gorilla-shaped body, blue-lit eyes, worn Ubik chest logo, old coat and hat. 현우: He is now standing after rising from his seat, with his outer shirt on. His facial injuries and untreated dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 험상궂은 표정으로 찰리의 팔을 거칠게 움켜쥔 현우의 밀착된 상체.\n\nLOCATION (lock): Inside the container home's shared living area, at the spot where the robot has been handling household objects in daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior (Same interior as the blocking exchange) — Only a narrow portion of the interior remains behind the two upper bodies; used as Maintain location continuity without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the container's neutral ambient illumination and restrained contrast, keeping the face and gripping hand equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same container interior, daylight, and household fixtures, including the broken vase fragments that remain on the floor. Exclude an intact replacement vase and any shards suspended in the air.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The broken vase fragments remain on the floor where they were pushed aside. Charlie retains his dirty, worn gorilla-shaped body, blue-lit eyes, worn Ubik chest logo, old coat and hat. 현우: He is now standing after rising from his seat, with his outer shirt on. His facial injuries and untreated dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 험상궂은 표정으로 찰리의 팔을 거칠게 움켜쥔 현우의 밀착된 상체.\n\nLOCATION (lock): Inside the container home's shared living area, at the spot where the robot has been handling household objects in daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior (Same interior as the blocking exchange) — Only a narrow portion of the interior remains behind the two upper bodies; used as Maintain location continuity without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the container's neutral ambient illumination and restrained contrast, keeping the face and gripping hand equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same container interior, daylight, and household fixtures, including the broken vase fragments that remain on the floor. Exclude an intact replacement vase and any shards suspended in the air.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The broken vase fragments remain on the floor where they were pushed aside. Charlie retains his dirty, worn gorilla-shaped body, blue-lit eyes, worn Ubik chest logo, old coat and hat. 현우: He is now standing after rising from his seat, with his outer shirt on. His facial injuries and untreated dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우는 매서운 표정으로 정면 약간 측면을 응시하며 찰리의 팔을 강하게 쥐고 있고, 찰리는 현우 쪽으로 고개를 숙여 내려다보고 있다.",
    "built_space": "컨테이너 내부의 벽면과 문, 파이프가 보이며 바닥에 깨진 도자기 조각이 놓여 있다. 그러나 두 인물 사이의 배경 바닥에 있어야 할 이유가 없는 부츠가 놓여 있다.",
    "entities": "현우의 얼굴 묘사와 상처, 옷차림은 일치하나, 찰리의 머리에 있어야 할 모자와 가슴의 Ubik 로고가 누락되었다. 또한 배경에 프롬프트에 없는 정체불명의 다리(이전 샷 인물의 잔재)가 등장한다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에서 제외를 명시한 이전 샷 인물의 신체 일부(배경에 서 있는 부츠 신은 두 다리)가 화면에 나타남.",
     "[gpt-high] 두 주인공 뒤 바닥에 제삼자의 바지 다리와 부츠가 남아 있어, 현우와 찰리 외 인물 및 이전 장면 인물의 신체를 배제하라는 지시를 위반한다."
    ],
    "physics": "현우가 찰리의 팔을 잡는 동작 자체는 성립하나, 배경에 등장한 다리는 상체나 다른 지지 구조 없이 바닥에 덩그러니 서 있어 물리적으로 불가능한 상태이다."
   },
   {
    "label": "B",
    "direction": "현우는 얼굴을 찌푸린 채 찰리의 팔뚝과 가슴 쪽으로 강한 시선을 던지며, 양손으로 팔을 거칠게 움켜쥐고 있다.",
    "built_space": "컨테이너 내부의 창문과 가구들이 자연스럽게 배치되어 있으며, 인물들 뒤편 바닥에 깨진 도자기 조각들이 지시된 대로 위치해 있다.",
    "entities": "현우의 앳된 얼굴, 헝클어진 머리, 얼굴의 상처가 지시사항과 완벽히 일치한다. 찰리 역시 낡은 코트와 모자(챙 부분 노출), 모래색 장갑판, 그리고 가슴에 새겨진 Ubik 로고까지 정확하게 반영되어 있다.",
    "hard_violations": [
     "[gpt-high] 찰리의 가슴에 판독 가능한 'Ubik' 문자와 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
    ],
    "physics": "현우가 찰리의 육중한 팔에 매달리듯 힘을 주어 움켜쥔 양손의 근육과 자세가 물리적 무게감을 잘 전달하며 자연스럽게 지탱되고 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우가 찰리의 팔을 거칠게 움켜쥔 텐션을 클로즈업 샷으로 사실적으로 묘사했으며, 찰리의 모자, Ubik 로고, 바닥의 깨진 도자기 등 세부 디테일을 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 인물 사이의 배경 바닥에 프롬프트에서 제외해야 할 이전 샷 인물의 부츠 신은 다리가 남아있는 치명적인 오류가 발생했으며, 찰리의 모자와 로고도 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 매서운 표정으로 정면 약간 측면을 응시하며 찰리의 팔을 강하게 쥐고 있고, 찰리는 현우 쪽으로 고개를 숙여 내려다보고 있다.",
        "built_space": "컨테이너 내부의 벽면과 문, 파이프가 보이며 바닥에 깨진 도자기 조각이 놓여 있다. 그러나 두 인물 사이의 배경 바닥에 있어야 할 이유가 없는 부츠가 놓여 있다.",
        "entities": "현우의 얼굴 묘사와 상처, 옷차림은 일치하나, 찰리의 머리에 있어야 할 모자와 가슴의 Ubik 로고가 누락되었다. 또한 배경에 프롬프트에 없는 정체불명의 다리(이전 샷 인물의 잔재)가 등장한다.",
        "hard_violations": [
         "프롬프트에서 제외를 명시한 이전 샷 인물의 신체 일부(배경에 서 있는 부츠 신은 두 다리)가 화면에 나타남."
        ],
        "physics": "현우가 찰리의 팔을 잡는 동작 자체는 성립하나, 배경에 등장한 다리는 상체나 다른 지지 구조 없이 바닥에 덩그러니 서 있어 물리적으로 불가능한 상태이다."
       },
       {
        "label": "B",
        "direction": "현우는 얼굴을 찌푸린 채 찰리의 팔뚝과 가슴 쪽으로 강한 시선을 던지며, 양손으로 팔을 거칠게 움켜쥐고 있다.",
        "built_space": "컨테이너 내부의 창문과 가구들이 자연스럽게 배치되어 있으며, 인물들 뒤편 바닥에 깨진 도자기 조각들이 지시된 대로 위치해 있다.",
        "entities": "현우의 앳된 얼굴, 헝클어진 머리, 얼굴의 상처가 지시사항과 완벽히 일치한다. 찰리 역시 낡은 코트와 모자(챙 부분 노출), 모래색 장갑판, 그리고 가슴에 새겨진 Ubik 로고까지 정확하게 반영되어 있다.",
        "hard_violations": [],
        "physics": "현우가 찰리의 육중한 팔에 매달리듯 힘을 주어 움켜쥔 양손의 근육과 자세가 물리적 무게감을 잘 전달하며 자연스럽게 지탱되고 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우가 찰리의 팔을 거칠게 움켜쥔 텐션을 클로즈업 샷으로 사실적으로 묘사했으며, 찰리의 모자, Ubik 로고, 바닥의 깨진 도자기 등 세부 디테일을 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 인물 사이의 배경 바닥에 프롬프트에서 제외해야 할 이전 샷 인물의 부츠 신은 다리가 남아있는 치명적인 오류가 발생했으며, 찰리의 모자와 로고도 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 매서운 표정으로 정면 약간 측면을 응시하며 찰리의 팔을 강하게 쥐고 있고, 찰리는 현우 쪽으로 고개를 숙여 내려다보고 있다.",
        "built_space": "컨테이너 내부의 벽면과 문, 파이프가 보이며 바닥에 깨진 도자기 조각이 놓여 있다. 그러나 두 인물 사이의 배경 바닥에 있어야 할 이유가 없는 부츠가 놓여 있다.",
        "entities": "현우의 얼굴 묘사와 상처, 옷차림은 일치하나, 찰리의 머리에 있어야 할 모자와 가슴의 Ubik 로고가 누락되었다. 또한 배경에 프롬프트에 없는 정체불명의 다리(이전 샷 인물의 잔재)가 등장한다.",
        "hard_violations": [
         "프롬프트에서 제외를 명시한 이전 샷 인물의 신체 일부(배경에 서 있는 부츠 신은 두 다리)가 화면에 나타남."
        ],
        "physics": "현우가 찰리의 팔을 잡는 동작 자체는 성립하나, 배경에 등장한 다리는 상체나 다른 지지 구조 없이 바닥에 덩그러니 서 있어 물리적으로 불가능한 상태이다."
       },
       {
        "label": "B",
        "direction": "현우는 얼굴을 찌푸린 채 찰리의 팔뚝과 가슴 쪽으로 강한 시선을 던지며, 양손으로 팔을 거칠게 움켜쥐고 있다.",
        "built_space": "컨테이너 내부의 창문과 가구들이 자연스럽게 배치되어 있으며, 인물들 뒤편 바닥에 깨진 도자기 조각들이 지시된 대로 위치해 있다.",
        "entities": "현우의 앳된 얼굴, 헝클어진 머리, 얼굴의 상처가 지시사항과 완벽히 일치한다. 찰리 역시 낡은 코트와 모자(챙 부분 노출), 모래색 장갑판, 그리고 가슴에 새겨진 Ubik 로고까지 정확하게 반영되어 있다.",
        "hard_violations": [],
        "physics": "현우가 찰리의 육중한 팔에 매달리듯 힘을 주어 움켜쥔 양손의 근육과 자세가 물리적 무게감을 잘 전달하며 자연스럽게 지탱되고 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리를 향한 험상궂은 시선과 실제 팔 움켜쥠은 더 정확하지만, 구도가 요구된 밀착 클로즈업보다 넓고 읽히는 가슴 로고가 명시적 금지사항을 위반한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "상체와 움켜쥔 손의 밀착 구도는 더 가깝지만, 배경에 제삼자의 다리와 부츠가 남아 있고 현우가 찰리 대신 카메라를 응시한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸을 앞으로 숙이고 화면 오른쪽 찰리의 상체와 얼굴 쪽을 노려본다. 양손은 찰리의 전완 장갑을 위와 아래에서 붙잡는다. 찰리의 얼굴은 현우 쪽으로 내려가 있으나 눈은 상단 잘림으로 확인하기 어렵다. 무기나 이동 중인 물체는 없다.",
        "built_space": "왼쪽에 창 하나와 문 하나, 뒤쪽 오른편에 작업대 하나와 의자 일부, 작업대 아래 상자들이 보인다. 골진 금속 벽과 천장, 낡은 목재 바닥은 장소 참고와 대체로 이어진다. 두 인물 사이로 바닥과 긴 실내가 상당히 드러나므로 배경을 좁게 남기는 밀착 클로즈업보다 넓다. 바닥에는 깨진 화병 조각 한 무더기가 있다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 회색 겉셔츠, 얼굴 찰과상이 참고 및 지시와 대체로 맞는다. 찰리는 육중한 베이지색 기계 팔, 마모된 장갑, 낡은 코트와 모자 일부, 흰 기계 얼굴을 유지한다. 눈과 다리 부상은 구도상 확인할 수 없다. 추가 인물이나 온전한 대체 화병은 보이지 않는다. 다만 찰리 가슴의 'Ubik' 문자와 상징은 선명하게 읽힌다.",
        "hard_violations": [
         "찰리의 가슴에 판독 가능한 'Ubik' 문자와 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "현우의 한 손은 찰리 전완 장갑의 윗면을 감싸고 다른 손은 아래쪽을 받쳐 실제 접촉하며 움켜쥔다. 찰리의 팔은 팔꿈치와 어깨로 몸통에 연결되어 있고 현우의 손도 이를 지지한다. 두 몸통은 화면 아래로 이어져 있으며 공중에 떠 있다는 징후는 없다. 화병 조각들은 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "현우는 눈썹을 찌푸렸지만 시선은 찰리가 아니라 카메라 쪽을 향한다. 찰리는 머리를 왼쪽 아래, 현우와 맞잡힌 팔 쪽으로 기울인다. 현우의 손은 가슴 앞을 가로지르는 찰리의 전완을 움켜쥐고 있다. 팔 접촉 방향은 맞지만 현우의 응시 대상이 대치 상대와 어긋난다.",
        "built_space": "왼쪽 금속 벽에 세로 배관과 전기함, 문 하나가 보이고, 뒤로 골진 천장과 목재 바닥이 이어진다. 참고 장소의 왼쪽 벽 배치와 재질은 잘 유지된다. 두 상체와 전경의 팔이 화면 대부분을 차지해 A보다 밀착 구도에 가깝다. 하단에는 화병 조각 한 무더기와, 두 주인공 뒤 중앙에 별도의 바지 다리와 부츠 하나가 보인다. 중복 설비나 반사는 보이지 않는다.",
        "entities": "현우의 젊은 동아시아계 남성 외형, 검은 머리, 회색 셔츠와 얼굴 상처는 대체로 맞는다. 찰리의 베이지 장갑, 흰 기계 얼굴, 파란 눈과 낡은 코트가 보인다. 머리 윗부분은 잘렸지만 드러난 두상 주변에 요구된 모자의 챙은 보이지 않는다. 읽히는 가슴 문자는 없다. 뒤쪽 중앙의 갈색 바지와 부츠는 두 주인공의 위치 및 복장과 맞지 않는 제삼자의 신체 일부다. 다리 부상은 구도 밖이다.",
        "hard_violations": [
         "두 주인공 뒤 바닥에 제삼자의 바지 다리와 부츠가 남아 있어, 현우와 찰리 외 인물 및 이전 장면 인물의 신체를 배제하라는 지시를 위반한다."
        ],
        "physics": "찰리의 가로놓인 전완은 오른쪽 팔꿈치와 상완으로 이어지고, 현우의 손가락은 장갑과 소매 부위를 실제로 감싼다. 팔이 현우의 몸 앞에 밀착된 자세는 물리적으로 가능하다. 두 인물의 하체는 화면 밖이므로 발의 접지는 확인할 수 없지만 부유를 뜻하는 모습은 없다. 배경의 추가 부츠와 화병 조각은 바닥에 닿아 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리를 향한 험상궂은 시선과 실제 팔 움켜쥠은 더 정확하지만, 구도가 요구된 밀착 클로즈업보다 넓고 읽히는 가슴 로고가 명시적 금지사항을 위반한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "상체와 움켜쥔 손의 밀착 구도는 더 가깝지만, 배경에 제삼자의 다리와 부츠가 남아 있고 현우가 찰리 대신 카메라를 응시한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸을 앞으로 숙이고 화면 오른쪽 찰리의 상체와 얼굴 쪽을 노려본다. 양손은 찰리의 전완 장갑을 위와 아래에서 붙잡는다. 찰리의 얼굴은 현우 쪽으로 내려가 있으나 눈은 상단 잘림으로 확인하기 어렵다. 무기나 이동 중인 물체는 없다.",
        "built_space": "왼쪽에 창 하나와 문 하나, 뒤쪽 오른편에 작업대 하나와 의자 일부, 작업대 아래 상자들이 보인다. 골진 금속 벽과 천장, 낡은 목재 바닥은 장소 참고와 대체로 이어진다. 두 인물 사이로 바닥과 긴 실내가 상당히 드러나므로 배경을 좁게 남기는 밀착 클로즈업보다 넓다. 바닥에는 깨진 화병 조각 한 무더기가 있다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 회색 겉셔츠, 얼굴 찰과상이 참고 및 지시와 대체로 맞는다. 찰리는 육중한 베이지색 기계 팔, 마모된 장갑, 낡은 코트와 모자 일부, 흰 기계 얼굴을 유지한다. 눈과 다리 부상은 구도상 확인할 수 없다. 추가 인물이나 온전한 대체 화병은 보이지 않는다. 다만 찰리 가슴의 'Ubik' 문자와 상징은 선명하게 읽힌다.",
        "hard_violations": [
         "찰리의 가슴에 판독 가능한 'Ubik' 문자와 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "현우의 한 손은 찰리 전완 장갑의 윗면을 감싸고 다른 손은 아래쪽을 받쳐 실제 접촉하며 움켜쥔다. 찰리의 팔은 팔꿈치와 어깨로 몸통에 연결되어 있고 현우의 손도 이를 지지한다. 두 몸통은 화면 아래로 이어져 있으며 공중에 떠 있다는 징후는 없다. 화병 조각들은 바닥에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "현우는 눈썹을 찌푸렸지만 시선은 찰리가 아니라 카메라 쪽을 향한다. 찰리는 머리를 왼쪽 아래, 현우와 맞잡힌 팔 쪽으로 기울인다. 현우의 손은 가슴 앞을 가로지르는 찰리의 전완을 움켜쥐고 있다. 팔 접촉 방향은 맞지만 현우의 응시 대상이 대치 상대와 어긋난다.",
        "built_space": "왼쪽 금속 벽에 세로 배관과 전기함, 문 하나가 보이고, 뒤로 골진 천장과 목재 바닥이 이어진다. 참고 장소의 왼쪽 벽 배치와 재질은 잘 유지된다. 두 상체와 전경의 팔이 화면 대부분을 차지해 A보다 밀착 구도에 가깝다. 하단에는 화병 조각 한 무더기와, 두 주인공 뒤 중앙에 별도의 바지 다리와 부츠 하나가 보인다. 중복 설비나 반사는 보이지 않는다.",
        "entities": "현우의 젊은 동아시아계 남성 외형, 검은 머리, 회색 셔츠와 얼굴 상처는 대체로 맞는다. 찰리의 베이지 장갑, 흰 기계 얼굴, 파란 눈과 낡은 코트가 보인다. 머리 윗부분은 잘렸지만 드러난 두상 주변에 요구된 모자의 챙은 보이지 않는다. 읽히는 가슴 문자는 없다. 뒤쪽 중앙의 갈색 바지와 부츠는 두 주인공의 위치 및 복장과 맞지 않는 제삼자의 신체 일부다. 다리 부상은 구도 밖이다.",
        "hard_violations": [
         "두 주인공 뒤 바닥에 제삼자의 바지 다리와 부츠가 남아 있어, 현우와 찰리 외 인물 및 이전 장면 인물의 신체를 배제하라는 지시를 위반한다."
        ],
        "physics": "찰리의 가로놓인 전완은 오른쪽 팔꿈치와 상완으로 이어지고, 현우의 손가락은 장갑과 소매 부위를 실제로 감싼다. 팔이 현우의 몸 앞에 밀착된 자세는 물리적으로 가능하다. 두 인물의 하체는 화면 밖이므로 발의 접지는 확인할 수 없지만 부유를 뜻하는 모습은 없다. 배경의 추가 부츠와 화병 조각은 바닥에 닿아 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에서 제외를 명시한 이전 샷 인물의 신체 일부(배경에 서 있는 부츠 신은 두 다리)가 화면에 나타남.",
     "[gpt-high] 두 주인공 뒤 바닥에 제삼자의 바지 다리와 부츠가 남아 있어, 현우와 찰리 외 인물 및 이전 장면 인물의 신체를 배제하라는 지시를 위반한다."
    ],
    "B": [
     "[gpt-high] 찰리의 가슴에 판독 가능한 'Ubik' 문자와 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "현우가 찰리의 팔을 거칠게 움켜쥔 텐션을 클로즈업 샷으로 사실적으로 묘사했으며, 찰리의 모자, Ubik 로고, 바닥의 깨진 도자기 등 세부 디테일을 완벽하게 구현했습니다.  ★위반: [gpt-high] 찰리의 가슴에 판독 가능한 'Ubik' 문자와 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "두 인물 사이의 배경 바닥에 프롬프트에서 제외해야 할 이전 샷 인물의 부츠 신은 다리가 남아있는 치명적인 오류가 발생했으며, 찰리의 모자와 로고도 누락되었습니다.  ★위반: [gemini-pro] 프롬프트에서 제외를 명시한 이전 샷 인물의 신체 일부(배경에 서 있는 부츠 신은 두 다리)가 화면에 나타남. / [gpt-high] 두 주인공 뒤 바닥에 제삼자의 바지 다리와 부츠가 남아 있어, 현우와 찰리 외 인물 및 이전 장면 인물의 신체를 배제하라는 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S14sh5_sel.png",
    "asset_id": "d1099389-4eb2-4b7a-9608-4caf1dc563f0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-d10e-7757-959e-c8b3d73c9bf1",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S14sh5"
  }
 },
 "S14sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:32:25.692739+00:00",
  "fingerprint": "2cef9edb91a45367dadd3de1704b34d5b408b4fe44afd7811ba3ef8ad3d77f0d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S14sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S14sh9_sel.png",
  "source_sha256": "205baa2cc9d51f7810bd4d45674a207c25253d7fed3b84226f8bed860e5ae023",
  "file": "S14sh9_cine.png",
  "staged_sha256": "385dde77d9db754c112328b6d8d825a82511d85d83fb06bd2e2950ddad2b6b3f",
  "latency_ms": 10429
 },
 "S15sh4::signage": {
  "fp": "507064929f8094a7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::c685222b67bf4235": {
  "subjects": [],
  "subject_text": "마츠다의 고물상 내부\n온갖 낡은 가전제품과 고물이 빽빽하게 쌓인 상점. 전기제품 수리용 책상과 텔레비전, 컴퓨터가 비좁은 실내에 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L168",
  "scope_role": "location_interior",
  "scope_sha": "49a46d2906d2c018"
 },
 "S15sh4::bgfirst_bg": {
  "input_fingerprint": "0c214deaf052740c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4__bgfirst_bg.png",
  "asset_id": "08f633d6-c3f3-4ebc-a4c1-0fb688cbc402",
  "input_asset_ids": [
   "8a256e3b-52f6-4af6-aa77-ef9e992d170a",
   "cb825975-b374-47cb-9da4-9a5bcc3765b2"
  ]
 },
 "S15sh4": {
  "input_fingerprint": "f8605918a29e7fbe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appliances and scrap are piled throughout the shop. Charlie remains dirty and worn, with blue-lit eyes, an old coat and hat, and a visibly worn Ubik logo on his chest. 마츠다: He has moved away from his repair work and is standing in the shop's inspection area with an arm extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appliances and scrap are piled throughout the shop. Charlie remains dirty and worn, with blue-lit eyes, an old coat and hat, and a visibly worn Ubik logo on his chest. 마츠다: He has moved away from his repair work and is standing in the shop's inspection area with an arm extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appliances and scrap are piled throughout the shop. Charlie remains dirty and worn, with blue-lit eyes, an old coat and hat, and a visibly worn Ubik logo on his chest. 마츠다: He has moved away from his repair work and is standing in the shop's inspection area with an arm extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4__bgfirst_bg.png",
     "asset_id": "08f633d6-c3f3-4ebc-a4c1-0fb688cbc402",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S15sh4.png",
     "asset_id": "8a256e3b-52f6-4af6-aa77-ef9e992d170a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:971613>",
     "asset_id": "945cc191-1d4f-461d-b482-4157f3024b75",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L168B03.png",
     "asset_id": "cb825975-b374-47cb-9da4-9a5bcc3765b2",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:971613>",
     "asset_id": "945cc191-1d4f-461d-b482-4157f3024b75",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "마츠다의 손가락이 찰리의 가슴에 있는 의미 불명의 원형 문양을 가리킴.",
    "built_space": "배경에 낡은 가전제품이 쌓인 고물상 내부가 묘사됨.",
    "entities": "마츠다의 얼굴과 상체가 보임. 찰리는 회색 장갑을 착용하고 있으며 지정된 코트와 모자가 없고 '유빅' 글자도 없음.",
    "hard_violations": [
     "[gemini-pro] 찰리의 외형(장갑 색상, 코트, 모자 누락) 불일치",
     "[gemini-pro] 필수 텍스트 '유빅' 누락"
    ],
    "physics": "두 인물 모두 지면에 서 있으며 팔의 자세와 지탱이 자연스러움."
   },
   {
    "label": "B",
    "direction": "마츠다의 손가락이 찰리의 가슴 장갑에 새겨진 '유빅' 글자를 정확히 가리킴.",
    "built_space": "참조 이미지와 일치하는 가전제품이 쌓인 고물상 내부가 표현됨.",
    "entities": "지시문대로 마츠다의 팔만 프레임에 나타남. 찰리는 샌드 베이지 장갑, 낡은 코트와 모자를 착용하고 '유빅' 로고를 정확히 표시함.",
    "hard_violations": [
     "[gpt-high] 가슴의 ‘유빅’ 문자가 명확히 판독되어, 마지막에 명시된 판독 가능한 글자 금지를 위반한다."
    ],
    "physics": "화면 밖의 몸통에서 뻗어 나온 팔과 찰리의 스탠딩 자세가 물리적으로 자연스럽게 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '유빅' 텍스트가 누락되었고 찰리의 장갑 색상 및 복장(코트, 모자)이 참조와 일치하지 않습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 '유빅' 텍스트와 찰리의 복장(장갑, 코트, 모자)을 클로즈업 구도에 맞춰 충실히 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "마츠다의 손가락이 찰리의 가슴에 있는 의미 불명의 원형 문양을 가리킴.",
        "built_space": "배경에 낡은 가전제품이 쌓인 고물상 내부가 묘사됨.",
        "entities": "마츠다의 얼굴과 상체가 보임. 찰리는 회색 장갑을 착용하고 있으며 지정된 코트와 모자가 없고 '유빅' 글자도 없음.",
        "hard_violations": [
         "찰리의 외형(장갑 색상, 코트, 모자 누락) 불일치",
         "필수 텍스트 '유빅' 누락"
        ],
        "physics": "두 인물 모두 지면에 서 있으며 팔의 자세와 지탱이 자연스러움."
       },
       {
        "label": "B",
        "direction": "마츠다의 손가락이 찰리의 가슴 장갑에 새겨진 '유빅' 글자를 정확히 가리킴.",
        "built_space": "참조 이미지와 일치하는 가전제품이 쌓인 고물상 내부가 표현됨.",
        "entities": "지시문대로 마츠다의 팔만 프레임에 나타남. 찰리는 샌드 베이지 장갑, 낡은 코트와 모자를 착용하고 '유빅' 로고를 정확히 표시함.",
        "hard_violations": [],
        "physics": "화면 밖의 몸통에서 뻗어 나온 팔과 찰리의 스탠딩 자세가 물리적으로 자연스럽게 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '유빅' 텍스트가 누락되었고 찰리의 장갑 색상 및 복장(코트, 모자)이 참조와 일치하지 않습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 '유빅' 텍스트와 찰리의 복장(장갑, 코트, 모자)을 클로즈업 구도에 맞춰 충실히 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "마츠다의 손가락이 찰리의 가슴에 있는 의미 불명의 원형 문양을 가리킴.",
        "built_space": "배경에 낡은 가전제품이 쌓인 고물상 내부가 묘사됨.",
        "entities": "마츠다의 얼굴과 상체가 보임. 찰리는 회색 장갑을 착용하고 있으며 지정된 코트와 모자가 없고 '유빅' 글자도 없음.",
        "hard_violations": [
         "찰리의 외형(장갑 색상, 코트, 모자 누락) 불일치",
         "필수 텍스트 '유빅' 누락"
        ],
        "physics": "두 인물 모두 지면에 서 있으며 팔의 자세와 지탱이 자연스러움."
       },
       {
        "label": "B",
        "direction": "마츠다의 손가락이 찰리의 가슴 장갑에 새겨진 '유빅' 글자를 정확히 가리킴.",
        "built_space": "참조 이미지와 일치하는 가전제품이 쌓인 고물상 내부가 표현됨.",
        "entities": "지시문대로 마츠다의 팔만 프레임에 나타남. 찰리는 샌드 베이지 장갑, 낡은 코트와 모자를 착용하고 '유빅' 로고를 정확히 표시함.",
        "hard_violations": [],
        "physics": "화면 밖의 몸통에서 뻗어 나온 팔과 찰리의 스탠딩 자세가 물리적으로 자연스럽게 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "팔과 가슴 중심의 근접 구도 및 베이지 장갑·낡은 코트는 더 충실하지만, ‘유빅’ 글자가 읽혀 최종 지시의 판독 가능한 문자 금지를 위반한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "손끝이 마모된 가슴 표식을 정확히 짚고 글자도 판독되지 않지만, 마츠다까지 넓어진 구도와 회백색 장갑·걸친 천은 지정 구도와 찰리 참조에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 아래에서 뻗은 마츠다의 검지가 찰리 가슴의 ‘유빅’ 글자 아래쪽에 직접 닿는다. 지시 대상은 정확하다. 찰리의 얼굴 아래쪽만 보이므로 눈의 시선은 확인할 수 없다.",
        "built_space": "왼쪽에 금속 선반, 냉장고 한 대, 그 오른쪽에 소형 브라운관 한 대와 네 층가량의 상자형 가전 더미가 보인다. 뒤쪽의 부품 선반도 장소 참조와 대응한다. 찰리는 오른쪽 전경, 마츠다는 왼쪽 프레임 밖에 있고 팔만 검사 공간으로 들어온다. 중복 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리의 각진 샌드 베이지 장갑, 흰 마스크형 턱, 낡은 올리브색 코트와 모자 일부가 참조와 맞는다. 눈과 하체는 잘려 평가할 수 없다. 마츠다의 주름지고 혈관이 드러난 손과 걷어 올린 작업복 소매는 노년 남성 설정에 부합하나 얼굴과 국적은 확인할 수 없다. 가슴 글자는 표면의 긁힘을 따르지만 ‘유빅’으로 읽힌다. 별도의 인물은 없다.",
        "hard_violations": [
         "가슴의 ‘유빅’ 문자가 명확히 판독되어, 마지막에 명시된 판독 가능한 글자 금지를 위반한다."
        ],
        "physics": "손가락은 손과 손목, 전완에 자연스럽게 연결되며 소매 밖으로 이어진 팔이 이를 지지한다. 손끝과 가슴 장갑의 접촉도 가능하다. 장갑은 몸통에 부착되고 코트는 어깨에서 내려온다. 발은 화면 밖이지만 몸이 떠 있다는 징후나 지지 없는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "왼쪽 마츠다의 검지가 가슴 중앙의 닳은 원형 표식 왼쪽 부분에 닿으며, 마츠다도 고개를 숙여 그 부위를 보고 있다. 찰리의 눈은 화면 위에서 잘려 시선 방향을 확정할 수 없다.",
        "built_space": "뒤쪽에 브라운관 한 대와 그 아래 상자형 가전 더미, 중앙 통로, 후면 목재 카운터 한 곳, 공구판 한 곳과 상부 기기 선반이 보인다. 참조 장소의 배열과 재료가 유지된다. 마츠다는 왼쪽, 찰리는 오른쪽에서 통로를 사이에 두기보다 가까이 마주하여 검사한다. 설비의 부당한 중복이나 불가능한 반사는 없다.",
        "entities": "마츠다는 주름과 희끗한 머리의 동아시아계 노년 남성으로 보이지만, 얼굴 일부만으로 일본인 설정이나 참조 인물과의 정확한 일치를 확인하기는 어렵다. 찰리의 흰 기계식 마스크와 육중한 상체는 맞지만 장갑은 샌드 베이지보다 회백색이며, 코트는 참조의 착용 형태보다 어깨에 걸친 천처럼 보인다. 가슴의 마모된 표식은 물리적 도장으로 보이고 읽을 수 있는 글자는 없으나, 남은 흔적만으로 ‘유빅’임을 식별하기 어렵다. 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "마츠다의 팔은 어깨에서 팔꿈치와 손목으로 자연스럽게 이어지고 검지만 펴서 가슴에 접촉한다. 찰리의 장갑판은 몸통과 어깨 구조에 연결되어 있으며 천은 어깨가 지지한다. 두 인물의 발은 잘렸지만 공중 부양이나 지지 없는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "팔과 가슴 중심의 근접 구도 및 베이지 장갑·낡은 코트는 더 충실하지만, ‘유빅’ 글자가 읽혀 최종 지시의 판독 가능한 문자 금지를 위반한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "손끝이 마모된 가슴 표식을 정확히 짚고 글자도 판독되지 않지만, 마츠다까지 넓어진 구도와 회백색 장갑·걸친 천은 지정 구도와 찰리 참조에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 아래에서 뻗은 마츠다의 검지가 찰리 가슴의 ‘유빅’ 글자 아래쪽에 직접 닿는다. 지시 대상은 정확하다. 찰리의 얼굴 아래쪽만 보이므로 눈의 시선은 확인할 수 없다.",
        "built_space": "왼쪽에 금속 선반, 냉장고 한 대, 그 오른쪽에 소형 브라운관 한 대와 네 층가량의 상자형 가전 더미가 보인다. 뒤쪽의 부품 선반도 장소 참조와 대응한다. 찰리는 오른쪽 전경, 마츠다는 왼쪽 프레임 밖에 있고 팔만 검사 공간으로 들어온다. 중복 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리의 각진 샌드 베이지 장갑, 흰 마스크형 턱, 낡은 올리브색 코트와 모자 일부가 참조와 맞는다. 눈과 하체는 잘려 평가할 수 없다. 마츠다의 주름지고 혈관이 드러난 손과 걷어 올린 작업복 소매는 노년 남성 설정에 부합하나 얼굴과 국적은 확인할 수 없다. 가슴 글자는 표면의 긁힘을 따르지만 ‘유빅’으로 읽힌다. 별도의 인물은 없다.",
        "hard_violations": [
         "가슴의 ‘유빅’ 문자가 명확히 판독되어, 마지막에 명시된 판독 가능한 글자 금지를 위반한다."
        ],
        "physics": "손가락은 손과 손목, 전완에 자연스럽게 연결되며 소매 밖으로 이어진 팔이 이를 지지한다. 손끝과 가슴 장갑의 접촉도 가능하다. 장갑은 몸통에 부착되고 코트는 어깨에서 내려온다. 발은 화면 밖이지만 몸이 떠 있다는 징후나 지지 없는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "왼쪽 마츠다의 검지가 가슴 중앙의 닳은 원형 표식 왼쪽 부분에 닿으며, 마츠다도 고개를 숙여 그 부위를 보고 있다. 찰리의 눈은 화면 위에서 잘려 시선 방향을 확정할 수 없다.",
        "built_space": "뒤쪽에 브라운관 한 대와 그 아래 상자형 가전 더미, 중앙 통로, 후면 목재 카운터 한 곳, 공구판 한 곳과 상부 기기 선반이 보인다. 참조 장소의 배열과 재료가 유지된다. 마츠다는 왼쪽, 찰리는 오른쪽에서 통로를 사이에 두기보다 가까이 마주하여 검사한다. 설비의 부당한 중복이나 불가능한 반사는 없다.",
        "entities": "마츠다는 주름과 희끗한 머리의 동아시아계 노년 남성으로 보이지만, 얼굴 일부만으로 일본인 설정이나 참조 인물과의 정확한 일치를 확인하기는 어렵다. 찰리의 흰 기계식 마스크와 육중한 상체는 맞지만 장갑은 샌드 베이지보다 회백색이며, 코트는 참조의 착용 형태보다 어깨에 걸친 천처럼 보인다. 가슴의 마모된 표식은 물리적 도장으로 보이고 읽을 수 있는 글자는 없으나, 남은 흔적만으로 ‘유빅’임을 식별하기 어렵다. 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "마츠다의 팔은 어깨에서 팔꿈치와 손목으로 자연스럽게 이어지고 검지만 펴서 가슴에 접촉한다. 찰리의 장갑판은 몸통과 어깨 구조에 연결되어 있으며 천은 어깨가 지지한다. 두 인물의 발은 잘렸지만 공중 부양이나 지지 없는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.321
   },
   "violations": {
    "A": [
     "[gemini-pro] 찰리의 외형(장갑 색상, 코트, 모자 누락) 불일치",
     "[gemini-pro] 필수 텍스트 '유빅' 누락"
    ],
    "B": [
     "[gpt-high] 가슴의 ‘유빅’ 문자가 명확히 판독되어, 마지막에 명시된 판독 가능한 글자 금지를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "지정된 '유빅' 텍스트가 누락되었고 찰리의 장갑 색상 및 복장(코트, 모자)이 참조와 일치하지 않습니다.  ★위반: [gemini-pro] 찰리의 외형(장갑 색상, 코트, 모자 누락) 불일치 / [gemini-pro] 필수 텍스트 '유빅' 누락"
   },
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "요구된 '유빅' 텍스트와 찰리의 복장(장갑, 코트, 모자)을 클로즈업 구도에 맞춰 충실히 구현했습니다.  ★위반: [gpt-high] 가슴의 ‘유빅’ 문자가 명확히 판독되어, 마지막에 명시된 판독 가능한 글자 금지를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L168B03.png",
    "asset_id": "cb825975-b374-47cb-9da4-9a5bcc3765b2",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:971613>",
    "asset_id": "945cc191-1d4f-461d-b482-4157f3024b75",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-d2bb-781b-ac6a-69068ae10b52",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4__bgfirst_bg.png",
   "bg_asset_id": "08f633d6-c3f3-4ebc-a4c1-0fb688cbc402",
   "bg_record_key": "S15sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S15sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:33:35.656804+00:00",
  "fingerprint": "cd196a15f76a408317eddf871fb60ca642bf6cceee655236932e013f4f503ed2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S15sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S15sh4_sel.png",
  "source_sha256": "b48a1e19f2dfc4aed257ac5de20fabfc8cfd89924530d7312c02f967e63f6ea1",
  "file": "S15sh4_cine.png",
  "staged_sha256": "e92985845ad8ec51a558eee66424b954d53f98eb97f3f19b8fd3268f52cbb2fc",
  "latency_ms": 12149
 },
 "S15sh10::signage": {
  "fp": "44823a96f0df2315",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S15sh10": {
  "input_fingerprint": "3107928d7b25be0b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 어두운 상점 안, 낡은 전화기 수화기를 귀에 바짝 댄 채 입꼬리를 올린 마츠다의 얼굴.\n\nLOCATION (lock): At the telephone area inside the cluttered junk shop, in its dim daytime interior. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old telephone receiver (Held tightly against Matsuda's ear during the call) — Seen along his near cheek, with the mouthpiece beside rather than across his smile; used as Connect the private expression to the act of reporting Charlie while remaining small relative to the face; Piled appliances and scrap (Still accumulated inside the shop) — Indistinct portions remain behind Matsuda; used as Anchor the private call in the same shop without drawing attention away from his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the described dim shop ambience with controlled facial detail and no newly introduced practical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains piled with appliances and scrap, and the computer browser has the retrieved news about the unique missing Ubik robot and its 500-million-won reward. Charlie is outside the shop, still dirty and worn, with the old coat and hat and worn chest logo. 마츠다: He remains inside the shop, using the telephone.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 마츠다 right now, so 마츠다's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 마츠다: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 어두운 상점 안, 낡은 전화기 수화기를 귀에 바짝 댄 채 입꼬리를 올린 마츠다의 얼굴.\n\nLOCATION (lock): At the telephone area inside the cluttered junk shop, in its dim daytime interior. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old telephone receiver (Held tightly against Matsuda's ear during the call) — Seen along his near cheek, with the mouthpiece beside rather than across his smile; used as Connect the private expression to the act of reporting Charlie while remaining small relative to the face; Piled appliances and scrap (Still accumulated inside the shop) — Indistinct portions remain behind Matsuda; used as Anchor the private call in the same shop without drawing attention away from his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the described dim shop ambience with controlled facial detail and no newly introduced practical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains piled with appliances and scrap, and the computer browser has the retrieved news about the unique missing Ubik robot and its 500-million-won reward. Charlie is outside the shop, still dirty and worn, with the old coat and hat and worn chest logo. 마츠다: He remains inside the shop, using the telephone.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 마츠다 right now, so 마츠다's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 마츠다: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 어두운 상점 안, 낡은 전화기 수화기를 귀에 바짝 댄 채 입꼬리를 올린 마츠다의 얼굴.\n\nLOCATION (lock): At the telephone area inside the cluttered junk shop, in its dim daytime interior. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old telephone receiver (Held tightly against Matsuda's ear during the call) — Seen along his near cheek, with the mouthpiece beside rather than across his smile; used as Connect the private expression to the act of reporting Charlie while remaining small relative to the face; Piled appliances and scrap (Still accumulated inside the shop) — Indistinct portions remain behind Matsuda; used as Anchor the private call in the same shop without drawing attention away from his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the described dim shop ambience with controlled facial detail and no newly introduced practical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains piled with appliances and scrap, and the computer browser has the retrieved news about the unique missing Ubik robot and its 500-million-won reward. Charlie is outside the shop, still dirty and worn, with the old coat and hat and worn chest logo. 마츠다: He remains inside the shop, using the telephone.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 마츠다 right now, so 마츠다's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 마츠다: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "마츠다의 시선은 화면 왼쪽 밖을 향하며, 미소를 지은 채 오른쪽 귀에 전화기를 대고 있습니다.",
    "built_space": "배경에 구형 모니터들이 쌓여 있고 밝은 빛이 들어오는 창문이 위치하여, 장소 기준 이미지의 선반과 기기 배치 및 어두운 조명 분위기와 일치하지 않습니다.",
    "entities": "마츠다의 얼굴, 의상, 파란색 모자가 캐릭터 레퍼런스와 잘 일치합니다. 검은색 낡은 수화기를 들고 있습니다.",
    "hard_violations": [],
    "physics": "오른손이 수화기를 자연스럽게 쥐고 있으며, 물리적으로 불가능한 자세나 중력 오류는 없습니다."
   },
   {
    "label": "B",
    "direction": "마츠다의 시선은 화면 왼쪽 밖을 향하며, 입꼬리를 올린 채 오른쪽 귀에 낡은 수화기를 바짝 대고 있습니다.",
    "built_space": "어두운 실내 환경이며, 배경 왼쪽의 금속 선반과 그 위에 쌓인 흰색 세탁기 및 전자레인지 등의 배치가 이전 숏 기준 이미지와 정확하게 일치합니다.",
    "entities": "마츠다의 얼굴과 의상은 일치하나 레퍼런스에 있는 파란색 모자를 착용하지 않았습니다. 베이지색의 낡은 수화기가 프롬프트의 요구에 맞게 뺨 옆에 묘사되었습니다.",
    "hard_violations": [],
    "physics": "손가락이 수화기를 안정적으로 파지하고 있으며, 꼬인 전화선이 아래로 자연스럽게 늘어져 있어 물리적인 어색함이 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "캐릭터의 복장(모자 포함)은 잘 일치하나, 기준 이미지에 없는 창문과 다른 종류의 폐가전들을 배경에 생성하여 더 높은 우선순위인 장소 고정(Location lock) 조건을 위반했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "모자가 누락되는 의상 묘사 오류가 있으나, 이전 숏 기준 이미지의 배경(금속 선반과 층층이 쌓인 가전제품) 및 어두운 분위기를 완벽하게 재현하여 장소 우선순위를 충족시킨 가장 우수한 결과물입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "마츠다의 시선은 화면 왼쪽 밖을 향하며, 미소를 지은 채 오른쪽 귀에 전화기를 대고 있습니다.",
        "built_space": "배경에 구형 모니터들이 쌓여 있고 밝은 빛이 들어오는 창문이 위치하여, 장소 기준 이미지의 선반과 기기 배치 및 어두운 조명 분위기와 일치하지 않습니다.",
        "entities": "마츠다의 얼굴, 의상, 파란색 모자가 캐릭터 레퍼런스와 잘 일치합니다. 검은색 낡은 수화기를 들고 있습니다.",
        "hard_violations": [],
        "physics": "오른손이 수화기를 자연스럽게 쥐고 있으며, 물리적으로 불가능한 자세나 중력 오류는 없습니다."
       },
       {
        "label": "B",
        "direction": "마츠다의 시선은 화면 왼쪽 밖을 향하며, 입꼬리를 올린 채 오른쪽 귀에 낡은 수화기를 바짝 대고 있습니다.",
        "built_space": "어두운 실내 환경이며, 배경 왼쪽의 금속 선반과 그 위에 쌓인 흰색 세탁기 및 전자레인지 등의 배치가 이전 숏 기준 이미지와 정확하게 일치합니다.",
        "entities": "마츠다의 얼굴과 의상은 일치하나 레퍼런스에 있는 파란색 모자를 착용하지 않았습니다. 베이지색의 낡은 수화기가 프롬프트의 요구에 맞게 뺨 옆에 묘사되었습니다.",
        "hard_violations": [],
        "physics": "손가락이 수화기를 안정적으로 파지하고 있으며, 꼬인 전화선이 아래로 자연스럽게 늘어져 있어 물리적인 어색함이 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "캐릭터의 복장(모자 포함)은 잘 일치하나, 기준 이미지에 없는 창문과 다른 종류의 폐가전들을 배경에 생성하여 더 높은 우선순위인 장소 고정(Location lock) 조건을 위반했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "모자가 누락되는 의상 묘사 오류가 있으나, 이전 숏 기준 이미지의 배경(금속 선반과 층층이 쌓인 가전제품) 및 어두운 분위기를 완벽하게 재현하여 장소 우선순위를 충족시킨 가장 우수한 결과물입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "마츠다의 시선은 화면 왼쪽 밖을 향하며, 미소를 지은 채 오른쪽 귀에 전화기를 대고 있습니다.",
        "built_space": "배경에 구형 모니터들이 쌓여 있고 밝은 빛이 들어오는 창문이 위치하여, 장소 기준 이미지의 선반과 기기 배치 및 어두운 조명 분위기와 일치하지 않습니다.",
        "entities": "마츠다의 얼굴, 의상, 파란색 모자가 캐릭터 레퍼런스와 잘 일치합니다. 검은색 낡은 수화기를 들고 있습니다.",
        "hard_violations": [],
        "physics": "오른손이 수화기를 자연스럽게 쥐고 있으며, 물리적으로 불가능한 자세나 중력 오류는 없습니다."
       },
       {
        "label": "B",
        "direction": "마츠다의 시선은 화면 왼쪽 밖을 향하며, 입꼬리를 올린 채 오른쪽 귀에 낡은 수화기를 바짝 대고 있습니다.",
        "built_space": "어두운 실내 환경이며, 배경 왼쪽의 금속 선반과 그 위에 쌓인 흰색 세탁기 및 전자레인지 등의 배치가 이전 숏 기준 이미지와 정확하게 일치합니다.",
        "entities": "마츠다의 얼굴과 의상은 일치하나 레퍼런스에 있는 파란색 모자를 착용하지 않았습니다. 베이지색의 낡은 수화기가 프롬프트의 요구에 맞게 뺨 옆에 묘사되었습니다.",
        "hard_violations": [],
        "physics": "손가락이 수화기를 안정적으로 파지하고 있으며, 꼬인 전화선이 아래로 자연스럽게 늘어져 있어 물리적인 어색함이 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "모자가 빠진 외형 차이는 있지만, 얼굴 중심의 밀착 클로즈업과 귀에 붙인 수화기, 올라간 입꼬리, 이전 장면의 가전 적재 배치가 핵심 지시를 더 충실히 구현한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "모자와 작업복을 포함한 인물 연속성은 좋지만, 가슴까지 넓어진 구도와 두드러진 창·모니터 배경이 얼굴 중심 클로즈업 및 이전 장소의 시각적 연속성에서 밀린다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "눈은 카메라가 아닌 화면 왼쪽 바깥을 향하며, 그곳의 구체적인 대상은 보이지 않는다. 수화기의 위쪽은 화면 왼쪽 귀에 밀착하고 아래 송화부는 입 왼쪽 아래에 있어 미소를 가리지 않는다. 전화 사용 방향은 자연스럽다.",
        "built_space": "왼쪽에 금속 선반과 세로형 냉장고 한 대, 중앙 뒤에 작은 브라운관 모니터 한 대와 그 아래 직사각형 가전 네 대가 층층이 보인다. 왼쪽 전경에도 금속 가전이 쌓여 있다. 이전 장면의 냉장고·모니터·가전 적재 관계가 잘 이어진다. 인물은 그 앞에서 얼굴과 손 위주로 잡혔고, 새 조명기구나 판정할 반사는 없다. 다만 배경 가전의 윤곽은 요구된 흐릿한 부분 묘사보다 다소 선명하다.",
        "entities": "보이는 사람은 나이 든 동아시아계 남성 한 명이며, 눈가 주름과 거친 피부는 60대 마츠다 설정에 부합한다. 일본 국적 자체는 외관으로 확인할 수 없다. 얼굴은 참조와 대체로 유사하지만 머리카락이 드러나 참조의 남색 모자가 유지되지 않았다. 보이는 갈색 작업 셔츠와 어두운 조끼는 부합한다. 낡은 유선 수화기 한 개와 가전·고철 배경이 있으며, 다른 인물이나 읽을 수 있는 글자는 없다. 컴퓨터 기사와 바깥의 찰리는 이 클로즈업에 보이지 않아도 된다.",
        "hard_violations": [],
        "physics": "나이 든 맨손이 수화기 손잡이를 감싸 쥐고 귀 쪽으로 누르고 있다. 손목과 팔뚝은 화면 아래까지 자연스럽게 이어지며, 전화선은 수화기 하단에서 아래로 늘어진다. 가전은 아래 기기나 선반에 받쳐져 있다. 하체는 프레임 밖이므로 발의 지지는 확인할 수 없지만, 공중에 떠 있는 신체나 무지지 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "시선은 화면 왼쪽 바깥으로 향하고 렌즈를 직접 보지 않는다. 수화기 수화부는 화면 왼쪽 귀에 붙어 있고 송화부는 입 왼쪽 아래를 향한다. 입꼬리가 올라간 표정을 수화기가 가리지 않으며 통화 방향도 맞다.",
        "built_space": "뒤쪽 상단에 큰 분할 창 한 개, 왼쪽부터 중앙에 브라운관 화면 약 여섯 개, 왼쪽 중앙에 여러 층의 오디오 기기, 오른쪽에 케이블이 얹힌 가전 더미가 보인다. 인물은 이 적재물 앞에 있으며 가슴과 양어깨까지 포함된다. 고물상 재질과 혼잡함은 맞지만, 이전 장면의 냉장고와 세로 가전 더미 대신 창과 모니터 군집이 두드러져 동일 지점의 연속성은 덜 확실하다. 불가능한 반사나 명백히 중복된 고정 설비는 보이지 않는다.",
        "entities": "보이는 인물은 마츠다에 해당하는 나이 든 동아시아계 남성 한 명이다. 눈가 주름, 남색 작업 모자, 갈색 셔츠와 어두운 공구 조끼가 인물 참조와 잘 맞는다. 일본 국적은 이미지 자체로 판별할 수 없다. 검은 낡은 유선 수화기 한 개와 쌓인 전자제품·케이블이 보이며, 다른 사람이나 읽을 수 있는 글자는 없다. 기사 화면과 찰리는 프레임 밖으로 볼 수 있다.",
        "hard_violations": [],
        "physics": "맨손의 손가락이 수화기 중앙을 감싸고 팔뚝이 아래로 이어져 수화기의 무게를 지탱한다. 수화부와 귀의 접촉, 아래로 처진 전화선이 물리적으로 자연스럽다. 배경 모니터와 오디오 기기는 아래 적재물에 받쳐져 있고, 케이블은 가전 위에 놓여 있다. 공중에 떠 있거나 지지 없는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "모자가 빠진 외형 차이는 있지만, 얼굴 중심의 밀착 클로즈업과 귀에 붙인 수화기, 올라간 입꼬리, 이전 장면의 가전 적재 배치가 핵심 지시를 더 충실히 구현한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "모자와 작업복을 포함한 인물 연속성은 좋지만, 가슴까지 넓어진 구도와 두드러진 창·모니터 배경이 얼굴 중심 클로즈업 및 이전 장소의 시각적 연속성에서 밀린다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "눈은 카메라가 아닌 화면 왼쪽 바깥을 향하며, 그곳의 구체적인 대상은 보이지 않는다. 수화기의 위쪽은 화면 왼쪽 귀에 밀착하고 아래 송화부는 입 왼쪽 아래에 있어 미소를 가리지 않는다. 전화 사용 방향은 자연스럽다.",
        "built_space": "왼쪽에 금속 선반과 세로형 냉장고 한 대, 중앙 뒤에 작은 브라운관 모니터 한 대와 그 아래 직사각형 가전 네 대가 층층이 보인다. 왼쪽 전경에도 금속 가전이 쌓여 있다. 이전 장면의 냉장고·모니터·가전 적재 관계가 잘 이어진다. 인물은 그 앞에서 얼굴과 손 위주로 잡혔고, 새 조명기구나 판정할 반사는 없다. 다만 배경 가전의 윤곽은 요구된 흐릿한 부분 묘사보다 다소 선명하다.",
        "entities": "보이는 사람은 나이 든 동아시아계 남성 한 명이며, 눈가 주름과 거친 피부는 60대 마츠다 설정에 부합한다. 일본 국적 자체는 외관으로 확인할 수 없다. 얼굴은 참조와 대체로 유사하지만 머리카락이 드러나 참조의 남색 모자가 유지되지 않았다. 보이는 갈색 작업 셔츠와 어두운 조끼는 부합한다. 낡은 유선 수화기 한 개와 가전·고철 배경이 있으며, 다른 인물이나 읽을 수 있는 글자는 없다. 컴퓨터 기사와 바깥의 찰리는 이 클로즈업에 보이지 않아도 된다.",
        "hard_violations": [],
        "physics": "나이 든 맨손이 수화기 손잡이를 감싸 쥐고 귀 쪽으로 누르고 있다. 손목과 팔뚝은 화면 아래까지 자연스럽게 이어지며, 전화선은 수화기 하단에서 아래로 늘어진다. 가전은 아래 기기나 선반에 받쳐져 있다. 하체는 프레임 밖이므로 발의 지지는 확인할 수 없지만, 공중에 떠 있는 신체나 무지지 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "시선은 화면 왼쪽 바깥으로 향하고 렌즈를 직접 보지 않는다. 수화기 수화부는 화면 왼쪽 귀에 붙어 있고 송화부는 입 왼쪽 아래를 향한다. 입꼬리가 올라간 표정을 수화기가 가리지 않으며 통화 방향도 맞다.",
        "built_space": "뒤쪽 상단에 큰 분할 창 한 개, 왼쪽부터 중앙에 브라운관 화면 약 여섯 개, 왼쪽 중앙에 여러 층의 오디오 기기, 오른쪽에 케이블이 얹힌 가전 더미가 보인다. 인물은 이 적재물 앞에 있으며 가슴과 양어깨까지 포함된다. 고물상 재질과 혼잡함은 맞지만, 이전 장면의 냉장고와 세로 가전 더미 대신 창과 모니터 군집이 두드러져 동일 지점의 연속성은 덜 확실하다. 불가능한 반사나 명백히 중복된 고정 설비는 보이지 않는다.",
        "entities": "보이는 인물은 마츠다에 해당하는 나이 든 동아시아계 남성 한 명이다. 눈가 주름, 남색 작업 모자, 갈색 셔츠와 어두운 공구 조끼가 인물 참조와 잘 맞는다. 일본 국적은 이미지 자체로 판별할 수 없다. 검은 낡은 유선 수화기 한 개와 쌓인 전자제품·케이블이 보이며, 다른 사람이나 읽을 수 있는 글자는 없다. 기사 화면과 찰리는 프레임 밖으로 볼 수 있다.",
        "hard_violations": [],
        "physics": "맨손의 손가락이 수화기 중앙을 감싸고 팔뚝이 아래로 이어져 수화기의 무게를 지탱한다. 수화부와 귀의 접촉, 아래로 처진 전화선이 물리적으로 자연스럽다. 배경 모니터와 오디오 기기는 아래 적재물에 받쳐져 있고, 케이블은 가전 위에 놓여 있다. 공중에 떠 있거나 지지 없는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.542,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.542,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 1542,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1542,
    "verdict_ko": "캐릭터의 복장(모자 포함)은 잘 일치하나, 기준 이미지에 없는 창문과 다른 종류의 폐가전들을 배경에 생성하여 더 높은 우선순위인 장소 고정(Location lock) 조건을 위반했습니다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "모자가 누락되는 의상 묘사 오류가 있으나, 이전 숏 기준 이미지의 배경(금속 선반과 층층이 쌓인 가전제품) 및 어두운 분위기를 완벽하게 재현하여 장소 우선순위를 충족시킨 가장 우수한 결과물입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4_sel.png",
    "asset_id": "2e303e07-390f-47aa-892e-cba65c9d5b5a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:971613>",
    "asset_id": "945cc191-1d4f-461d-b482-4157f3024b75",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-d60f-7a15-97f3-0a7f8e05fa73",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S15sh4"
  }
 },
 "S15sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:34:50.227104+00:00",
  "fingerprint": "d91d6e3b35eda02357db1d1eb80add2c1ac55470eee6f6cbf2e57392b3fbbc4c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S15sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S15sh10_sel.png",
  "source_sha256": "e408a87f83a5d6335b5c94270d21ad5a272b2c96a3fcc4bc48e957d9ada7e024",
  "file": "S15sh10_cine.png",
  "staged_sha256": "041061ed0039a87eb268219c2049dec687e6b18380c4d216b3f89617535cc5d4",
  "latency_ms": 10564
 },
 "S16sh3::signage": {
  "fp": "a72c33f111f761d2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S16sh3": {
  "input_fingerprint": "1c8c3e6f072b31b5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A hose is spraying water outside Raul's container, leaving Charlie wet as the dirt washes off. His worn Ubik chest logo and blue-lit eyes remain; the previously established old coat and hat have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A hose is spraying water outside Raul's container, leaving Charlie wet as the dirt washes off. His worn Ubik chest logo and blue-lit eyes remain; the previously established old coat and hat have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A hose is spraying water outside Raul's container, leaving Charlie wet as the dirt washes off. His worn Ubik chest logo and blue-lit eyes remain; the previously established old coat and hat have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3__bgfirst_bg.png",
     "asset_id": "58f097e3-cb89-4739-96b1-20ba4de0ccc6",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S16sh3.png",
     "asset_id": "083f7cfb-ee18-4327-afb0-899eddfe6874",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L169B01.png",
     "asset_id": "a2e9a0db-ecb3-47e7-96eb-291c9f873ffb",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "우측 밖에서 가슴으로 향하는 물줄기가 있으나, 찰리의 왼쪽 팔목에서도 우측을 향해 알 수 없는 물줄기가 발사됨.",
    "built_space": "레퍼런스와 일치하는 컨테이너 주택 외벽과 창문이 배경에 적절한 비율로 묘사됨.",
    "entities": "찰리의 전반적인 외형은 레퍼런스와 일치하나, 가슴에 금지된 명확하게 읽을 수 있는 텍스트('Ubik')가 중앙에 배치됨.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 연출 (왼쪽 팔목에서 허공으로 뿜어져 나오는 출처 없는 물줄기)",
     "[gemini-pro] 읽을 수 있는 텍스트 노출 금지 위반 (가슴에 선명하게 적힌 'Ubik')",
     "[gpt-high] 가슴 중앙에 영문 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
    ],
    "physics": "가슴에 맞는 물줄기는 지지되나, 팔에서 뿜어져 나가는 두 번째 물줄기는 물리적 근거나 발사 지점이 전혀 없음."
   },
   {
    "label": "B",
    "direction": "화면 좌측에 보이는 호스 끝에서 찰리의 가슴을 향해 물줄기가 정확히 발사됨.",
    "built_space": "흙바닥과 컨테이너 주택의 열린 문, 창문 등 고정 요소들이 레퍼런스와 동일한 위치에 올바르게 배치됨.",
    "entities": "찰리의 마스크, 의상, 체형이 레퍼런스와 정확히 일치하며, 가슴의 로고 텍스트는 의도적으로 식별할 수 없게 찌그러져 있음.",
    "hard_violations": [
     "[gpt-high] 가슴의 영문 로고가 판독 가능하게 노출되어, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
    ],
    "physics": "호스에서 분사된 물이 몸에 부딪혀 자연스럽게 흩어지며, 물에 젖은 코트와 두 팔을 벌린 자세가 안정적으로 유지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리가 어깨를 치켜올리며 물을 맞는 모습을 잘 포착했으며, 요구된 프레이밍과 물리적 조건들을 오류 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "팔에서 알 수 없는 물줄기가 뿜어져 나오는 치명적인 물리적 오류가 있으며, 텍스트 노출 금지 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "화면 좌측에 보이는 호스 끝에서 찰리의 가슴을 향해 물줄기가 정확히 발사됨.",
        "built_space": "흙바닥과 컨테이너 주택의 열린 문, 창문 등 고정 요소들이 레퍼런스와 동일한 위치에 올바르게 배치됨.",
        "entities": "찰리의 마스크, 의상, 체형이 레퍼런스와 정확히 일치하며, 가슴의 로고 텍스트는 의도적으로 식별할 수 없게 찌그러져 있음.",
        "hard_violations": [],
        "physics": "호스에서 분사된 물이 몸에 부딪혀 자연스럽게 흩어지며, 물에 젖은 코트와 두 팔을 벌린 자세가 안정적으로 유지됨."
       },
       {
        "label": "A",
        "direction": "우측 밖에서 가슴으로 향하는 물줄기가 있으나, 찰리의 왼쪽 팔목에서도 우측을 향해 알 수 없는 물줄기가 발사됨.",
        "built_space": "레퍼런스와 일치하는 컨테이너 주택 외벽과 창문이 배경에 적절한 비율로 묘사됨.",
        "entities": "찰리의 전반적인 외형은 레퍼런스와 일치하나, 가슴에 금지된 명확하게 읽을 수 있는 텍스트('Ubik')가 중앙에 배치됨.",
        "hard_violations": [
         "물리적으로 불가능한 연출 (왼쪽 팔목에서 허공으로 뿜어져 나오는 출처 없는 물줄기)",
         "읽을 수 있는 텍스트 노출 금지 위반 (가슴에 선명하게 적힌 'Ubik')"
        ],
        "physics": "가슴에 맞는 물줄기는 지지되나, 팔에서 뿜어져 나가는 두 번째 물줄기는 물리적 근거나 발사 지점이 전혀 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리가 어깨를 치켜올리며 물을 맞는 모습을 잘 포착했으며, 요구된 프레이밍과 물리적 조건들을 오류 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "팔에서 알 수 없는 물줄기가 뿜어져 나오는 치명적인 물리적 오류가 있으며, 텍스트 노출 금지 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 좌측에 보이는 호스 끝에서 찰리의 가슴을 향해 물줄기가 정확히 발사됨.",
        "built_space": "흙바닥과 컨테이너 주택의 열린 문, 창문 등 고정 요소들이 레퍼런스와 동일한 위치에 올바르게 배치됨.",
        "entities": "찰리의 마스크, 의상, 체형이 레퍼런스와 정확히 일치하며, 가슴의 로고 텍스트는 의도적으로 식별할 수 없게 찌그러져 있음.",
        "hard_violations": [],
        "physics": "호스에서 분사된 물이 몸에 부딪혀 자연스럽게 흩어지며, 물에 젖은 코트와 두 팔을 벌린 자세가 안정적으로 유지됨."
       },
       {
        "label": "A",
        "direction": "우측 밖에서 가슴으로 향하는 물줄기가 있으나, 찰리의 왼쪽 팔목에서도 우측을 향해 알 수 없는 물줄기가 발사됨.",
        "built_space": "레퍼런스와 일치하는 컨테이너 주택 외벽과 창문이 배경에 적절한 비율로 묘사됨.",
        "entities": "찰리의 전반적인 외형은 레퍼런스와 일치하나, 가슴에 금지된 명확하게 읽을 수 있는 텍스트('Ubik')가 중앙에 배치됨.",
        "hard_violations": [
         "물리적으로 불가능한 연출 (왼쪽 팔목에서 허공으로 뿜어져 나오는 출처 없는 물줄기)",
         "읽을 수 있는 텍스트 노출 금지 위반 (가슴에 선명하게 적힌 'Ubik')"
        ],
        "physics": "가슴에 맞는 물줄기는 지지되나, 팔에서 뿜어져 나가는 두 번째 물줄기는 물리적 근거나 발사 지점이 전혀 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "상체 중심 구도와 몸통을 맞히는 측면 물줄기는 더 충실하지만, 판독 가능한 가슴 로고가 무문자 조건을 위반하며 어깨 움츠림보다 양팔을 든 동작이 강조된다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "젖은 외투와 몸통에 맞는 물줄기는 구현했지만, 가슴의 선명한 문자가 금지 조건을 위반하고 넓은 마당과 하체까지 드러내 상체 중심 미디엄 숏에서 더 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 푸른 눈은 거의 카메라 정면을 향한다. 양손은 위쪽과 바깥쪽으로 벌어져 있다. 왼쪽 아래 가장자리의 노즐에서 나온 물줄기가 찰리의 가슴 아래와 복부를 향해 실제로 부딪히며 얼굴은 가리지 않는다. 다만 호스 위치를 화면 밖에 두라는 지시와 달리 노즐 끝이 보인다.",
        "built_space": "회색 골강판 컨테이너 바로 앞에 찰리가 서 있다. 뒤에는 열린 출입문 하나와 문턱 계단, 좌우로 일부 보이는 창 두 개, 오른쪽 선반 하나와 여러 용기가 있다. 참조의 외벽 재료와 출입구 주변 구성을 대체로 유지하며, 컨테이너 일부가 상체 양옆의 배경으로 남는다. 불가능한 반사나 명백히 중복된 고정 설비는 보이지 않는다.",
        "entities": "찰리 한 명만 보인다. 육중한 장갑 몸체, 샌드 베이지색 각진 어깨와 팔 장갑, 흰 마스크형 얼굴, 낡은 모자와 외투가 참조와 대응한다. 눈은 본문의 지정대로 푸르게 빛난다. 얼굴 세부는 참조와 조금 다르며, 가슴에는 읽을 수 있는 영문 로고가 있다. 외투와 장갑에 물방울과 젖은 흔적이 있지만 흙이 씻겨 나가는 변화 자체는 뚜렷하지 않다. 일반 인간의 나이·민족성·성별을 판단할 피부는 노출되지 않는다.",
        "hard_violations": [
         "가슴의 영문 로고가 판독 가능하게 노출되어, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
        ],
        "physics": "양팔은 어깨와 팔꿈치 관절에서 굽혀 들어 올려져 있어 지지 관계가 자연스럽다. 다리와 지면 접촉은 프레임 아래에 가려져 있으나 몸이 공중에 떠 있다는 증거는 없다. 물은 가장자리 노즐에서 연속적으로 분사되어 몸통에 충돌하고 아래로 떨어진다. 노즐의 지지부는 화면 밖이라 확인할 수 없으며, 독립적으로 떠 있는 물체로 보이지는 않는다. 양어깨를 움츠리는 감각보다 팔을 크게 벌리는 제스처가 우세하다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴과 시선은 화면 오른쪽 위를 향한다. 양팔을 좌우로 벌리고 손바닥을 위로 들어 올렸다. 주 물줄기는 오른쪽 아래 화면 밖에서 왼쪽 위의 가슴으로 들어와 몸통에 충돌하며 얼굴을 피한다. 뒤쪽 오른편에는 별도로 마당 쪽을 향해 분사되는 작은 물줄기도 보인다.",
        "built_space": "찰리 뒤로 컨테이너 전면 대부분과 넓은 진흙 마당이 보인다. 식별되는 창은 왼쪽과 오른쪽에 하나씩이며 중앙 출입구와 그 주변은 몸에 상당 부분 가려진다. 왼쪽 작업대, 오른쪽 급수대와 바닥 호스, 맨 오른쪽 물탱크가 보이고 참조 장소의 주요 재료와 배치를 따른다. 다만 건물에서 더 떨어진 위치와 넓은 배경 때문에 지정된 부분 외벽 배경보다 장소 전경의 비중이 커졌다.",
        "entities": "추가 인물 없이 찰리 한 명이다. 흰 마스크형 얼굴, 푸른 눈, 베이지색 장갑, 큰 팔, 낡은 모자와 외투는 지정된 정체성과 대응한다. 가슴 중앙의 원형 부품 안에는 선명하게 읽히는 영문 로고가 들어가 참조의 부품 표현과도 달라졌다. 몸과 외투의 젖은 광택은 분명하지만 흙먼지가 씻겨 내려가는 흔적은 제한적이다. 얼굴이 가려져 일반 인간의 나이·민족성·성별은 확인할 수 없다.",
        "hard_violations": [
         "가슴 중앙에 영문 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
        ],
        "physics": "몸통은 바로 서 있고 팔은 연결된 관절을 통해 들려 있다. 발은 프레임 밖이므로 접지 상태는 확인할 수 없지만 부유 자세는 아니다. 주 물줄기는 화면 밖 공급원에서 가슴으로 이어지고 충돌한 물은 튀거나 아래로 흐른다. 배경 물줄기는 급수대 부근에서 시작하는 것으로 보인다. 물리적으로 불가능한 지지 관계는 없으나, 어깨를 치켜올리는 반응보다 양팔을 펼친 자세가 더 강하게 읽힌다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "상체 중심 구도와 몸통을 맞히는 측면 물줄기는 더 충실하지만, 판독 가능한 가슴 로고가 무문자 조건을 위반하며 어깨 움츠림보다 양팔을 든 동작이 강조된다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "젖은 외투와 몸통에 맞는 물줄기는 구현했지만, 가슴의 선명한 문자가 금지 조건을 위반하고 넓은 마당과 하체까지 드러내 상체 중심 미디엄 숏에서 더 멀어진다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 푸른 눈은 거의 카메라 정면을 향한다. 양손은 위쪽과 바깥쪽으로 벌어져 있다. 왼쪽 아래 가장자리의 노즐에서 나온 물줄기가 찰리의 가슴 아래와 복부를 향해 실제로 부딪히며 얼굴은 가리지 않는다. 다만 호스 위치를 화면 밖에 두라는 지시와 달리 노즐 끝이 보인다.",
        "built_space": "회색 골강판 컨테이너 바로 앞에 찰리가 서 있다. 뒤에는 열린 출입문 하나와 문턱 계단, 좌우로 일부 보이는 창 두 개, 오른쪽 선반 하나와 여러 용기가 있다. 참조의 외벽 재료와 출입구 주변 구성을 대체로 유지하며, 컨테이너 일부가 상체 양옆의 배경으로 남는다. 불가능한 반사나 명백히 중복된 고정 설비는 보이지 않는다.",
        "entities": "찰리 한 명만 보인다. 육중한 장갑 몸체, 샌드 베이지색 각진 어깨와 팔 장갑, 흰 마스크형 얼굴, 낡은 모자와 외투가 참조와 대응한다. 눈은 본문의 지정대로 푸르게 빛난다. 얼굴 세부는 참조와 조금 다르며, 가슴에는 읽을 수 있는 영문 로고가 있다. 외투와 장갑에 물방울과 젖은 흔적이 있지만 흙이 씻겨 나가는 변화 자체는 뚜렷하지 않다. 일반 인간의 나이·민족성·성별을 판단할 피부는 노출되지 않는다.",
        "hard_violations": [
         "가슴의 영문 로고가 판독 가능하게 노출되어, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
        ],
        "physics": "양팔은 어깨와 팔꿈치 관절에서 굽혀 들어 올려져 있어 지지 관계가 자연스럽다. 다리와 지면 접촉은 프레임 아래에 가려져 있으나 몸이 공중에 떠 있다는 증거는 없다. 물은 가장자리 노즐에서 연속적으로 분사되어 몸통에 충돌하고 아래로 떨어진다. 노즐의 지지부는 화면 밖이라 확인할 수 없으며, 독립적으로 떠 있는 물체로 보이지는 않는다. 양어깨를 움츠리는 감각보다 팔을 크게 벌리는 제스처가 우세하다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴과 시선은 화면 오른쪽 위를 향한다. 양팔을 좌우로 벌리고 손바닥을 위로 들어 올렸다. 주 물줄기는 오른쪽 아래 화면 밖에서 왼쪽 위의 가슴으로 들어와 몸통에 충돌하며 얼굴을 피한다. 뒤쪽 오른편에는 별도로 마당 쪽을 향해 분사되는 작은 물줄기도 보인다.",
        "built_space": "찰리 뒤로 컨테이너 전면 대부분과 넓은 진흙 마당이 보인다. 식별되는 창은 왼쪽과 오른쪽에 하나씩이며 중앙 출입구와 그 주변은 몸에 상당 부분 가려진다. 왼쪽 작업대, 오른쪽 급수대와 바닥 호스, 맨 오른쪽 물탱크가 보이고 참조 장소의 주요 재료와 배치를 따른다. 다만 건물에서 더 떨어진 위치와 넓은 배경 때문에 지정된 부분 외벽 배경보다 장소 전경의 비중이 커졌다.",
        "entities": "추가 인물 없이 찰리 한 명이다. 흰 마스크형 얼굴, 푸른 눈, 베이지색 장갑, 큰 팔, 낡은 모자와 외투는 지정된 정체성과 대응한다. 가슴 중앙의 원형 부품 안에는 선명하게 읽히는 영문 로고가 들어가 참조의 부품 표현과도 달라졌다. 몸과 외투의 젖은 광택은 분명하지만 흙먼지가 씻겨 내려가는 흔적은 제한적이다. 얼굴이 가려져 일반 인간의 나이·민족성·성별은 확인할 수 없다.",
        "hard_violations": [
         "가슴 중앙에 영문 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
        ],
        "physics": "몸통은 바로 서 있고 팔은 연결된 관절을 통해 들려 있다. 발은 프레임 밖이므로 접지 상태는 확인할 수 없지만 부유 자세는 아니다. 주 물줄기는 화면 밖 공급원에서 가슴으로 이어지고 충돌한 물은 튀거나 아래로 흐른다. 배경 물줄기는 급수대 부근에서 시작하는 것으로 보인다. 물리적으로 불가능한 지지 관계는 없으나, 어깨를 치켜올리는 반응보다 양팔을 펼친 자세가 더 강하게 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.179,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.929,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 연출 (왼쪽 팔목에서 허공으로 뿜어져 나오는 출처 없는 물줄기)",
     "[gemini-pro] 읽을 수 있는 텍스트 노출 금지 위반 (가슴에 선명하게 적힌 'Ubik')",
     "[gpt-high] 가슴 중앙에 영문 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
    ],
    "B": [
     "[gpt-high] 가슴의 영문 로고가 판독 가능하게 노출되어, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 929
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "찰리가 어깨를 치켜올리며 물을 맞는 모습을 잘 포착했으며, 요구된 프레이밍과 물리적 조건들을 오류 없이 훌륭하게 구현했습니다.  ★위반: [gpt-high] 가슴의 영문 로고가 판독 가능하게 노출되어, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 929,
    "verdict_ko": "팔에서 알 수 없는 물줄기가 뿜어져 나오는 치명적인 물리적 오류가 있으며, 텍스트 노출 금지 지침을 위반했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 연출 (왼쪽 팔목에서 허공으로 뿜어져 나오는 출처 없는 물줄기) / [gemini-pro] 읽을 수 있는 텍스트 노출 금지 위반 (가슴에 선명하게 적힌 'Ubik') / [gpt-high] 가슴 중앙에 영문 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 명시적 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L169B01.png",
    "asset_id": "a2e9a0db-ecb3-47e7-96eb-291c9f873ffb",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-d7be-776c-abdf-a722f3b7b8fc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3__bgfirst_bg.png",
   "bg_asset_id": "58f097e3-cb89-4739-96b1-20ba4de0ccc6",
   "bg_record_key": "S16sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S16sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:36:06.220401+00:00",
  "fingerprint": "914945cf0033b0bcab2be94e896c05da0548179c62b05cab2a49c9aba898888e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S16sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S16sh3_sel.png",
  "source_sha256": "8a2233dcc2a63655282e7b367e435a8883bfb0fe427aca9df01ab69ecf26b9aa",
  "file": "S16sh3_cine.png",
  "staged_sha256": "96019df7883bea5606265f8f5fea71ec39e25a29bd0dad7386cf0219107ca14e",
  "latency_ms": 13863
 },
 "S16sh6::signage": {
  "fp": "5b8a99dbccdef863",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S16sh6": {
  "input_fingerprint": "f15c809599196ab2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 한 손으로 턱을 괸 채 확신에 찬 눈빛으로 찰리 쪽을 주시하는 현우의 얼굴.\n\nLOCATION (lock): At the edge of the open ground outside the container home, a short distance from the robot-washing activity. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container-home exterior surroundings (Same exterior setting as the washing action) — Only soft, partial surroundings remain behind Hyunwoo; used as Maintain spatial continuity while preserving clear look room toward off-screen Charlie.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the neutral daylight of the exterior and restrained facial contrast, allowing the assured expression to carry the change in mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The washing area remains outside Raul's container, with the hose in use and Charlie's washed body wet. Charlie retains the worn chest logo, blue-lit eyes and established coat-and-hat disguise. 현우: He continues watching from a distance, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 한 손으로 턱을 괸 채 확신에 찬 눈빛으로 찰리 쪽을 주시하는 현우의 얼굴.\n\nLOCATION (lock): At the edge of the open ground outside the container home, a short distance from the robot-washing activity. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container-home exterior surroundings (Same exterior setting as the washing action) — Only soft, partial surroundings remain behind Hyunwoo; used as Maintain spatial continuity while preserving clear look room toward off-screen Charlie.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the neutral daylight of the exterior and restrained facial contrast, allowing the assured expression to carry the change in mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The washing area remains outside Raul's container, with the hose in use and Charlie's washed body wet. Charlie retains the worn chest logo, blue-lit eyes and established coat-and-hat disguise. 현우: He continues watching from a distance, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 한 손으로 턱을 괸 채 확신에 찬 눈빛으로 찰리 쪽을 주시하는 현우의 얼굴.\n\nLOCATION (lock): At the edge of the open ground outside the container home, a short distance from the robot-washing activity. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container-home exterior surroundings (Same exterior setting as the washing action) — Only soft, partial surroundings remain behind Hyunwoo; used as Maintain spatial continuity while preserving clear look room toward off-screen Charlie.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the neutral daylight of the exterior and restrained facial contrast, allowing the assured expression to carry the change in mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The washing area remains outside Raul's container, with the hose in use and Charlie's washed body wet. Charlie retains the worn chest logo, blue-lit eyes and established coat-and-hat disguise. 현우: He continues watching from a distance, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "화면 밖 왼쪽 오프스크린 대상을 향해 확신에 찬 시선을 고정하고 있음.",
    "built_space": "컨테이너 벽과 문 가장자리가 아웃포커싱되어 부드럽고 부분적인 배경 지침을 잘 충족함.",
    "entities": "현우의 앳된 얼굴, 상처, 회색 셔츠, 그리고 오른쪽 귀에 착용한 이어피스가 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "오른손이 턱과 목 부분을 자연스럽게 받치고 자세를 유지함."
   },
   {
    "label": "B",
    "direction": "화면 밖 왼쪽 오프스크린 대상을 향해 시선을 고정하고 있음.",
    "built_space": "컨테이너 문과 금속 테이블, 양동이 등 배경이 너무 선명하게 넓게 노출되어 지침에 어긋남.",
    "entities": "얼굴 상처와 의상은 일치하나, 이어피스가 레퍼런스와 반대인 왼쪽 귀에 착용됨.",
    "hard_violations": [],
    "physics": "왼손을 들어 턱을 온전히 받치고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 프레이밍과 아웃포커싱된 배경 지침을 잘 따랐으며, 레퍼런스와 동일하게 오른쪽 귀에 이어피스를 착용한 점이 우수합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경이 너무 선명하고 넓게 프레이밍되어 샷 지침을 위반했으며, 이어피스를 왼쪽 귀에 잘못 착용하였습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 밖 왼쪽 오프스크린 대상을 향해 확신에 찬 시선을 고정하고 있음.",
        "built_space": "컨테이너 벽과 문 가장자리가 아웃포커싱되어 부드럽고 부분적인 배경 지침을 잘 충족함.",
        "entities": "현우의 앳된 얼굴, 상처, 회색 셔츠, 그리고 오른쪽 귀에 착용한 이어피스가 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "오른손이 턱과 목 부분을 자연스럽게 받치고 자세를 유지함."
       },
       {
        "label": "B",
        "direction": "화면 밖 왼쪽 오프스크린 대상을 향해 시선을 고정하고 있음.",
        "built_space": "컨테이너 문과 금속 테이블, 양동이 등 배경이 너무 선명하게 넓게 노출되어 지침에 어긋남.",
        "entities": "얼굴 상처와 의상은 일치하나, 이어피스가 레퍼런스와 반대인 왼쪽 귀에 착용됨.",
        "hard_violations": [],
        "physics": "왼손을 들어 턱을 온전히 받치고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 프레이밍과 아웃포커싱된 배경 지침을 잘 따랐으며, 레퍼런스와 동일하게 오른쪽 귀에 이어피스를 착용한 점이 우수합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경이 너무 선명하고 넓게 프레이밍되어 샷 지침을 위반했으며, 이어피스를 왼쪽 귀에 잘못 착용하였습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "화면 밖 왼쪽 오프스크린 대상을 향해 확신에 찬 시선을 고정하고 있음.",
        "built_space": "컨테이너 벽과 문 가장자리가 아웃포커싱되어 부드럽고 부분적인 배경 지침을 잘 충족함.",
        "entities": "현우의 앳된 얼굴, 상처, 회색 셔츠, 그리고 오른쪽 귀에 착용한 이어피스가 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "오른손이 턱과 목 부분을 자연스럽게 받치고 자세를 유지함."
       },
       {
        "label": "B",
        "direction": "화면 밖 왼쪽 오프스크린 대상을 향해 시선을 고정하고 있음.",
        "built_space": "컨테이너 문과 금속 테이블, 양동이 등 배경이 너무 선명하게 넓게 노출되어 지침에 어긋남.",
        "entities": "얼굴 상처와 의상은 일치하나, 이어피스가 레퍼런스와 반대인 왼쪽 귀에 착용됨.",
        "hard_violations": [],
        "physics": "왼손을 들어 턱을 온전히 받치고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "턱을 괴고 화면 밖 찰리를 바라보는 동작과 시선 여백은 맞지만, 배경 집기가 너무 많이 드러나고 얼굴의 명암이 강해 요구된 얼굴 중심 클로즈업과 절제된 낮 조명에서 B보다 멀다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "한 손으로 턱을 받친 현우의 얼굴을 밀착해 담고, 화면 밖을 향한 확신 있는 시선과 부드러운 컨테이너 배경, 중립적인 낮 조명을 가장 충실히 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 두 눈은 화면 왼쪽 바깥을 향한다. 찰리 자체는 보이지 않지만, 화면 밖 찰리를 주시한다는 연출에 맞으며 왼쪽 시선 여백도 충분하다. 무기나 방향을 확인할 별도 소품은 없다.",
        "built_space": "회색 골판 외벽, 열린 출입구 하나, 그 아래 낮은 계단, 머리 뒤 창틀 일부가 보인다. 화면 왼쪽에는 선반 하나와 여러 용기, 오른쪽에는 회색 수납함과 상자들이 보인다. 현우는 외벽 앞의 야외 전경에 있다. 이전 장면과 재질은 유사하지만 선반이 출입구 반대쪽에 나타나 공간 연속성이 약하며, 배경이 요구된 부드러운 일부 주변보다 넓고 구체적으로 드러난다. 반사는 없다.",
        "entities": "인물은 현우 한 명뿐이다. 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 회색 겉셔츠와 귀의 통신 장치가 인물 참조와 대체로 맞는다. 코와 뺨, 눈가 및 입술의 상처가 남아 있다. 다리 부상은 프레임 밖이므로 확인 대상이 아니다. 찰리와 세척 호스는 화면 밖에 있으며 다른 사람이나 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "한 손의 구부린 손가락과 주먹 윗부분이 턱 아래에 직접 닿아 머리를 받친다. 손목과 팔은 셔츠 소매로 자연스럽게 이어진다. 팔꿈치와 하체의 지지점은 화면 밖이지만, 보이는 부분에 공중 부양이나 불가능한 관절 배치는 없다. 배경 용기는 선반 또는 지면 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이지만 두 눈의 시선은 렌즈보다 화면 왼쪽 바깥으로 향한다. 실제 목표인 찰리는 프레임 밖이므로 직접 확인할 수 없으나, 그쪽을 바라보는 장면으로 자연스럽게 읽힌다. 왼쪽에 시선이 이어질 여백이 있다.",
        "built_space": "부드럽게 흐려진 회색 골판 외벽 뒤로 왼쪽 출입구 일부 하나와 오른쪽 위 창문 일부 하나만 보인다. 현우는 그 외부 전경에 위치하며, 얼굴과 어깨가 프레임 대부분을 차지한다. 이전 장면의 컨테이너 재질과 출입구·창문의 관계를 유지하면서 배경을 부분적으로만 남긴다. 중복된 고정 설비나 부자연스러운 반사는 보이지 않는다.",
        "entities": "현우 한 명만 등장한다. 참조에 부합하는 앳된 동아시아계 남성 얼굴, 헝클어진 검은 머리, 회색 겉셔츠와 나선형 통신 장치 선이 보인다. 코, 양쪽 뺨, 눈가와 입술에 낫지 않은 상처가 표현돼 있다. 눈은 정상적인 홍채와 동공을 유지하며, 다문 입과 안정된 눈빛이 확신 있는 표정으로 읽힌다. 다리와 찰리의 세척 상태는 올바르게 프레임에서 제외되어 있고 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "구부린 한 손이 턱 아래와 턱선을 실제로 받치고 있으며, 손가락·손목·팔의 연결이 자연스럽다. 팔은 화면 아래로 이어지고 팔꿈치와 좌석 여부는 보이지 않는다. 보이지 않는 지지점을 단정할 수는 없지만, 드러난 신체에 지지 없는 부양이나 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "턱을 괴고 화면 밖 찰리를 바라보는 동작과 시선 여백은 맞지만, 배경 집기가 너무 많이 드러나고 얼굴의 명암이 강해 요구된 얼굴 중심 클로즈업과 절제된 낮 조명에서 B보다 멀다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "한 손으로 턱을 받친 현우의 얼굴을 밀착해 담고, 화면 밖을 향한 확신 있는 시선과 부드러운 컨테이너 배경, 중립적인 낮 조명을 가장 충실히 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 두 눈은 화면 왼쪽 바깥을 향한다. 찰리 자체는 보이지 않지만, 화면 밖 찰리를 주시한다는 연출에 맞으며 왼쪽 시선 여백도 충분하다. 무기나 방향을 확인할 별도 소품은 없다.",
        "built_space": "회색 골판 외벽, 열린 출입구 하나, 그 아래 낮은 계단, 머리 뒤 창틀 일부가 보인다. 화면 왼쪽에는 선반 하나와 여러 용기, 오른쪽에는 회색 수납함과 상자들이 보인다. 현우는 외벽 앞의 야외 전경에 있다. 이전 장면과 재질은 유사하지만 선반이 출입구 반대쪽에 나타나 공간 연속성이 약하며, 배경이 요구된 부드러운 일부 주변보다 넓고 구체적으로 드러난다. 반사는 없다.",
        "entities": "인물은 현우 한 명뿐이다. 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 회색 겉셔츠와 귀의 통신 장치가 인물 참조와 대체로 맞는다. 코와 뺨, 눈가 및 입술의 상처가 남아 있다. 다리 부상은 프레임 밖이므로 확인 대상이 아니다. 찰리와 세척 호스는 화면 밖에 있으며 다른 사람이나 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "한 손의 구부린 손가락과 주먹 윗부분이 턱 아래에 직접 닿아 머리를 받친다. 손목과 팔은 셔츠 소매로 자연스럽게 이어진다. 팔꿈치와 하체의 지지점은 화면 밖이지만, 보이는 부분에 공중 부양이나 불가능한 관절 배치는 없다. 배경 용기는 선반 또는 지면 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이지만 두 눈의 시선은 렌즈보다 화면 왼쪽 바깥으로 향한다. 실제 목표인 찰리는 프레임 밖이므로 직접 확인할 수 없으나, 그쪽을 바라보는 장면으로 자연스럽게 읽힌다. 왼쪽에 시선이 이어질 여백이 있다.",
        "built_space": "부드럽게 흐려진 회색 골판 외벽 뒤로 왼쪽 출입구 일부 하나와 오른쪽 위 창문 일부 하나만 보인다. 현우는 그 외부 전경에 위치하며, 얼굴과 어깨가 프레임 대부분을 차지한다. 이전 장면의 컨테이너 재질과 출입구·창문의 관계를 유지하면서 배경을 부분적으로만 남긴다. 중복된 고정 설비나 부자연스러운 반사는 보이지 않는다.",
        "entities": "현우 한 명만 등장한다. 참조에 부합하는 앳된 동아시아계 남성 얼굴, 헝클어진 검은 머리, 회색 겉셔츠와 나선형 통신 장치 선이 보인다. 코, 양쪽 뺨, 눈가와 입술에 낫지 않은 상처가 표현돼 있다. 눈은 정상적인 홍채와 동공을 유지하며, 다문 입과 안정된 눈빛이 확신 있는 표정으로 읽힌다. 다리와 찰리의 세척 상태는 올바르게 프레임에서 제외되어 있고 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "구부린 한 손이 턱 아래와 턱선을 실제로 받치고 있으며, 손가락·손목·팔의 연결이 자연스럽다. 팔은 화면 아래로 이어지고 팔꿈치와 좌석 여부는 보이지 않는다. 보이지 않는 지지점을 단정할 수는 없지만, 드러난 신체에 지지 없는 부양이나 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.349
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.349
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1349
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 클로즈업 프레이밍과 아웃포커싱된 배경 지침을 잘 따랐으며, 레퍼런스와 동일하게 오른쪽 귀에 이어피스를 착용한 점이 우수합니다."
   },
   {
    "label": "B",
    "score": 1349,
    "verdict_ko": "배경이 너무 선명하고 넓게 프레이밍되어 샷 지침을 위반했으며, 이어피스를 왼쪽 귀에 잘못 착용하였습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3_sel.png",
    "asset_id": "f2009afb-8ddc-4a1d-9b8b-447d993f8193",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-db0b-74e1-ba70-74f96bf700e1",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S16sh3"
  }
 },
 "S16sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:37:01.638072+00:00",
  "fingerprint": "59d990b5d92a3f421a2eaf3d596cef68e250475f03f56c6ceb6a40a54fc3ef62",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S16sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S16sh6_sel.png",
  "source_sha256": "49a9dc107330061d689772e42d66323991184f0ad7f31cadf5a4d56feee28ad0",
  "file": "S16sh6_cine.png",
  "staged_sha256": "fa2a60940aaa783b84c38c4b47c1123ddd33cc4f262920b1b2b4a20eaafdcd49",
  "latency_ms": 11773
 },
 "S17sh1::signage": {
  "fp": "d571c1e172ce9549",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S17sh1": {
  "input_fingerprint": "371f6c7ffbb3e5d0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 상점 입구에 들어선 무장한 수하 1의 검은 실루엣 전신.\n\nLOCATION (lock): Just inside the junk shop entrance, silhouetted against the brighter daytime doorway with piled appliances farther inside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Shop entrance (Being crossed by the entering subordinate) — Seen obliquely from inside, with the entrance boundaries surrounding his full figure; used as Frame the arrival and establish the route toward Matsuda without implying a door-leaf position; Silenced rifle (Carried by the subordinate, not firing) — Its outline remains adjacent to his body rather than aimed into the lens; used as Make the armed arrival legible without staging the later shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the described dark silhouette through controlled entrance-to-interior tonal separation without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains cluttered with appliances and scrap, with a television currently under repair. 수하 1: He has entered the shop and is looking around.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 상점 입구에 들어선 무장한 수하 1의 검은 실루엣 전신.\n\nLOCATION (lock): Just inside the junk shop entrance, silhouetted against the brighter daytime doorway with piled appliances farther inside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Shop entrance (Being crossed by the entering subordinate) — Seen obliquely from inside, with the entrance boundaries surrounding his full figure; used as Frame the arrival and establish the route toward Matsuda without implying a door-leaf position; Silenced rifle (Carried by the subordinate, not firing) — Its outline remains adjacent to his body rather than aimed into the lens; used as Make the armed arrival legible without staging the later shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the described dark silhouette through controlled entrance-to-interior tonal separation without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains cluttered with appliances and scrap, with a television currently under repair. 수하 1: He has entered the shop and is looking around.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 상점 입구에 들어선 무장한 수하 1의 검은 실루엣 전신.\n\nLOCATION (lock): Just inside the junk shop entrance, silhouetted against the brighter daytime doorway with piled appliances farther inside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Shop entrance (Being crossed by the entering subordinate) — Seen obliquely from inside, with the entrance boundaries surrounding his full figure; used as Frame the arrival and establish the route toward Matsuda without implying a door-leaf position; Silenced rifle (Carried by the subordinate, not firing) — Its outline remains adjacent to his body rather than aimed into the lens; used as Make the armed arrival legible without staging the later shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the described dark silhouette through controlled entrance-to-interior tonal separation without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains cluttered with appliances and scrap, with a television currently under repair. 수하 1: He has entered the shop and is looking around.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물은 카메라가 위치한 상점 내부를 향해 걸어 들어오고 있으며 시선 역시 정면을 향함. 오른손에 든 총구는 바닥을 향하고 있음.",
    "built_space": "고물상 내부로, 좌우로 오래된 CRT TV, 전자레인지, 냉장고 등이 쌓여 있음. 인물의 등 뒤로 밝은 빛이 들어오는 출입구가 위치하여 정확한 역광 환경을 구성함.",
    "entities": "지정된 체형과 모자를 착용한 수하 1의 전신이 완전한 검은 실루엣으로 표현됨. 오른손에는 소음기가 달린 총(프롭 레퍼런스 형태)을 들고 있음.",
    "hard_violations": [],
    "physics": "오른발이 앞으로 향하고 왼발이 뒤에 있는 자연스러운 보행 자세로 바닥에 지지되어 있으며, 오른손으로 총기의 손잡이를 안정적으로 쥐고 있음."
   },
   {
    "label": "B",
    "direction": "인물은 상점 내부가 아닌 바깥쪽(출입구)을 향해 등지고 서 있음. 오른손에 든 총구는 바닥을 향함.",
    "built_space": "고물상 내부로, 오래된 TV와 가전제품들이 선반과 바닥에 배치되어 있음. 인물 앞쪽으로 바깥 풍경이 보이는 출입구가 있음.",
    "entities": "수하 1의 체형과 모자, 유니폼이 묘사되었으나 완전한 실루엣이 되지 못하고 오른팔의 붉은 완장 등 색상이 드러남. 소음기가 달린 총을 들고 있음.",
    "hard_violations": [
     "[gemini-pro] 프롬프트는 '상점 입구에 들어선(entering)' 인물을 내부에서 바라보는 시점(Seen obliquely from inside)으로 지정했으나, 인물이 내부에서 바깥으로 나가는 방향을 향해 서 있어 동선과 무대 연출이 완전히 반대로 적용됨."
    ],
    "physics": "두 발로 바닥을 딛고 서 있으며, 오른손으로 총을 자연스럽게 쥐고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "상점 내부에서 입구를 향해 들어오는 인물의 동선과 요구된 완벽한 검은 실루엣, 소음기 장착 무기 및 고물상 배경을 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "내부로 들어오는(entering) 상황임에도 인물이 바깥을 향해 서 있으며, 완전한 검은 실루엣이 아닌 유니폼의 색상(붉은 완장)이 노출되어 조명 및 연출 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 카메라가 위치한 상점 내부를 향해 걸어 들어오고 있으며 시선 역시 정면을 향함. 오른손에 든 총구는 바닥을 향하고 있음.",
        "built_space": "고물상 내부로, 좌우로 오래된 CRT TV, 전자레인지, 냉장고 등이 쌓여 있음. 인물의 등 뒤로 밝은 빛이 들어오는 출입구가 위치하여 정확한 역광 환경을 구성함.",
        "entities": "지정된 체형과 모자를 착용한 수하 1의 전신이 완전한 검은 실루엣으로 표현됨. 오른손에는 소음기가 달린 총(프롭 레퍼런스 형태)을 들고 있음.",
        "hard_violations": [],
        "physics": "오른발이 앞으로 향하고 왼발이 뒤에 있는 자연스러운 보행 자세로 바닥에 지지되어 있으며, 오른손으로 총기의 손잡이를 안정적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "인물은 상점 내부가 아닌 바깥쪽(출입구)을 향해 등지고 서 있음. 오른손에 든 총구는 바닥을 향함.",
        "built_space": "고물상 내부로, 오래된 TV와 가전제품들이 선반과 바닥에 배치되어 있음. 인물 앞쪽으로 바깥 풍경이 보이는 출입구가 있음.",
        "entities": "수하 1의 체형과 모자, 유니폼이 묘사되었으나 완전한 실루엣이 되지 못하고 오른팔의 붉은 완장 등 색상이 드러남. 소음기가 달린 총을 들고 있음.",
        "hard_violations": [
         "프롬프트는 '상점 입구에 들어선(entering)' 인물을 내부에서 바라보는 시점(Seen obliquely from inside)으로 지정했으나, 인물이 내부에서 바깥으로 나가는 방향을 향해 서 있어 동선과 무대 연출이 완전히 반대로 적용됨."
        ],
        "physics": "두 발로 바닥을 딛고 서 있으며, 오른손으로 총을 자연스럽게 쥐고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "상점 내부에서 입구를 향해 들어오는 인물의 동선과 요구된 완벽한 검은 실루엣, 소음기 장착 무기 및 고물상 배경을 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "내부로 들어오는(entering) 상황임에도 인물이 바깥을 향해 서 있으며, 완전한 검은 실루엣이 아닌 유니폼의 색상(붉은 완장)이 노출되어 조명 및 연출 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 카메라가 위치한 상점 내부를 향해 걸어 들어오고 있으며 시선 역시 정면을 향함. 오른손에 든 총구는 바닥을 향하고 있음.",
        "built_space": "고물상 내부로, 좌우로 오래된 CRT TV, 전자레인지, 냉장고 등이 쌓여 있음. 인물의 등 뒤로 밝은 빛이 들어오는 출입구가 위치하여 정확한 역광 환경을 구성함.",
        "entities": "지정된 체형과 모자를 착용한 수하 1의 전신이 완전한 검은 실루엣으로 표현됨. 오른손에는 소음기가 달린 총(프롭 레퍼런스 형태)을 들고 있음.",
        "hard_violations": [],
        "physics": "오른발이 앞으로 향하고 왼발이 뒤에 있는 자연스러운 보행 자세로 바닥에 지지되어 있으며, 오른손으로 총기의 손잡이를 안정적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "인물은 상점 내부가 아닌 바깥쪽(출입구)을 향해 등지고 서 있음. 오른손에 든 총구는 바닥을 향함.",
        "built_space": "고물상 내부로, 오래된 TV와 가전제품들이 선반과 바닥에 배치되어 있음. 인물 앞쪽으로 바깥 풍경이 보이는 출입구가 있음.",
        "entities": "수하 1의 체형과 모자, 유니폼이 묘사되었으나 완전한 실루엣이 되지 못하고 오른팔의 붉은 완장 등 색상이 드러남. 소음기가 달린 총을 들고 있음.",
        "hard_violations": [
         "프롬프트는 '상점 입구에 들어선(entering)' 인물을 내부에서 바라보는 시점(Seen obliquely from inside)으로 지정했으나, 인물이 내부에서 바깥으로 나가는 방향을 향해 서 있어 동선과 무대 연출이 완전히 반대로 적용됨."
        ],
        "physics": "두 발로 바닥을 딛고 서 있으며, 오른손으로 총을 자연스럽게 쥐고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "전신과 낮의 입구 역광은 맞지만, 인물이 등을 보이고 바깥으로 걷고 있어 상점에 들어선 순간이라는 핵심 동선이 반대다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "밝은 출입구를 등지고 안으로 들어오는 무장 수하의 전신 실루엣을 구현했지만, 비스듬한 관찰 각도와 주변을 살피는 시선은 부족하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "머리와 몸의 앞면이 밝은 실외를 향하고, 카메라에는 등이 보인다. 문턱 쪽으로 내디딘 발도 퇴장 방향으로 읽혀 입장 동선과 반대다. 오른손의 총구는 몸 옆에서 아래쪽 바닥을 향하며 렌즈나 사람을 겨누지 않는다.",
        "built_space": "실내에서 출입구 한 곳을 거의 정면으로 본다. 왼쪽에는 열린 문짝, 오른쪽에는 유리 구획이 있으며 전신이 입구 경계 안에 들어간다. 왼쪽 냉장고 한 대와 그 위의 브라운관 기기, 바닥 쪽에 모인 브라운관 세 대가 보이고, 오른쪽 작업대 한 곳에는 외장이 열린 수리 중 텔레비전 한 대가 놓여 있다. 낡은 가전과 금속 선반의 재질은 장소 참고와 유사하지만 요구한 사선 시점은 약하다.",
        "entities": "성인 남성 한 명만 있다. 모자, 어두운 작업복, 붉은 완장, 허리 장비와 부츠는 인물 참고에 부합한다. 뒷모습이라 얼굴과 한국인 인물의 정확한 동일성은 확인할 수 없다. 소음기가 달린 긴 총 한 정은 보이나, 참고의 개머리판 부착 권총형 몸체와 정확히 일치하는지는 불분명하다. 수리 중 텔레비전과 가전 더미가 보이며 이전 장면의 노인은 없다.",
        "hard_violations": [],
        "physics": "한쪽 부츠는 실내 바닥에, 다른 쪽은 문턱 부근에 닿아 보행 체중을 지지한다. 총은 오른손으로 잡고 있어 공중에 떠 있지 않다. 가전은 바닥·선반·작업대 또는 아래 가전 위에 놓여 있다. 보행 자체는 물리적으로 가능하지만 입장보다 퇴장 동작으로 읽힌다."
       },
       {
        "label": "B",
        "direction": "몸과 얼굴은 실외에서 실내 카메라 쪽으로 향하고 한 발을 안으로 내딛는다. 따라서 입장 방향은 맞는다. 얼굴이 어두워 정확한 눈동자 방향은 확인하기 어렵고, 고개를 돌려 주변을 살피는 행동은 뚜렷하지 않다. 화면 왼쪽 손에 든 총의 소음기 끝은 아래 바닥을 향하며 렌즈를 겨누지 않는다.",
        "built_space": "출입구 한 곳의 양옆 유리 구획과 상부 창이 전신을 둘러싼다. 카메라는 실내에 있으나 입구를 거의 정면으로 본다. 왼쪽에는 냉장고 한 대와 여러 층의 낡은 가전, 오른쪽에는 선반과 작업면 및 브라운관 기기들이 있다. 이 적층 가전 배치는 참고 장소와 더 가깝다. 수리 대상처럼 놓인 텔레비전은 있으나 분해·수리 상태는 A만큼 명확하지 않다. 불가능한 반사나 출입구 중복은 보이지 않는다.",
        "entities": "모자를 쓴 성인 남성 한 명이며 참고와 비슷한 체격, 작업복, 허리 파우치, 부츠의 윤곽이 보인다. 짙은 역광 때문에 얼굴·머리색·완장 색의 정확한 일치는 판별할 수 없다. 몸에는 인간의 머리와 사지, 옷의 입체적인 윤곽이 유지된다. 몸 옆 총 한 정에 긴 소음기가 보이지만 참고 총의 세부 구조와 표면 마모는 식별되지 않는다. 낡은 가전과 잡동사니가 있고 다른 사람은 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 부츠가 실내 바닥에 닿아 체중을 받으며 반대쪽 다리는 무릎이 굽혀진 보행 중간 상태다. 든 발이 떠 있어도 지지 발이 있으므로 부유가 아니다. 총은 손으로 잡혀 있고 가전들은 선반·바닥·다른 가전 위에 받쳐져 있다. 문턱을 넘어 들어오는 동작으로 물리적으로 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "전신과 낮의 입구 역광은 맞지만, 인물이 등을 보이고 바깥으로 걷고 있어 상점에 들어선 순간이라는 핵심 동선이 반대다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "밝은 출입구를 등지고 안으로 들어오는 무장 수하의 전신 실루엣을 구현했지만, 비스듬한 관찰 각도와 주변을 살피는 시선은 부족하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "머리와 몸의 앞면이 밝은 실외를 향하고, 카메라에는 등이 보인다. 문턱 쪽으로 내디딘 발도 퇴장 방향으로 읽혀 입장 동선과 반대다. 오른손의 총구는 몸 옆에서 아래쪽 바닥을 향하며 렌즈나 사람을 겨누지 않는다.",
        "built_space": "실내에서 출입구 한 곳을 거의 정면으로 본다. 왼쪽에는 열린 문짝, 오른쪽에는 유리 구획이 있으며 전신이 입구 경계 안에 들어간다. 왼쪽 냉장고 한 대와 그 위의 브라운관 기기, 바닥 쪽에 모인 브라운관 세 대가 보이고, 오른쪽 작업대 한 곳에는 외장이 열린 수리 중 텔레비전 한 대가 놓여 있다. 낡은 가전과 금속 선반의 재질은 장소 참고와 유사하지만 요구한 사선 시점은 약하다.",
        "entities": "성인 남성 한 명만 있다. 모자, 어두운 작업복, 붉은 완장, 허리 장비와 부츠는 인물 참고에 부합한다. 뒷모습이라 얼굴과 한국인 인물의 정확한 동일성은 확인할 수 없다. 소음기가 달린 긴 총 한 정은 보이나, 참고의 개머리판 부착 권총형 몸체와 정확히 일치하는지는 불분명하다. 수리 중 텔레비전과 가전 더미가 보이며 이전 장면의 노인은 없다.",
        "hard_violations": [],
        "physics": "한쪽 부츠는 실내 바닥에, 다른 쪽은 문턱 부근에 닿아 보행 체중을 지지한다. 총은 오른손으로 잡고 있어 공중에 떠 있지 않다. 가전은 바닥·선반·작업대 또는 아래 가전 위에 놓여 있다. 보행 자체는 물리적으로 가능하지만 입장보다 퇴장 동작으로 읽힌다."
       },
       {
        "label": "A",
        "direction": "몸과 얼굴은 실외에서 실내 카메라 쪽으로 향하고 한 발을 안으로 내딛는다. 따라서 입장 방향은 맞는다. 얼굴이 어두워 정확한 눈동자 방향은 확인하기 어렵고, 고개를 돌려 주변을 살피는 행동은 뚜렷하지 않다. 화면 왼쪽 손에 든 총의 소음기 끝은 아래 바닥을 향하며 렌즈를 겨누지 않는다.",
        "built_space": "출입구 한 곳의 양옆 유리 구획과 상부 창이 전신을 둘러싼다. 카메라는 실내에 있으나 입구를 거의 정면으로 본다. 왼쪽에는 냉장고 한 대와 여러 층의 낡은 가전, 오른쪽에는 선반과 작업면 및 브라운관 기기들이 있다. 이 적층 가전 배치는 참고 장소와 더 가깝다. 수리 대상처럼 놓인 텔레비전은 있으나 분해·수리 상태는 A만큼 명확하지 않다. 불가능한 반사나 출입구 중복은 보이지 않는다.",
        "entities": "모자를 쓴 성인 남성 한 명이며 참고와 비슷한 체격, 작업복, 허리 파우치, 부츠의 윤곽이 보인다. 짙은 역광 때문에 얼굴·머리색·완장 색의 정확한 일치는 판별할 수 없다. 몸에는 인간의 머리와 사지, 옷의 입체적인 윤곽이 유지된다. 몸 옆 총 한 정에 긴 소음기가 보이지만 참고 총의 세부 구조와 표면 마모는 식별되지 않는다. 낡은 가전과 잡동사니가 있고 다른 사람은 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 부츠가 실내 바닥에 닿아 체중을 받으며 반대쪽 다리는 무릎이 굽혀진 보행 중간 상태다. 든 발이 떠 있어도 지지 발이 있으므로 부유가 아니다. 총은 손으로 잡혀 있고 가전들은 선반·바닥·다른 가전 위에 받쳐져 있다. 문턱을 넘어 들어오는 동작으로 물리적으로 자연스럽다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.833
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.583
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트는 '상점 입구에 들어선(entering)' 인물을 내부에서 바라보는 시점(Seen obliquely from inside)으로 지정했으나, 인물이 내부에서 바깥으로 나가는 방향을 향해 서 있어 동선과 무대 연출이 완전히 반대로 적용됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 583
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "상점 내부에서 입구를 향해 들어오는 인물의 동선과 요구된 완벽한 검은 실루엣, 소음기 장착 무기 및 고물상 배경을 매우 충실하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 583,
    "verdict_ko": "내부로 들어오는(entering) 상황임에도 인물이 바깥을 향해 서 있으며, 완전한 검은 실루엣이 아닌 유니폼의 색상(붉은 완장)이 노출되어 조명 및 연출 지침을 위반했습니다.  ★위반: [gemini-pro] 프롬프트는 '상점 입구에 들어선(entering)' 인물을 내부에서 바라보는 시점(Seen obliquely from inside)으로 지정했으나, 인물이 내부에서 바깥으로 나가는 방향을 향해 서 있어 동선과 무대 연출이 완전히 반대로 적용됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh10_sel.png",
    "asset_id": "ab5c1c26-cb72-4e8e-a3c4-2b6936d7eb42",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:798636>",
    "asset_id": "840c8473-e2ef-41b8-ac3f-988134fc14e7",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-dcac-7647-8b70-20714e18142a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S15sh10"
  }
 },
 "S17sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:38:04.742168+00:00",
  "fingerprint": "38c6832f03b36ef1b2c7355c6926029581c2717b78cb3f817b79c90d3d64a48d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S17sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S17sh1_sel.png",
  "source_sha256": "4362e9bfc5f78bcd7aeb996a98e4dc61d141041a76dc961d807a04719ce32404",
  "file": "S17sh1_cine.png",
  "staged_sha256": "975651cbdf48d69226c8570806d9aabecebf14bb132000880b20bbe5741d38ff",
  "latency_ms": 11674
 },
 "S17sh7::signage": {
  "fp": "695a031b3956ab20",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S17sh7": {
  "input_fingerprint": "ba789aaa88850292",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰러진 마츠다를 등진 채 무전기를 입가에 바짝 댄 수하 1의 차가운 옆얼굴.\n\nLOCATION (lock): Beside the repair desk inside the junk shop, with the fallen owner behind it in subdued daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Desk supporting the collapsed Matsuda, behind the subordinate's back in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Desk beneath Matsuda (Supporting Matsuda's collapsed upper body) — The desktop is visible obliquely in the right background behind the subordinate; used as Provide an unmistakable shared-space anchor for the aftermath; Radio (Held at the subordinate's mouth during his report) — Seen from the side, below the nose and beside the lips; used as Identify the reporting action without concealing the cold profile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained shop ambience with enough tonal separation to distinguish the speaking profile from Matsuda's collapsed body.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Matsuda is slumped forward over the desk, his collapsed upper body resting on the desktop after being shot. The source does not specify his head's turn, the placement of his arms, or the position of his lower body.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The television repair and surrounding appliance-and-scrap clutter remain in place. 마츠다: He is slumped over the desk after being shot. 수하 1: He remains inside the shop, holding and using a radio.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수하 1 right now, so 수하 1's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수하 1: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰러진 마츠다를 등진 채 무전기를 입가에 바짝 댄 수하 1의 차가운 옆얼굴.\n\nLOCATION (lock): Beside the repair desk inside the junk shop, with the fallen owner behind it in subdued daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Desk supporting the collapsed Matsuda, behind the subordinate's back in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Desk beneath Matsuda (Supporting Matsuda's collapsed upper body) — The desktop is visible obliquely in the right background behind the subordinate; used as Provide an unmistakable shared-space anchor for the aftermath; Radio (Held at the subordinate's mouth during his report) — Seen from the side, below the nose and beside the lips; used as Identify the reporting action without concealing the cold profile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained shop ambience with enough tonal separation to distinguish the speaking profile from Matsuda's collapsed body.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Matsuda is slumped forward over the desk, his collapsed upper body resting on the desktop after being shot. The source does not specify his head's turn, the placement of his arms, or the position of his lower body.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The television repair and surrounding appliance-and-scrap clutter remain in place. 마츠다: He is slumped over the desk after being shot. 수하 1: He remains inside the shop, holding and using a radio.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수하 1 right now, so 수하 1's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수하 1: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰러진 마츠다를 등진 채 무전기를 입가에 바짝 댄 수하 1의 차가운 옆얼굴.\n\nLOCATION (lock): Beside the repair desk inside the junk shop, with the fallen owner behind it in subdued daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Desk supporting the collapsed Matsuda, behind the subordinate's back in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Desk beneath Matsuda (Supporting Matsuda's collapsed upper body) — The desktop is visible obliquely in the right background behind the subordinate; used as Provide an unmistakable shared-space anchor for the aftermath; Radio (Held at the subordinate's mouth during his report) — Seen from the side, below the nose and beside the lips; used as Identify the reporting action without concealing the cold profile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained shop ambience with enough tonal separation to distinguish the speaking profile from Matsuda's collapsed body.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Matsuda is slumped forward over the desk, his collapsed upper body resting on the desktop after being shot. The source does not specify his head's turn, the placement of his arms, or the position of his lower body.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The television repair and surrounding appliance-and-scrap clutter remain in place. 마츠다: He is slumped over the desk after being shot. 수하 1: He remains inside the shop, holding and using a radio.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수하 1 right now, so 수하 1's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수하 1: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수하 1은 화면 왼쪽을 향해 시선을 두고 무전기를 입가에 대고 있으며, 마츠다는 그의 등 뒤인 화면 우측 방향에 위치함.",
    "built_space": "고물상 내부로 배경에 가전제품과 선반이 보이며, 우측 배경에 쓰러진 마츠다를 받치고 있는 책상이 올바르게 배치됨.",
    "entities": "수하 1(모자 누락), 마츠다(파란색 모자 누락)의 의상이 일부 부족하나 신원과 역할은 인지 가능함. 무전기의 묘사는 적절함.",
    "hard_violations": [],
    "physics": "수하 1이 손으로 무전기를 자연스럽게 쥐고 있으며, 마츠다는 상체를 책상 위에 지지한 채 중력에 맞게 늘어져 있음."
   },
   {
    "label": "B",
    "direction": "수하 1이 화면 왼쪽을 바라보는데, 마츠다가 그의 시선 앞쪽(화면 좌측)에 위치하여 '등진' 방향을 어김.",
    "built_space": "고물상 내부로 오래된 모니터와 기기들이 보이나, 책상이 지시와 반대로 화면 좌측에 배치됨.",
    "entities": "수하 1(모자 누락), 마츠다(모자 착용)가 묘사됨. 무전기의 형태와 파지법은 정상적임.",
    "hard_violations": [
     "[gpt-high] 마츠다와 수리 책상을 수하의 등 뒤 중간 오른쪽 배경이 아니라 왼쪽 얼굴 방향에 배치하여 명시된 인물·공간 배치를 위반했다."
    ],
    "physics": "수하 1의 손이 무전기를 잡고 있고, 마츠다 역시 상체가 책상에 기대어 지지받고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "화면 우측 배경에 마츠다를 배치하여 '등진 채'라는 핵심 구도와 위치 지정(우측 배경)을 잘 준수했으나, 두 인물의 모자가 모두 누락된 점이 아쉬움."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "마츠다를 수하의 시선 앞쪽인 좌측 배경에 배치하여 '등진 채'라는 프롬프트의 지시와 화면 우측 배치라는 프레이밍을 크게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수하 1은 화면 왼쪽을 향해 시선을 두고 무전기를 입가에 대고 있으며, 마츠다는 그의 등 뒤인 화면 우측 방향에 위치함.",
        "built_space": "고물상 내부로 배경에 가전제품과 선반이 보이며, 우측 배경에 쓰러진 마츠다를 받치고 있는 책상이 올바르게 배치됨.",
        "entities": "수하 1(모자 누락), 마츠다(파란색 모자 누락)의 의상이 일부 부족하나 신원과 역할은 인지 가능함. 무전기의 묘사는 적절함.",
        "hard_violations": [],
        "physics": "수하 1이 손으로 무전기를 자연스럽게 쥐고 있으며, 마츠다는 상체를 책상 위에 지지한 채 중력에 맞게 늘어져 있음."
       },
       {
        "label": "B",
        "direction": "수하 1이 화면 왼쪽을 바라보는데, 마츠다가 그의 시선 앞쪽(화면 좌측)에 위치하여 '등진' 방향을 어김.",
        "built_space": "고물상 내부로 오래된 모니터와 기기들이 보이나, 책상이 지시와 반대로 화면 좌측에 배치됨.",
        "entities": "수하 1(모자 누락), 마츠다(모자 착용)가 묘사됨. 무전기의 형태와 파지법은 정상적임.",
        "hard_violations": [],
        "physics": "수하 1의 손이 무전기를 잡고 있고, 마츠다 역시 상체가 책상에 기대어 지지받고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "화면 우측 배경에 마츠다를 배치하여 '등진 채'라는 핵심 구도와 위치 지정(우측 배경)을 잘 준수했으나, 두 인물의 모자가 모두 누락된 점이 아쉬움."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "마츠다를 수하의 시선 앞쪽인 좌측 배경에 배치하여 '등진 채'라는 프롬프트의 지시와 화면 우측 배치라는 프레이밍을 크게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수하 1은 화면 왼쪽을 향해 시선을 두고 무전기를 입가에 대고 있으며, 마츠다는 그의 등 뒤인 화면 우측 방향에 위치함.",
        "built_space": "고물상 내부로 배경에 가전제품과 선반이 보이며, 우측 배경에 쓰러진 마츠다를 받치고 있는 책상이 올바르게 배치됨.",
        "entities": "수하 1(모자 누락), 마츠다(파란색 모자 누락)의 의상이 일부 부족하나 신원과 역할은 인지 가능함. 무전기의 묘사는 적절함.",
        "hard_violations": [],
        "physics": "수하 1이 손으로 무전기를 자연스럽게 쥐고 있으며, 마츠다는 상체를 책상 위에 지지한 채 중력에 맞게 늘어져 있음."
       },
       {
        "label": "B",
        "direction": "수하 1이 화면 왼쪽을 바라보는데, 마츠다가 그의 시선 앞쪽(화면 좌측)에 위치하여 '등진' 방향을 어김.",
        "built_space": "고물상 내부로 오래된 모니터와 기기들이 보이나, 책상이 지시와 반대로 화면 좌측에 배치됨.",
        "entities": "수하 1(모자 누락), 마츠다(모자 착용)가 묘사됨. 무전기의 형태와 파지법은 정상적임.",
        "hard_violations": [],
        "physics": "수하 1의 손이 무전기를 잡고 있고, 마츠다 역시 상체가 책상에 기대어 지지받고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "입가의 무전기와 차가운 옆얼굴은 구현했지만, 마츠다와 책상이 수하의 등 뒤 오른쪽이 아니라 얼굴이 향한 왼쪽에 놓여 핵심 배치를 어겼다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업과 수하의 등 뒤 오른쪽에 쓰러진 마츠다를 둔 구도가 정확하지만, 두 인물의 모자 누락과 책상에 충분히 엎드리지 않은 마츠다의 상체는 아쉽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수하는 화면 왼쪽을 바라보며 무전기를 입 바로 아래에 댄다. 안테나는 위로 향하고 손이 본체를 잡고 있다. 마츠다는 왼쪽 배경에 있어 수하의 시선 방향 쪽에 놓이며, 요구된 등 뒤 관계가 성립하지 않는다. 마츠다의 얼굴은 책상 위에서 카메라 쪽으로 돌아가 있다.",
        "built_space": "왼쪽에 유리 창호, 뒤쪽에 폐가전 선반, 왼쪽 중경에 수리 책상 한 개와 그 위의 대형 브라운관 한 대가 보인다. 낡은 가전과 목재 표면은 장소의 성격에 맞지만, 책상과 마츠다가 화면 중간 오른쪽 배경이라는 지정 위치를 벗어났다. 거울이나 반사에 따른 문제는 보이지 않는다.",
        "entities": "성인 동아시아계 남성 수하 한 명과 노년 동아시아계 남성 마츠다 한 명이 보인다. 국적은 외모만으로 확인할 수 없다. 수하의 짧은 검은 머리, 회색 작업복, 붉은 완장과 무전기는 맞지만 참조의 모자가 없다. 마츠다의 푸른 모자, 베이지색 셔츠, 어두운 작업 조끼는 참조와 가깝다. 얼굴의 세부 동일성과 상처는 흐린 배경 때문에 확정하기 어렵다.",
        "hard_violations": [
         "마츠다와 수리 책상을 수하의 등 뒤 중간 오른쪽 배경이 아니라 왼쪽 얼굴 방향에 배치하여 명시된 인물·공간 배치를 위반했다."
        ],
        "physics": "무전기는 수하의 손가락과 손바닥으로 지지되며 손목과 소매도 자연스럽게 이어진다. 마츠다의 머리는 책상에 닿고 어깨와 상체 일부가 책상 가장자리에 걸쳐 있으며, 보이는 팔은 아래로 늘어진다. 공중에 떠 있거나 스스로 들어 올린 신체 부위는 보이지 않는다. 하체의 지지 상태는 프레임 밖이다."
       },
       {
        "label": "B",
        "direction": "수하는 왼쪽 화면 밖을 차갑게 응시하며 무전기를 입술 바로 옆에 댄다. 마츠다는 수하의 등 뒤 오른쪽에 있어 등을 돌린 관계가 명확하다. 마츠다는 눈을 감고 얼굴 옆면을 책상에 댄다. 무전기의 전면 격자와 버튼은 카메라 쪽으로 비스듬히 드러나며, 입에 정면으로 향하기보다는 옆으로 돌아가 있다.",
        "built_space": "수하는 왼쪽 전경의 클로즈업이고, 수리 책상 한 개와 마츠다는 오른쪽 배경에 있다. 책상 상판은 비스듬히 보이며 오른쪽 뒤에 브라운관 한 대가 놓였다. 중앙 유리 출입구 한 조, 왼쪽 냉장고와 적층 가전, 양쪽 선반이 이전 장면의 공간을 알아볼 수 있게 유지한다. 낮의 출입구 빛도 이어지며, 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "성인 동아시아계 남성 수하 한 명과 노년 동아시아계 남성 마츠다 한 명만 보인다. 국적은 시각적으로 확인할 수 없다. 수하의 짧은 검은 머리, 회색 작업복과 일부 보이는 붉은 완장은 맞지만 참조의 모자가 없다. 마츠다의 나이 든 얼굴, 회색 머리, 베이지색 셔츠와 작업 조끼는 대체로 맞지만 푸른 모자가 빠졌다. 무전기, 수리 책상, 텔레비전과 수리 공구가 보인다.",
        "hard_violations": [],
        "physics": "수하의 손이 무전기를 확실히 감싸 쥐며 손목과 팔이 연결된다. 마츠다의 머리는 책상 상판에 지지되고 팔은 아래로 처져 있다. 다만 몸통 대부분은 책상 밖에서 앞으로 굽어 있어, 상체가 상판에 푹 엎어진 정해진 자세보다는 머리를 책상에 기댄 모습에 가깝다. 하체와 좌석의 접촉은 가려져 확인할 수 없지만, 보이는 부분에 명백한 부유나 능동적으로 들어 올린 팔다리는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "입가의 무전기와 차가운 옆얼굴은 구현했지만, 마츠다와 책상이 수하의 등 뒤 오른쪽이 아니라 얼굴이 향한 왼쪽에 놓여 핵심 배치를 어겼다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업과 수하의 등 뒤 오른쪽에 쓰러진 마츠다를 둔 구도가 정확하지만, 두 인물의 모자 누락과 책상에 충분히 엎드리지 않은 마츠다의 상체는 아쉽다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "수하는 화면 왼쪽을 바라보며 무전기를 입 바로 아래에 댄다. 안테나는 위로 향하고 손이 본체를 잡고 있다. 마츠다는 왼쪽 배경에 있어 수하의 시선 방향 쪽에 놓이며, 요구된 등 뒤 관계가 성립하지 않는다. 마츠다의 얼굴은 책상 위에서 카메라 쪽으로 돌아가 있다.",
        "built_space": "왼쪽에 유리 창호, 뒤쪽에 폐가전 선반, 왼쪽 중경에 수리 책상 한 개와 그 위의 대형 브라운관 한 대가 보인다. 낡은 가전과 목재 표면은 장소의 성격에 맞지만, 책상과 마츠다가 화면 중간 오른쪽 배경이라는 지정 위치를 벗어났다. 거울이나 반사에 따른 문제는 보이지 않는다.",
        "entities": "성인 동아시아계 남성 수하 한 명과 노년 동아시아계 남성 마츠다 한 명이 보인다. 국적은 외모만으로 확인할 수 없다. 수하의 짧은 검은 머리, 회색 작업복, 붉은 완장과 무전기는 맞지만 참조의 모자가 없다. 마츠다의 푸른 모자, 베이지색 셔츠, 어두운 작업 조끼는 참조와 가깝다. 얼굴의 세부 동일성과 상처는 흐린 배경 때문에 확정하기 어렵다.",
        "hard_violations": [
         "마츠다와 수리 책상을 수하의 등 뒤 중간 오른쪽 배경이 아니라 왼쪽 얼굴 방향에 배치하여 명시된 인물·공간 배치를 위반했다."
        ],
        "physics": "무전기는 수하의 손가락과 손바닥으로 지지되며 손목과 소매도 자연스럽게 이어진다. 마츠다의 머리는 책상에 닿고 어깨와 상체 일부가 책상 가장자리에 걸쳐 있으며, 보이는 팔은 아래로 늘어진다. 공중에 떠 있거나 스스로 들어 올린 신체 부위는 보이지 않는다. 하체의 지지 상태는 프레임 밖이다."
       },
       {
        "label": "A",
        "direction": "수하는 왼쪽 화면 밖을 차갑게 응시하며 무전기를 입술 바로 옆에 댄다. 마츠다는 수하의 등 뒤 오른쪽에 있어 등을 돌린 관계가 명확하다. 마츠다는 눈을 감고 얼굴 옆면을 책상에 댄다. 무전기의 전면 격자와 버튼은 카메라 쪽으로 비스듬히 드러나며, 입에 정면으로 향하기보다는 옆으로 돌아가 있다.",
        "built_space": "수하는 왼쪽 전경의 클로즈업이고, 수리 책상 한 개와 마츠다는 오른쪽 배경에 있다. 책상 상판은 비스듬히 보이며 오른쪽 뒤에 브라운관 한 대가 놓였다. 중앙 유리 출입구 한 조, 왼쪽 냉장고와 적층 가전, 양쪽 선반이 이전 장면의 공간을 알아볼 수 있게 유지한다. 낮의 출입구 빛도 이어지며, 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "성인 동아시아계 남성 수하 한 명과 노년 동아시아계 남성 마츠다 한 명만 보인다. 국적은 시각적으로 확인할 수 없다. 수하의 짧은 검은 머리, 회색 작업복과 일부 보이는 붉은 완장은 맞지만 참조의 모자가 없다. 마츠다의 나이 든 얼굴, 회색 머리, 베이지색 셔츠와 작업 조끼는 대체로 맞지만 푸른 모자가 빠졌다. 무전기, 수리 책상, 텔레비전과 수리 공구가 보인다.",
        "hard_violations": [],
        "physics": "수하의 손이 무전기를 확실히 감싸 쥐며 손목과 팔이 연결된다. 마츠다의 머리는 책상 상판에 지지되고 팔은 아래로 처져 있다. 다만 몸통 대부분은 책상 밖에서 앞으로 굽어 있어, 상체가 상판에 푹 엎어진 정해진 자세보다는 머리를 책상에 기댄 모습에 가깝다. 하체와 좌석의 접촉은 가려져 확인할 수 없지만, 보이는 부분에 명백한 부유나 능동적으로 들어 올린 팔다리는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.143
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.893
   },
   "violations": {
    "B": [
     "[gpt-high] 마츠다와 수리 책상을 수하의 등 뒤 중간 오른쪽 배경이 아니라 왼쪽 얼굴 방향에 배치하여 명시된 인물·공간 배치를 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 893
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "화면 우측 배경에 마츠다를 배치하여 '등진 채'라는 핵심 구도와 위치 지정(우측 배경)을 잘 준수했으나, 두 인물의 모자가 모두 누락된 점이 아쉬움."
   },
   {
    "label": "B",
    "score": 893,
    "verdict_ko": "마츠다를 수하의 시선 앞쪽인 좌측 배경에 배치하여 '등진 채'라는 프롬프트의 지시와 화면 우측 배치라는 프레이밍을 크게 위반함.  ★위반: [gpt-high] 마츠다와 수리 책상을 수하의 등 뒤 중간 오른쪽 배경이 아니라 왼쪽 얼굴 방향에 배치하여 명시된 인물·공간 배치를 위반했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S17sh1_sel.png",
    "asset_id": "d82cc84e-eccf-48ee-95b2-8dbb5a4226df",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:798636>",
    "asset_id": "840c8473-e2ef-41b8-ac3f-988134fc14e7",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1356643>",
    "asset_id": "b9c02101-a6c1-4859-84a5-c3269b6eef27",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-de61-7e1d-99ba-6a4c6eda68f5",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S17sh1"
  },
  "staged_characters_added": [
   "C15"
  ]
 },
 "S17sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:39:18.392953+00:00",
  "fingerprint": "b00c0d0521c8328be9223a50ff09391e0e2fe9d7eb61f33cc1303257fdb6d247",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S17sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S17sh7_sel.png",
  "source_sha256": "38c00d0eb11e91f14407890cb39145240959c0efe518b0cc050caa8965269426",
  "file": "S17sh7_cine.png",
  "staged_sha256": "088e0126236f68ac1be9cd93ce04f9f2ae996a03ce55e38dc6f8eca9a112d199",
  "latency_ms": 14294
 },
 "S18sh1::signage": {
  "fp": "9ffc554368356fea",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::a44191158dcd7174": {
  "subjects": [],
  "subject_text": "박철진의 민병대 사무실\n책상과 의자, 스피커폰이 놓인 사무실. 책상 맞은편에 빈 공간이 있고 그 뒤로 벽면이 드러나 있다.",
  "identity": "canonical",
  "scope_id": "L171",
  "scope_role": "location_interior",
  "scope_sha": "ecbb2eee6c55995e"
 },
 "S18sh1::bgfirst_bg": {
  "input_fingerprint": "4944a65c5e16b663",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1__bgfirst_bg.png",
  "asset_id": "5a3734d6-9057-4f50-afa3-4307c7c976a2",
  "input_asset_ids": [
   "86ca7f7a-7c3a-4769-b837-fca22d49aac1",
   "64c1729a-dc88-4c0d-aef7-ac2fc3c15d65"
  ]
 },
 "S18sh1": {
  "input_fingerprint": "492a4f92624639e0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The speakerphone is in use in the militia office. 박철진: He has a knife in hand, playing with it during the speakerphone conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The speakerphone is in use in the militia office. 박철진: He has a knife in hand, playing with it during the speakerphone conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The speakerphone is in use in the militia office. 박철진: He has a knife in hand, playing with it during the speakerphone conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1__bgfirst_bg.png",
     "asset_id": "5a3734d6-9057-4f50-afa3-4307c7c976a2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S18sh1.png",
     "asset_id": "86ca7f7a-7c3a-4769-b837-fca22d49aac1",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L171B01.png",
     "asset_id": "64c1729a-dc88-4c0d-aef7-ac2fc3c15d65",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물의 시선과 고개가 책상 중앙에 놓인 스피커폰을 향해 자연스럽게 아래로 고정되어 있음.",
    "built_space": "참조 이미지의 사무실 구조와 일치함. 창문, 파일 캐비닛, 화이트보드가 제자리에 있고 인물은 책상 뒤 의자에 착석한 상태임.",
    "entities": "박철진의 얼굴, 헤어스타일, 어두운 셔츠는 일치하나 우측 팔의 붉은 완장이 누락됨. 책상 위 스피커폰과 손에 든 칼은 명확히 식별됨.",
    "hard_violations": [
     "[gpt-high] 사용 중인 스피커폰의 버튼·표시부가 사용자 반대쪽인 카메라를 향해 있어, 기능면을 사용자 자신의 눈 쪽으로 두라는 명시적 소품 방향 지시를 위반한다."
    ],
    "physics": "의자에 앉은 상태로 왼팔 하박을 책상에 기대어 체중을 싣고, 오른손으로 칼을 들고 있는 안정적이고 현실적인 자세를 유지함."
   },
   {
    "label": "B",
    "direction": "인물의 시선이 책상 위의 스피커폰을 향해 아래로 향하고 있음.",
    "built_space": "사무실 창문과 책상이 배치되어 있으며, 인물이 책상 위로 상체를 깊게 기울인 공간감을 보여줌.",
    "entities": "박철진의 이목구비와 우측 팔의 붉은 완장 등 복장 요소가 잘 반영되었으며, 스피커폰이 책상 위에 있음.",
    "hard_violations": [
     "[gemini-pro] 칼을 쥐고 있는 양손의 손가락 개수와 관절이 기형적으로 융합된 물리적 불가능 해부학 구조",
     "[gemini-pro] 손에 든 칼이 양방향으로 날이 뻗어 있는 불가능한 형태의 사물로 렌더링됨",
     "[gpt-high] 사용 중인 스피커폰의 경사진 조작면이 박철진 반대쪽인 카메라를 향해 있어, 소품의 기능면을 사용자 쪽으로 두라는 명시적 방향 지시를 위반한다."
    ],
    "physics": "양팔을 책상 면에 기대어 상체를 지지하고 있으나, 손과 칼이 닿는 부위의 형태와 물리적 상호작용이 완전히 붕괴됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "우측 완장 누락과 책상 모서리의 사선 구도 미적용 등 일부 아쉬움이 있으나, 치명적인 물리적·해부학적 오류가 없어 지시문을 안정적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "복장 디테일과 책상의 사선 프레이밍은 훌륭하나, 칼을 쥔 양손의 손가락이 융합되고 칼날 형태가 붕괴되는 심각한 해부학적/물리적 오류(Hard Violation)로 인해 실격됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선과 고개가 책상 중앙에 놓인 스피커폰을 향해 자연스럽게 아래로 고정되어 있음.",
        "built_space": "참조 이미지의 사무실 구조와 일치함. 창문, 파일 캐비닛, 화이트보드가 제자리에 있고 인물은 책상 뒤 의자에 착석한 상태임.",
        "entities": "박철진의 얼굴, 헤어스타일, 어두운 셔츠는 일치하나 우측 팔의 붉은 완장이 누락됨. 책상 위 스피커폰과 손에 든 칼은 명확히 식별됨.",
        "hard_violations": [],
        "physics": "의자에 앉은 상태로 왼팔 하박을 책상에 기대어 체중을 싣고, 오른손으로 칼을 들고 있는 안정적이고 현실적인 자세를 유지함."
       },
       {
        "label": "B",
        "direction": "인물의 시선이 책상 위의 스피커폰을 향해 아래로 향하고 있음.",
        "built_space": "사무실 창문과 책상이 배치되어 있으며, 인물이 책상 위로 상체를 깊게 기울인 공간감을 보여줌.",
        "entities": "박철진의 이목구비와 우측 팔의 붉은 완장 등 복장 요소가 잘 반영되었으며, 스피커폰이 책상 위에 있음.",
        "hard_violations": [
         "칼을 쥐고 있는 양손의 손가락 개수와 관절이 기형적으로 융합된 물리적 불가능 해부학 구조",
         "손에 든 칼이 양방향으로 날이 뻗어 있는 불가능한 형태의 사물로 렌더링됨"
        ],
        "physics": "양팔을 책상 면에 기대어 상체를 지지하고 있으나, 손과 칼이 닿는 부위의 형태와 물리적 상호작용이 완전히 붕괴됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "우측 완장 누락과 책상 모서리의 사선 구도 미적용 등 일부 아쉬움이 있으나, 치명적인 물리적·해부학적 오류가 없어 지시문을 안정적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "복장 디테일과 책상의 사선 프레이밍은 훌륭하나, 칼을 쥔 양손의 손가락이 융합되고 칼날 형태가 붕괴되는 심각한 해부학적/물리적 오류(Hard Violation)로 인해 실격됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선과 고개가 책상 중앙에 놓인 스피커폰을 향해 자연스럽게 아래로 고정되어 있음.",
        "built_space": "참조 이미지의 사무실 구조와 일치함. 창문, 파일 캐비닛, 화이트보드가 제자리에 있고 인물은 책상 뒤 의자에 착석한 상태임.",
        "entities": "박철진의 얼굴, 헤어스타일, 어두운 셔츠는 일치하나 우측 팔의 붉은 완장이 누락됨. 책상 위 스피커폰과 손에 든 칼은 명확히 식별됨.",
        "hard_violations": [],
        "physics": "의자에 앉은 상태로 왼팔 하박을 책상에 기대어 체중을 싣고, 오른손으로 칼을 들고 있는 안정적이고 현실적인 자세를 유지함."
       },
       {
        "label": "B",
        "direction": "인물의 시선이 책상 위의 스피커폰을 향해 아래로 향하고 있음.",
        "built_space": "사무실 창문과 책상이 배치되어 있으며, 인물이 책상 위로 상체를 깊게 기울인 공간감을 보여줌.",
        "entities": "박철진의 이목구비와 우측 팔의 붉은 완장 등 복장 요소가 잘 반영되었으며, 스피커폰이 책상 위에 있음.",
        "hard_violations": [
         "칼을 쥐고 있는 양손의 손가락 개수와 관절이 기형적으로 융합된 물리적 불가능 해부학 구조",
         "손에 든 칼이 양방향으로 날이 뻗어 있는 불가능한 형태의 사물로 렌더링됨"
        ],
        "physics": "양팔을 책상 면에 기대어 상체를 지지하고 있으나, 손과 칼이 닿는 부위의 형태와 물리적 상호작용이 완전히 붕괴됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "상체를 기울인 자세와 붉은 완장은 맞지만, 시선이 스피커폰이 아닌 왼쪽 전방을 향하고 전화기 조작면도 사용자 반대쪽으로 돌아가 있다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "스피커폰을 내려다보며 능글맞게 웃는 상체와 칼을 만지작거리는 순간은 더 정확하지만, 전화기 조작면이 카메라를 향한 공통 오류가 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "몸통과 얼굴은 책상 위 스피커폰 쪽으로 기울어 있지만, 눈은 전화기보다 높은 화면 왼쪽 전방을 본다. 아래쪽 전화기에 주의를 둔 시선은 아니다. 양손 근처의 칼끝은 화면 왼쪽 위를 향하며 특정 대상을 겨누지는 않는다. 전화기의 경사진 버튼 면은 박철진이 아니라 카메라 쪽을 향한다.",
        "built_space": "목재 책상 하나와 그 뒤 검은 사무용 의자 하나가 보인다. 박철진은 의자 앞에서 책상에 기대고 있다. 왼쪽 창, 금속 선반, 낮은 수납장, 뒤쪽 세로형 서류함과 흰색·회색 벽이 참고 공간과 대체로 연결된다. 책상 위에는 스피커폰 하나, 펜통 하나, 서류 묶음들이 있다. 책상 앞 모서리는 하단을 사선으로 가로지르지만, 앞판까지 크게 보여 전경 비중이 다소 크다. 반사상이나 중복된 주요 설비는 보이지 않는다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명으로, 참고 인물의 얼굴과 체격에 대체로 부합한다. 어두운 군복, 붉은 오른팔 완장과 허리 장구가 보인다. 능글맞은 미소, 손에 든 칼 하나, 책상 위 회의용 스피커폰 하나가 확인된다. 추가 인물은 없으며 뚜렷하게 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [
         "사용 중인 스피커폰의 경사진 조작면이 박철진 반대쪽인 카메라를 향해 있어, 소품의 기능면을 사용자 쪽으로 두라는 명시적 방향 지시를 위반한다."
        ],
        "physics": "앞으로 숙인 몸은 책상에 닿은 팔과 손으로 지지된다. 하체의 접지는 프레임 밖이지만 상체가 공중에 떠 있는 모습은 아니다. 칼 손잡이는 손에 잡혀 있으며, 전화기와 펜통 및 서류는 책상에 놓여 있다. 칼을 낮게 들고 만지는 자세는 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "얼굴과 시선이 아래쪽 중앙의 스피커폰을 향하고, 몸통도 그쪽으로 기울어 있다. 미소와 아래로 향한 주의가 통화 순간에 잘 연결된다. 오른손의 칼끝은 화면 왼쪽 위를 향하며 공격 대상을 겨누지 않는다. 전화기의 버튼과 표시부는 박철진 반대편인 카메라 쪽으로 기울어 있다.",
        "built_space": "목재 책상 하나와 검은 사무용 의자 하나가 있고, 박철진은 등받이를 뒤에 두고 앉아 있다. 왼쪽 창과 금속 선반, 낮은 수납장, 높은 서류함 하나, 오른쪽 게시판 하나, 뒤 벽의 작은 걸이 물체가 참고 공간의 배치와 잘 맞는다. 책상에는 스피커폰 하나, 펜통 하나, 좌우 서류 묶음이 있다. 상체 중심의 미디엄 숏이며 책상 앞 모서리가 하단을 비스듬히 지난다. 불가능한 반사나 주요 설비의 중복은 보이지 않는다.",
        "entities": "짧은 검은 머리와 중년 얼굴을 가진 동아시아계 남성 한 명으로 참고 인물과 대체로 일치한다. 어두운 군복은 맞지만 보이는 오른쪽 소매에서 참고의 붉은 완장은 확인되지 않는다. 손에 든 칼 하나와 통화 표시등이 켜진 회의용 스피커폰 하나가 있다. 미소는 요청한 능글맞은 표정에 가깝다. 추가 인물은 없고 게시판 문서는 읽기 어렵게 처리되어 있다.",
        "hard_violations": [
         "사용 중인 스피커폰의 버튼·표시부가 사용자 반대쪽인 카메라를 향해 있어, 기능면을 사용자 자신의 눈 쪽으로 두라는 명시적 소품 방향 지시를 위반한다."
        ],
        "physics": "몸은 사무용 의자에 앉아 지지되며 양팔은 책상에 기대고 있다. 오른손이 칼 손잡이를 확실히 잡아 들어 올리고 있어 칼의 지지가 명확하다. 왼팔과 손은 책상에 자연스럽게 놓여 있다. 전화기와 서류, 펜통 모두 책상 표면에 지지되며, 지지 없이 떠 있는 물체나 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "상체를 기울인 자세와 붉은 완장은 맞지만, 시선이 스피커폰이 아닌 왼쪽 전방을 향하고 전화기 조작면도 사용자 반대쪽으로 돌아가 있다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "스피커폰을 내려다보며 능글맞게 웃는 상체와 칼을 만지작거리는 순간은 더 정확하지만, 전화기 조작면이 카메라를 향한 공통 오류가 남는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "몸통과 얼굴은 책상 위 스피커폰 쪽으로 기울어 있지만, 눈은 전화기보다 높은 화면 왼쪽 전방을 본다. 아래쪽 전화기에 주의를 둔 시선은 아니다. 양손 근처의 칼끝은 화면 왼쪽 위를 향하며 특정 대상을 겨누지는 않는다. 전화기의 경사진 버튼 면은 박철진이 아니라 카메라 쪽을 향한다.",
        "built_space": "목재 책상 하나와 그 뒤 검은 사무용 의자 하나가 보인다. 박철진은 의자 앞에서 책상에 기대고 있다. 왼쪽 창, 금속 선반, 낮은 수납장, 뒤쪽 세로형 서류함과 흰색·회색 벽이 참고 공간과 대체로 연결된다. 책상 위에는 스피커폰 하나, 펜통 하나, 서류 묶음들이 있다. 책상 앞 모서리는 하단을 사선으로 가로지르지만, 앞판까지 크게 보여 전경 비중이 다소 크다. 반사상이나 중복된 주요 설비는 보이지 않는다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명으로, 참고 인물의 얼굴과 체격에 대체로 부합한다. 어두운 군복, 붉은 오른팔 완장과 허리 장구가 보인다. 능글맞은 미소, 손에 든 칼 하나, 책상 위 회의용 스피커폰 하나가 확인된다. 추가 인물은 없으며 뚜렷하게 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [
         "사용 중인 스피커폰의 경사진 조작면이 박철진 반대쪽인 카메라를 향해 있어, 소품의 기능면을 사용자 쪽으로 두라는 명시적 방향 지시를 위반한다."
        ],
        "physics": "앞으로 숙인 몸은 책상에 닿은 팔과 손으로 지지된다. 하체의 접지는 프레임 밖이지만 상체가 공중에 떠 있는 모습은 아니다. 칼 손잡이는 손에 잡혀 있으며, 전화기와 펜통 및 서류는 책상에 놓여 있다. 칼을 낮게 들고 만지는 자세는 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "얼굴과 시선이 아래쪽 중앙의 스피커폰을 향하고, 몸통도 그쪽으로 기울어 있다. 미소와 아래로 향한 주의가 통화 순간에 잘 연결된다. 오른손의 칼끝은 화면 왼쪽 위를 향하며 공격 대상을 겨누지 않는다. 전화기의 버튼과 표시부는 박철진 반대편인 카메라 쪽으로 기울어 있다.",
        "built_space": "목재 책상 하나와 검은 사무용 의자 하나가 있고, 박철진은 등받이를 뒤에 두고 앉아 있다. 왼쪽 창과 금속 선반, 낮은 수납장, 높은 서류함 하나, 오른쪽 게시판 하나, 뒤 벽의 작은 걸이 물체가 참고 공간의 배치와 잘 맞는다. 책상에는 스피커폰 하나, 펜통 하나, 좌우 서류 묶음이 있다. 상체 중심의 미디엄 숏이며 책상 앞 모서리가 하단을 비스듬히 지난다. 불가능한 반사나 주요 설비의 중복은 보이지 않는다.",
        "entities": "짧은 검은 머리와 중년 얼굴을 가진 동아시아계 남성 한 명으로 참고 인물과 대체로 일치한다. 어두운 군복은 맞지만 보이는 오른쪽 소매에서 참고의 붉은 완장은 확인되지 않는다. 손에 든 칼 하나와 통화 표시등이 켜진 회의용 스피커폰 하나가 있다. 미소는 요청한 능글맞은 표정에 가깝다. 추가 인물은 없고 게시판 문서는 읽기 어렵게 처리되어 있다.",
        "hard_violations": [
         "사용 중인 스피커폰의 버튼·표시부가 사용자 반대쪽인 카메라를 향해 있어, 기능면을 사용자 자신의 눈 쪽으로 두라는 명시적 소품 방향 지시를 위반한다."
        ],
        "physics": "몸은 사무용 의자에 앉아 지지되며 양팔은 책상에 기대고 있다. 오른손이 칼 손잡이를 확실히 잡아 들어 올리고 있어 칼의 지지가 명확하다. 왼팔과 손은 책상에 자연스럽게 놓여 있다. 전화기와 서류, 펜통 모두 책상 표면에 지지되며, 지지 없이 떠 있는 물체나 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.229
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.979
   },
   "violations": {
    "B": [
     "[gemini-pro] 칼을 쥐고 있는 양손의 손가락 개수와 관절이 기형적으로 융합된 물리적 불가능 해부학 구조",
     "[gemini-pro] 손에 든 칼이 양방향으로 날이 뻗어 있는 불가능한 형태의 사물로 렌더링됨",
     "[gpt-high] 사용 중인 스피커폰의 경사진 조작면이 박철진 반대쪽인 카메라를 향해 있어, 소품의 기능면을 사용자 쪽으로 두라는 명시적 방향 지시를 위반한다."
    ],
    "A": [
     "[gpt-high] 사용 중인 스피커폰의 버튼·표시부가 사용자 반대쪽인 카메라를 향해 있어, 기능면을 사용자 자신의 눈 쪽으로 두라는 명시적 소품 방향 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 979
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "우측 완장 누락과 책상 모서리의 사선 구도 미적용 등 일부 아쉬움이 있으나, 치명적인 물리적·해부학적 오류가 없어 지시문을 안정적으로 구현함.  ★위반: [gpt-high] 사용 중인 스피커폰의 버튼·표시부가 사용자 반대쪽인 카메라를 향해 있어, 기능면을 사용자 자신의 눈 쪽으로 두라는 명시적 소품 방향 지시를 위반한다."
   },
   {
    "label": "B",
    "score": 979,
    "verdict_ko": "복장 디테일과 책상의 사선 프레이밍은 훌륭하나, 칼을 쥔 양손의 손가락이 융합되고 칼날 형태가 붕괴되는 심각한 해부학적/물리적 오류(Hard Violation)로 인해 실격됨.  ★위반: [gemini-pro] 칼을 쥐고 있는 양손의 손가락 개수와 관절이 기형적으로 융합된 물리적 불가능 해부학 구조 / [gemini-pro] 손에 든 칼이 양방향으로 날이 뻗어 있는 불가능한 형태의 사물로 렌더링됨 / [gpt-high] 사용 중인 스피커폰의 경사진 조작면이 박철진 반대쪽인 카메라를 향해 있어, 소품의 기능면을 사용자 쪽으로 두라는 명시적 방향 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L171B01.png",
    "asset_id": "64c1729a-dc88-4c0d-aef7-ac2fc3c15d65",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-e00d-7844-993e-c6e4f6027ecf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1__bgfirst_bg.png",
   "bg_asset_id": "5a3734d6-9057-4f50-afa3-4307c7c976a2",
   "bg_record_key": "S18sh1::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S18sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:40:38.894708+00:00",
  "fingerprint": "47df47e677e69f51dee62b23cbf7fbeecd89456487906e084ca9f81821375719",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S18sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S18sh1_sel.png",
  "source_sha256": "6c6f3194f2be80e1bc9326c6a4a7f75bec7db09a1fb04dffa0c5bd4ce0b3f00f",
  "file": "S18sh1_cine.png",
  "staged_sha256": "689264d1b176447a7835aef2b8cbf69a92eaac5880a4895c121e0c7fd3f96827",
  "latency_ms": 11586
 },
 "S18sh5::signage": {
  "fp": "fa92ba9eec020c5a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S18sh5": {
  "input_fingerprint": "27dab96b9ec59248",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 벽에 박힌 단검 바로 옆에서 몸을 움츠린 채 두 눈을 질끈 감은 윤성찬의 붉어진 얼굴.\n\nLOCATION (lock): Inside the militia commander's office, directly beside the wall struck by a thrown dagger, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Wall section with the embedded dagger immediately beside and behind 윤성찬's head in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Wall behind 윤성찬 (The thrown dagger is embedded beside him) — An adjacent section is visible behind the right side of his head; used as Make the near miss spatially legible without placing the weapon in front of his face; Embedded dagger (Lodged in the wall after being thrown) — The exposed handle projects from the wall and is seen obliquely beside, not overlapping, his head; used as Provide a small but readable cause for his recoil.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral office illumination with restrained contrast, preserving the explicitly flushed skin without turning it into a colored lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The thrown knife remains lodged in the office wall; the speakerphone remains in place. 윤성찬: He remains standing with his handkerchief, his face flushed from the fright.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 벽에 박힌 단검 바로 옆에서 몸을 움츠린 채 두 눈을 질끈 감은 윤성찬의 붉어진 얼굴.\n\nLOCATION (lock): Inside the militia commander's office, directly beside the wall struck by a thrown dagger, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Wall section with the embedded dagger immediately beside and behind 윤성찬's head in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Wall behind 윤성찬 (The thrown dagger is embedded beside him) — An adjacent section is visible behind the right side of his head; used as Make the near miss spatially legible without placing the weapon in front of his face; Embedded dagger (Lodged in the wall after being thrown) — The exposed handle projects from the wall and is seen obliquely beside, not overlapping, his head; used as Provide a small but readable cause for his recoil.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral office illumination with restrained contrast, preserving the explicitly flushed skin without turning it into a colored lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The thrown knife remains lodged in the office wall; the speakerphone remains in place. 윤성찬: He remains standing with his handkerchief, his face flushed from the fright.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 벽에 박힌 단검 바로 옆에서 몸을 움츠린 채 두 눈을 질끈 감은 윤성찬의 붉어진 얼굴.\n\nLOCATION (lock): Inside the militia commander's office, directly beside the wall struck by a thrown dagger, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Wall section with the embedded dagger immediately beside and behind 윤성찬's head in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Wall behind 윤성찬 (The thrown dagger is embedded beside him) — An adjacent section is visible behind the right side of his head; used as Make the near miss spatially legible without placing the weapon in front of his face; Embedded dagger (Lodged in the wall after being thrown) — The exposed handle projects from the wall and is seen obliquely beside, not overlapping, his head; used as Provide a small but readable cause for his recoil.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral office illumination with restrained contrast, preserving the explicitly flushed skin without turning it into a colored lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The thrown knife remains lodged in the office wall; the speakerphone remains in place. 윤성찬: He remains standing with his handkerchief, his face flushed from the fright.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "윤성찬은 두 눈을 질끈 감고 오른쪽 벽에 박힌 단검을 피해 몸을 움츠리고 있다.",
    "built_space": "사무실 내부. 화면 오른쪽에 단검이 박힌 벽면이 있고, 배경 왼쪽에는 스피커폰이 놓인 책상이 위치한다.",
    "entities": "윤성찬(고령의 남성, 홍조 띤 얼굴, 감은 눈, 네이비 정장), 손에 든 손수건, 벽에 박힌 단검이 모두 확인된다.",
    "hard_violations": [],
    "physics": "단검은 오른쪽 벽에 단단히 박혀 지탱되고 있으며, 남성의 손은 자연스럽게 손수건을 쥐고 있다."
   },
   {
    "label": "B",
    "direction": "윤성찬은 두 눈을 감고 손수건으로 코와 입을 가린 채 움츠리고 있다.",
    "built_space": "사무실 내부. 왼쪽에 콘크리트 기둥이 있고 배경 오른쪽에 책상과 게시판이 보인다.",
    "entities": "윤성찬(고령의 남성, 붉어진 얼굴, 감은 눈, 네이비 정장), 손수건, 단검이 확인된다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 묘사 (단검이 어떤 표면에도 박히지 않고 허공에 떠 있음)",
     "[gemini-pro] 물리적으로 불가능한 신체 구조 (손수건을 쥔 손가락이 비정상적으로 많고 뭉개짐)"
    ],
    "physics": "단검이 지지대 없이 허공에 떠 있으며, 손가락의 관절과 형태가 비정상적으로 엉켜 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트의 클로즈업 구도, 붉어진 얼굴로 움츠러든 표정, 벽에 박힌 단검의 위치를 정확히 구현했으나, 단검의 디자인이 레퍼런스와 다릅니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "단검이 허공에 떠 있고 손가락 구조가 왜곡되는 등 심각한 물리적 오류(Hard Violation)가 발생하여 지시를 충족하지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬은 두 눈을 질끈 감고 오른쪽 벽에 박힌 단검을 피해 몸을 움츠리고 있다.",
        "built_space": "사무실 내부. 화면 오른쪽에 단검이 박힌 벽면이 있고, 배경 왼쪽에는 스피커폰이 놓인 책상이 위치한다.",
        "entities": "윤성찬(고령의 남성, 홍조 띤 얼굴, 감은 눈, 네이비 정장), 손에 든 손수건, 벽에 박힌 단검이 모두 확인된다.",
        "hard_violations": [],
        "physics": "단검은 오른쪽 벽에 단단히 박혀 지탱되고 있으며, 남성의 손은 자연스럽게 손수건을 쥐고 있다."
       },
       {
        "label": "B",
        "direction": "윤성찬은 두 눈을 감고 손수건으로 코와 입을 가린 채 움츠리고 있다.",
        "built_space": "사무실 내부. 왼쪽에 콘크리트 기둥이 있고 배경 오른쪽에 책상과 게시판이 보인다.",
        "entities": "윤성찬(고령의 남성, 붉어진 얼굴, 감은 눈, 네이비 정장), 손수건, 단검이 확인된다.",
        "hard_violations": [
         "물리적으로 불가능한 묘사 (단검이 어떤 표면에도 박히지 않고 허공에 떠 있음)",
         "물리적으로 불가능한 신체 구조 (손수건을 쥔 손가락이 비정상적으로 많고 뭉개짐)"
        ],
        "physics": "단검이 지지대 없이 허공에 떠 있으며, 손가락의 관절과 형태가 비정상적으로 엉켜 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트의 클로즈업 구도, 붉어진 얼굴로 움츠러든 표정, 벽에 박힌 단검의 위치를 정확히 구현했으나, 단검의 디자인이 레퍼런스와 다릅니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "단검이 허공에 떠 있고 손가락 구조가 왜곡되는 등 심각한 물리적 오류(Hard Violation)가 발생하여 지시를 충족하지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬은 두 눈을 질끈 감고 오른쪽 벽에 박힌 단검을 피해 몸을 움츠리고 있다.",
        "built_space": "사무실 내부. 화면 오른쪽에 단검이 박힌 벽면이 있고, 배경 왼쪽에는 스피커폰이 놓인 책상이 위치한다.",
        "entities": "윤성찬(고령의 남성, 홍조 띤 얼굴, 감은 눈, 네이비 정장), 손에 든 손수건, 벽에 박힌 단검이 모두 확인된다.",
        "hard_violations": [],
        "physics": "단검은 오른쪽 벽에 단단히 박혀 지탱되고 있으며, 남성의 손은 자연스럽게 손수건을 쥐고 있다."
       },
       {
        "label": "B",
        "direction": "윤성찬은 두 눈을 감고 손수건으로 코와 입을 가린 채 움츠리고 있다.",
        "built_space": "사무실 내부. 왼쪽에 콘크리트 기둥이 있고 배경 오른쪽에 책상과 게시판이 보인다.",
        "entities": "윤성찬(고령의 남성, 붉어진 얼굴, 감은 눈, 네이비 정장), 손수건, 단검이 확인된다.",
        "hard_violations": [
         "물리적으로 불가능한 묘사 (단검이 어떤 표면에도 박히지 않고 허공에 떠 있음)",
         "물리적으로 불가능한 신체 구조 (손수건을 쥔 손가락이 비정상적으로 많고 뭉개짐)"
        ],
        "physics": "단검이 지지대 없이 허공에 떠 있으며, 손가락의 관절과 형태가 비정상적으로 엉켜 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "단검을 머리와 겹치지 않게 배치하고 움츠림을 잘 표현했지만, 구도가 더 넓고 손수건이 얼굴 하반부를 가려 핵심인 붉어진 얼굴의 클로즈업이 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "질끈 감은 두 눈과 붉어진 얼굴을 온전히 보여주는 클로즈업이 더 충실하지만, 단검이 지나치게 두드러지고 손잡이가 머리카락과 일부 겹치며 참조의 안경이 빠졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남성은 고개를 약간 숙이고 두 눈을 강하게 감아 특정 대상을 바라보지 않는다. 단검의 날끝은 화면 왼쪽 아래 방향으로 벽에 들어가고 손잡이는 오른쪽 위로 돌출된다. 칼은 머리 오른쪽 바로 옆에 있으며 얼굴이나 머리와 겹치지 않아 빗나간 위치가 읽힌다.",
        "built_space": "인물 뒤에는 거친 밝은색 벽과 오른쪽 수직 모서리가 보인다. 오른쪽 배경에 게시판 하나, 사무용 의자 하나, 책상 하나, 책상 위 스피커폰 하나와 서류 묶음 하나가 보인다. 참조의 사무실 요소는 유지하지만 벽 표면은 더 거칠고 돌출 벽 구획이 강조된다. 인물은 벽에 등을 가까이 댄 상반신 자세이며, 거울이나 반사는 없다.",
        "entities": "고령의 한국인 남성으로 읽히는 인물 한 명만 있다. 회색 옆가르마 머리, 깊은 이마와 눈가 주름, 짙은 남색 외투와 흰 셔츠는 참조와 대체로 맞지만 안경이 없다. 얼굴은 뚜렷하게 붉고 두 눈은 감겨 있다. 양손에 든 어두운 손수건이 코와 입을 가린다. 벽에 박힌 단검 하나와 배경의 스피커폰이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "단검은 날끝이 벽에 삽입된 지점으로 지지되며 떠 있지 않다. 손수건은 양손으로 잡고 있다. 어깨를 올리고 팔을 모아 움츠린 자세는 자연스럽고 등은 벽에 가까이 놓인다. 하체와 발은 프레임 밖이므로 서 있는 상태의 접지까지 확인할 수는 없지만, 공중에 뜬 신체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "남성은 턱을 당기고 두 눈을 질끈 감고 있으며 시선 대상은 없다. 단검의 날끝은 오른쪽 아래 방향으로 벽에 박혀 있고 손잡이는 왼쪽 위, 머리 쪽으로 돌출된다. 얼굴 앞을 가로막지는 않지만 손잡이 끝이 머리카락 윤곽과 일부 겹쳐, 머리와 완전히 분리하라는 배치에는 미달한다.",
        "built_space": "인물 바로 뒤 오른쪽에 밝은 상부와 회색 하부로 나뉜 벽 하나가 보이고 칼이 그 벽에 박혀 있다. 왼쪽 배경에는 책상 하나, 의자 하나, 스피커폰 하나가 있으며 위쪽에는 게시판 일부가 보인다. 참조의 투톤 벽과 목재 책상, 사무실 조명 분위기는 이어지지만 벽의 균열과 마모가 더 강하다. 반사면이나 중복된 고정 설비는 보이지 않는다.",
        "entities": "인물은 한 명이며 고령의 한국인 남성으로 읽힌다. 회색 머리, 깊은 주름, 남색 외투와 흰 셔츠는 참조에 부합하지만 안경은 없다. 붉어진 볼과 이마, 감긴 눈, 긴장된 입이 가려지지 않고 보인다. 한 손에 밝은 손수건을 쥐고 있다. 단검은 하나지만 참조의 칼보다 넓고 각진 형태이며 화면에서 크게 두드러진다. 스피커폰은 배경에 남아 있고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "칼날이 벽 안으로 들어간 접점과 주변 균열이 보여 벽이 칼을 지지한다. 손수건은 턱 아래의 손이 확실히 쥐고 있다. 인물은 어깨를 움츠리고 등과 오른쪽 어깨를 벽에 가깝게 둔 자세로, 놀라 물러난 동작으로 가능하다. 발과 하체는 보이지 않아 서 있는 상태를 직접 검증할 수 없으나 부유하거나 해부학적으로 불가능한 부분은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "단검을 머리와 겹치지 않게 배치하고 움츠림을 잘 표현했지만, 구도가 더 넓고 손수건이 얼굴 하반부를 가려 핵심인 붉어진 얼굴의 클로즈업이 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "질끈 감은 두 눈과 붉어진 얼굴을 온전히 보여주는 클로즈업이 더 충실하지만, 단검이 지나치게 두드러지고 손잡이가 머리카락과 일부 겹치며 참조의 안경이 빠졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "남성은 고개를 약간 숙이고 두 눈을 강하게 감아 특정 대상을 바라보지 않는다. 단검의 날끝은 화면 왼쪽 아래 방향으로 벽에 들어가고 손잡이는 오른쪽 위로 돌출된다. 칼은 머리 오른쪽 바로 옆에 있으며 얼굴이나 머리와 겹치지 않아 빗나간 위치가 읽힌다.",
        "built_space": "인물 뒤에는 거친 밝은색 벽과 오른쪽 수직 모서리가 보인다. 오른쪽 배경에 게시판 하나, 사무용 의자 하나, 책상 하나, 책상 위 스피커폰 하나와 서류 묶음 하나가 보인다. 참조의 사무실 요소는 유지하지만 벽 표면은 더 거칠고 돌출 벽 구획이 강조된다. 인물은 벽에 등을 가까이 댄 상반신 자세이며, 거울이나 반사는 없다.",
        "entities": "고령의 한국인 남성으로 읽히는 인물 한 명만 있다. 회색 옆가르마 머리, 깊은 이마와 눈가 주름, 짙은 남색 외투와 흰 셔츠는 참조와 대체로 맞지만 안경이 없다. 얼굴은 뚜렷하게 붉고 두 눈은 감겨 있다. 양손에 든 어두운 손수건이 코와 입을 가린다. 벽에 박힌 단검 하나와 배경의 스피커폰이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "단검은 날끝이 벽에 삽입된 지점으로 지지되며 떠 있지 않다. 손수건은 양손으로 잡고 있다. 어깨를 올리고 팔을 모아 움츠린 자세는 자연스럽고 등은 벽에 가까이 놓인다. 하체와 발은 프레임 밖이므로 서 있는 상태의 접지까지 확인할 수는 없지만, 공중에 뜬 신체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "남성은 턱을 당기고 두 눈을 질끈 감고 있으며 시선 대상은 없다. 단검의 날끝은 오른쪽 아래 방향으로 벽에 박혀 있고 손잡이는 왼쪽 위, 머리 쪽으로 돌출된다. 얼굴 앞을 가로막지는 않지만 손잡이 끝이 머리카락 윤곽과 일부 겹쳐, 머리와 완전히 분리하라는 배치에는 미달한다.",
        "built_space": "인물 바로 뒤 오른쪽에 밝은 상부와 회색 하부로 나뉜 벽 하나가 보이고 칼이 그 벽에 박혀 있다. 왼쪽 배경에는 책상 하나, 의자 하나, 스피커폰 하나가 있으며 위쪽에는 게시판 일부가 보인다. 참조의 투톤 벽과 목재 책상, 사무실 조명 분위기는 이어지지만 벽의 균열과 마모가 더 강하다. 반사면이나 중복된 고정 설비는 보이지 않는다.",
        "entities": "인물은 한 명이며 고령의 한국인 남성으로 읽힌다. 회색 머리, 깊은 주름, 남색 외투와 흰 셔츠는 참조에 부합하지만 안경은 없다. 붉어진 볼과 이마, 감긴 눈, 긴장된 입이 가려지지 않고 보인다. 한 손에 밝은 손수건을 쥐고 있다. 단검은 하나지만 참조의 칼보다 넓고 각진 형태이며 화면에서 크게 두드러진다. 스피커폰은 배경에 남아 있고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "칼날이 벽 안으로 들어간 접점과 주변 균열이 보여 벽이 칼을 지지한다. 손수건은 턱 아래의 손이 확실히 쥐고 있다. 인물은 어깨를 움츠리고 등과 오른쪽 어깨를 벽에 가깝게 둔 자세로, 놀라 물러난 동작으로 가능하다. 발과 하체는 보이지 않아 서 있는 상태를 직접 검증할 수 없으나 부유하거나 해부학적으로 불가능한 부분은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.179
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.929
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 묘사 (단검이 어떤 표면에도 박히지 않고 허공에 떠 있음)",
     "[gemini-pro] 물리적으로 불가능한 신체 구조 (손수건을 쥔 손가락이 비정상적으로 많고 뭉개짐)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트의 클로즈업 구도, 붉어진 얼굴로 움츠러든 표정, 벽에 박힌 단검의 위치를 정확히 구현했으나, 단검의 디자인이 레퍼런스와 다릅니다."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "단검이 허공에 떠 있고 손가락 구조가 왜곡되는 등 심각한 물리적 오류(Hard Violation)가 발생하여 지시를 충족하지 못했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 묘사 (단검이 어떤 표면에도 박히지 않고 허공에 떠 있음) / [gemini-pro] 물리적으로 불가능한 신체 구조 (손수건을 쥔 손가락이 비정상적으로 많고 뭉개짐)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1_sel.png",
    "asset_id": "1f4214c9-bdce-4ba8-b37d-181fdf81ae8a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-e360-7e41-8b4f-a21609cc7a1e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S18sh1"
  }
 },
 "S18sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:41:43.945187+00:00",
  "fingerprint": "7b163f9d132543ebc0ab6cbd01999646113198ec46e3de33ab44870abd452d3f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S18sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S18sh5_sel.png",
  "source_sha256": "9fdc7378f0b12c9175f516af4943f57f9e49e27e696a922f98e50725cfe2b4da",
  "file": "S18sh5_cine.png",
  "staged_sha256": "7b066d78274f0d41d745811e67ffc1e3f7b463320754b012d2590ff0dc1235df",
  "latency_ms": 10468
 },
 "S18sh6::signage": {
  "fp": "fbada9e8e6a451d6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S18sh6": {
  "input_fingerprint": "bded17df5e0b6fcf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 겁먹은 윤성찬을 향해 허리를 90도로 굽힌 채 고개를 들어 입꼬리를 길게 찢어 웃고 있는 박철진의 자세.\n\nLOCATION (lock): In the open space between the commander's desk and the visitor inside the militia office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office interior (Only a limited portion remains visible around the bowed figure); used as Preserve spatial context around the bend of the torso without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination neutral and the facial contrast controlled, allowing the threatening smile to remain plainly visible without an expressionistic lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains lodged in the wall, and the speakerphone remains in the office. 박철진: He no longer holds the knife and bends into a deep, ninety-degree bow.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 겁먹은 윤성찬을 향해 허리를 90도로 굽힌 채 고개를 들어 입꼬리를 길게 찢어 웃고 있는 박철진의 자세.\n\nLOCATION (lock): In the open space between the commander's desk and the visitor inside the militia office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office interior (Only a limited portion remains visible around the bowed figure); used as Preserve spatial context around the bend of the torso without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination neutral and the facial contrast controlled, allowing the threatening smile to remain plainly visible without an expressionistic lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains lodged in the wall, and the speakerphone remains in the office. 박철진: He no longer holds the knife and bends into a deep, ninety-degree bow.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 겁먹은 윤성찬을 향해 허리를 90도로 굽힌 채 고개를 들어 입꼬리를 길게 찢어 웃고 있는 박철진의 자세.\n\nLOCATION (lock): In the open space between the commander's desk and the visitor inside the militia office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office interior (Only a limited portion remains visible around the bowed figure); used as Preserve spatial context around the bend of the torso without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination neutral and the facial contrast controlled, allowing the threatening smile to remain plainly visible without an expressionistic lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains lodged in the wall, and the speakerphone remains in the office. 박철진: He no longer holds the knife and bends into a deep, ninety-degree bow.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진이 허리를 90도로 굽히고 고개를 들어 화면 밖을 보며 미소 짓고 있음.",
    "built_space": "사무실 내부. 배경에 책상과 전화기가 있으며 오른쪽 벽면에 칼이 꽂혀 있음.",
    "entities": "박철진(검은 전술복, 붉은 완장) 단독 등장. 벽에 꽂힌 거대한 칼.",
    "hard_violations": [],
    "physics": "두 발로 서서 허리를 깊게 숙이고 뒷짐을 진 안정적인 자세. 칼은 벽에 단단히 고정됨."
   },
   {
    "label": "B",
    "direction": "박철진이 상체를 약간 숙인 채 왼쪽 전경에 있는 노인을 향해 시선을 두고 웃음.",
    "built_space": "사무실 내부. 뒤쪽 책상 위에 스피커폰이 있고 오른쪽 벽에 칼이 꽂혀 있음.",
    "entities": "박철진 등장. 프롬프트가 금지한 노인(이전 샷의 인물)이 좌측 전경에 등장함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트의 PEOPLE 목록에 명시되지 않은 금지된 인물(노인)이 프레임에 포함됨.",
     "[gpt-high] 이번 숏에서 제외하도록 명시한 이전 숏의 노인을 왼쪽 전경에 머리와 상체까지 재등장시켰다."
    ],
    "physics": "상체를 가볍게 숙인 자세. 칼은 벽에 정상적으로 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "허리를 90도로 굽힌 지정된 자세와 단독 등장 조건을 정확히 충족하여 프롬프트의 의도를 잘 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시적으로 금지한 이전 샷의 인물을 화면에 포함시키는 치명적인 오류를 범함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 허리를 90도로 굽히고 고개를 들어 화면 밖을 보며 미소 짓고 있음.",
        "built_space": "사무실 내부. 배경에 책상과 전화기가 있으며 오른쪽 벽면에 칼이 꽂혀 있음.",
        "entities": "박철진(검은 전술복, 붉은 완장) 단독 등장. 벽에 꽂힌 거대한 칼.",
        "hard_violations": [],
        "physics": "두 발로 서서 허리를 깊게 숙이고 뒷짐을 진 안정적인 자세. 칼은 벽에 단단히 고정됨."
       },
       {
        "label": "B",
        "direction": "박철진이 상체를 약간 숙인 채 왼쪽 전경에 있는 노인을 향해 시선을 두고 웃음.",
        "built_space": "사무실 내부. 뒤쪽 책상 위에 스피커폰이 있고 오른쪽 벽에 칼이 꽂혀 있음.",
        "entities": "박철진 등장. 프롬프트가 금지한 노인(이전 샷의 인물)이 좌측 전경에 등장함.",
        "hard_violations": [
         "프롬프트의 PEOPLE 목록에 명시되지 않은 금지된 인물(노인)이 프레임에 포함됨."
        ],
        "physics": "상체를 가볍게 숙인 자세. 칼은 벽에 정상적으로 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "허리를 90도로 굽힌 지정된 자세와 단독 등장 조건을 정확히 충족하여 프롬프트의 의도를 잘 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시적으로 금지한 이전 샷의 인물을 화면에 포함시키는 치명적인 오류를 범함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 허리를 90도로 굽히고 고개를 들어 화면 밖을 보며 미소 짓고 있음.",
        "built_space": "사무실 내부. 배경에 책상과 전화기가 있으며 오른쪽 벽면에 칼이 꽂혀 있음.",
        "entities": "박철진(검은 전술복, 붉은 완장) 단독 등장. 벽에 꽂힌 거대한 칼.",
        "hard_violations": [],
        "physics": "두 발로 서서 허리를 깊게 숙이고 뒷짐을 진 안정적인 자세. 칼은 벽에 단단히 고정됨."
       },
       {
        "label": "B",
        "direction": "박철진이 상체를 약간 숙인 채 왼쪽 전경에 있는 노인을 향해 시선을 두고 웃음.",
        "built_space": "사무실 내부. 뒤쪽 책상 위에 스피커폰이 있고 오른쪽 벽에 칼이 꽂혀 있음.",
        "entities": "박철진 등장. 프롬프트가 금지한 노인(이전 샷의 인물)이 좌측 전경에 등장함.",
        "hard_violations": [
         "프롬프트의 PEOPLE 목록에 명시되지 않은 금지된 인물(노인)이 프레임에 포함됨."
        ],
        "physics": "상체를 가볍게 숙인 자세. 칼은 벽에 정상적으로 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "상대를 바라보는 웃음은 명확하지만, 금지된 이전 숏의 노인을 전경에 재등장시켰고 허리도 90도로 굽히지 않았다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "박철진만 등장하는 미디엄 숏과 90도에 가까운 깊은 절은 충실하지만, 고개를 들어 상대를 위협하기보다 얼굴을 옆으로 돌려 카메라 쪽을 보는 인상이 강하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 눈과 웃는 얼굴은 왼쪽 전경 노인의 얼굴을 향한다. 상대를 향한 위협이라는 관계는 분명하다. 오른쪽 칼은 칼끝이 벽에 들어가고 손잡이가 왼쪽 위로 나와 있어 누구를 겨누거나 들고 있는 상태는 아니다.",
        "built_space": "뒤쪽에 나무 책상 하나, 그 위 전화기 하나, 일부 가려진 의자 등받이 하나가 보인다. 오른쪽에는 칼 하나가 박힌 균열 있는 회백색 이중 도장 벽과 게시판 일부가 있다. 박철진은 책상 앞에 있지만, 왼쪽 전경의 노인이 화면을 크게 가려 전경 장애물 없이 굽힌 몸을 보여주라는 구성을 어긴다.",
        "entities": "짧은 검은 머리의 중년 한국인 남성 박철진은 참조의 얼굴과 대체로 유사하며, 짙은 회색 군복·붉은 완장·장비 벨트를 착용했다. 그러나 이전 숏 노인의 회색 머리와 짙은 정장 차림까지 재등장했다. 벽의 칼과 책상 위 전화기는 유지되며 판독 가능한 글자는 없다.",
        "hard_violations": [
         "이번 숏에서 제외하도록 명시한 이전 숏의 노인을 왼쪽 전경에 머리와 상체까지 재등장시켰다."
        ],
        "physics": "박철진은 허리에서 앞으로 기울고 하체는 화면 아래로 이어진다. 발은 잘렸지만 떠 있는 신체는 아니며 자세 자체는 가능하다. 다만 몸통은 수평이 아니라 비스듬해 요구된 90도 절보다 얕다. 칼은 벽에 박힌 날끝으로 지지되고 전화기는 책상 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "몸통은 왼쪽 골반에서 오른쪽 벽 쪽으로 거의 수평으로 뻗는다. 얼굴은 옆으로 틀어져 눈이 카메라 근처를 향하며, 화면 밖 윤성찬에게 시선이 닿는지는 확인하기 어렵다. 칼끝은 오른쪽 벽에 박혀 있고 손잡이는 왼쪽 위를 향한다.",
        "built_space": "왼쪽 뒤 나무 책상 하나와 그 위 일부 가려진 전화기 하나, 오른쪽의 균열 있는 이중 도장 벽, 벽에 박힌 칼 하나, 뒤쪽 게시판 일부가 보인다. 박철진은 책상 앞 열린 공간을 차지하고 있으며 다른 사람이 전경을 막지 않는다. 참조의 벽 재질과 낮의 중립적인 조명이 유지된다.",
        "entities": "보이는 사람은 박철진 한 명이다. 중년 한국인 남성의 얼굴, 짧은 검은 머리, 짙은 회색 군복, 붉은 완장과 허리 장비가 참조와 대체로 맞는다. 입을 길게 벌린 웃음도 정상적인 사람의 표정으로 표현됐다. 완장에는 참조보다 두드러진 검은 마름모 장식이 있다. 칼은 손에 없고 벽에 남아 있으며 전화기도 사무실에 남아 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "허벅지가 화면 아래로 이어지고 골반에서 몸통을 거의 직각으로 접어 하체로 무게를 받는 자세다. 발은 프레임 밖이지만 몸이 공중에 떠 있는 모습은 아니다. 두 손을 등 뒤에서 맞잡은 자세도 가능하다. 목은 고개를 위로 든 모습보다 옆으로 돌린 모습에 가깝지만 해부학적으로 불가능하지 않다. 칼은 벽에 박힌 날끝으로, 전화기는 책상으로 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상대를 바라보는 웃음은 명확하지만, 금지된 이전 숏의 노인을 전경에 재등장시켰고 허리도 90도로 굽히지 않았다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "박철진만 등장하는 미디엄 숏과 90도에 가까운 깊은 절은 충실하지만, 고개를 들어 상대를 위협하기보다 얼굴을 옆으로 돌려 카메라 쪽을 보는 인상이 강하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 눈과 웃는 얼굴은 왼쪽 전경 노인의 얼굴을 향한다. 상대를 향한 위협이라는 관계는 분명하다. 오른쪽 칼은 칼끝이 벽에 들어가고 손잡이가 왼쪽 위로 나와 있어 누구를 겨누거나 들고 있는 상태는 아니다.",
        "built_space": "뒤쪽에 나무 책상 하나, 그 위 전화기 하나, 일부 가려진 의자 등받이 하나가 보인다. 오른쪽에는 칼 하나가 박힌 균열 있는 회백색 이중 도장 벽과 게시판 일부가 있다. 박철진은 책상 앞에 있지만, 왼쪽 전경의 노인이 화면을 크게 가려 전경 장애물 없이 굽힌 몸을 보여주라는 구성을 어긴다.",
        "entities": "짧은 검은 머리의 중년 한국인 남성 박철진은 참조의 얼굴과 대체로 유사하며, 짙은 회색 군복·붉은 완장·장비 벨트를 착용했다. 그러나 이전 숏 노인의 회색 머리와 짙은 정장 차림까지 재등장했다. 벽의 칼과 책상 위 전화기는 유지되며 판독 가능한 글자는 없다.",
        "hard_violations": [
         "이번 숏에서 제외하도록 명시한 이전 숏의 노인을 왼쪽 전경에 머리와 상체까지 재등장시켰다."
        ],
        "physics": "박철진은 허리에서 앞으로 기울고 하체는 화면 아래로 이어진다. 발은 잘렸지만 떠 있는 신체는 아니며 자세 자체는 가능하다. 다만 몸통은 수평이 아니라 비스듬해 요구된 90도 절보다 얕다. 칼은 벽에 박힌 날끝으로 지지되고 전화기는 책상 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "몸통은 왼쪽 골반에서 오른쪽 벽 쪽으로 거의 수평으로 뻗는다. 얼굴은 옆으로 틀어져 눈이 카메라 근처를 향하며, 화면 밖 윤성찬에게 시선이 닿는지는 확인하기 어렵다. 칼끝은 오른쪽 벽에 박혀 있고 손잡이는 왼쪽 위를 향한다.",
        "built_space": "왼쪽 뒤 나무 책상 하나와 그 위 일부 가려진 전화기 하나, 오른쪽의 균열 있는 이중 도장 벽, 벽에 박힌 칼 하나, 뒤쪽 게시판 일부가 보인다. 박철진은 책상 앞 열린 공간을 차지하고 있으며 다른 사람이 전경을 막지 않는다. 참조의 벽 재질과 낮의 중립적인 조명이 유지된다.",
        "entities": "보이는 사람은 박철진 한 명이다. 중년 한국인 남성의 얼굴, 짧은 검은 머리, 짙은 회색 군복, 붉은 완장과 허리 장비가 참조와 대체로 맞는다. 입을 길게 벌린 웃음도 정상적인 사람의 표정으로 표현됐다. 완장에는 참조보다 두드러진 검은 마름모 장식이 있다. 칼은 손에 없고 벽에 남아 있으며 전화기도 사무실에 남아 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "허벅지가 화면 아래로 이어지고 골반에서 몸통을 거의 직각으로 접어 하체로 무게를 받는 자세다. 발은 프레임 밖이지만 몸이 공중에 떠 있는 모습은 아니다. 두 손을 등 뒤에서 맞잡은 자세도 가능하다. 목은 고개를 위로 든 모습보다 옆으로 돌린 모습에 가깝지만 해부학적으로 불가능하지 않다. 칼은 벽에 박힌 날끝으로, 전화기는 책상으로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.857
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.607
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트의 PEOPLE 목록에 명시되지 않은 금지된 인물(노인)이 프레임에 포함됨.",
     "[gpt-high] 이번 숏에서 제외하도록 명시한 이전 숏의 노인을 왼쪽 전경에 머리와 상체까지 재등장시켰다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 607
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "허리를 90도로 굽힌 지정된 자세와 단독 등장 조건을 정확히 충족하여 프롬프트의 의도를 잘 구현함."
   },
   {
    "label": "B",
    "score": 607,
    "verdict_ko": "프롬프트에서 명시적으로 금지한 이전 샷의 인물을 화면에 포함시키는 치명적인 오류를 범함.  ★위반: [gemini-pro] 프롬프트의 PEOPLE 목록에 명시되지 않은 금지된 인물(노인)이 프레임에 포함됨. / [gpt-high] 이번 숏에서 제외하도록 명시한 이전 숏의 노인을 왼쪽 전경에 머리와 상체까지 재등장시켰다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh5_sel.png",
    "asset_id": "4708e371-18e5-48be-bd1f-e07bbf84f3cf",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-e505-7319-bb98-1e300bd25872",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S18sh5"
  }
 },
 "S18sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:42:36.690310+00:00",
  "fingerprint": "aad0e4666395b4e667b6711f2f105c2bc7567bbbfc4a10c1652225029fcdfbcf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S18sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S18sh6_sel.png",
  "source_sha256": "276221ea6b9cb80f68492203106355433f9401acac6e11ecf889292af63d272b",
  "file": "S18sh6_cine.png",
  "staged_sha256": "cd54c1353056f837478dbebe1d64c7d8b7ab46adca45ed744523b2724c2ceb5c",
  "latency_ms": 10142
 },
 "S19sh3::confined_fp_apt": {
  "applies": true,
  "reason_ko": "이 숏은 차량 내부라는 제한된 공간에서 진행되며, 뒷좌석에 앉은 인물이 앞좌석을 향해 몸을 기울이는 구체적인 행동을 묘사하고 있습니다. 따라서 앞뒤 좌석의 공간적 배치와 인물의 위치 관계를 정확하게 표현하지 않으면 상황을 이해하기 어려우므로 평면도 레이아웃 보조가 필요합니다.",
  "input_fingerprint": "5244c6375b6673fe"
 },
 "S19sh3::signage": {
  "fp": "a689f1a4fcea29e3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::3a46c2692d54": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_3a46c2692d54.png",
  "place_text": "Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows.",
  "input_fingerprint": "4fdd5d2a3a8abf12"
 },
 "S19sh3::confined_fp": {
  "reads": {
   "controls": "A steering wheel is located at the front-left driver's seat.",
   "mirrors": "No mirrors are present in the diagram.",
   "camera": "The camera is positioned between the front driver and passenger seats, pointing directly backward toward the rear center seat.",
   "occupants": "Yoon Sung-chan (윤성찬) is seated in the Rear Center position. All other seats are empty."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned between the front seats, looking straight back into the rear cabin. On the near left edge of the screen, the inner side of the empty driver's seat frames the view. On the near right edge, the inner side of the empty front passenger seat provides the opposite frame. In the center of the frame sits Yoon Sung-chan, occupying the rear center seat and facing forward toward the camera. In the background, flanking him on the far left and far right, are the empty rear left and rear right seats.",
  "fixed": true,
  "input_fingerprint": "9a6fa3d50bd62343"
 },
 "era_assess::1dd1249d0430d782": {
  "subjects": [],
  "subject_text": "윤성찬의 자동차 내부\n앞좌석과 뒷좌석이 구분된 고급 세단 실내. 운전대와 계기판이 앞쪽에 있고 뒷좌석 옆으로 창문과 도어 내장재가 이어진다.",
  "identity": "canonical",
  "scope_id": "L172",
  "scope_role": "location_interior",
  "scope_sha": "e79f96e1a718ce2b"
 },
 "S19sh3": {
  "input_fingerprint": "d33cad44275d28cd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 주먹을 꽉 쥔 채 앞좌석을 향해 상체를 바짝 내민 윤성찬의 격양된 얼굴.\n\nLOCATION (lock): Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Front seats (Their inner edges remain visible beside the camera's viewing gap) — The inner sides and partial rear faces bracket the view toward the rear passenger; used as Create a narrow spatial channel around the face and fist without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination within the car, retaining readable facial tension and restrained contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 윤성찬: He is seated inside the car and retains his handkerchief.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, looking straight back into the rear cabin. On the near left edge of the screen, the inner side of the empty driver's seat frames the view. On the near right edge, the inner side of the empty front passenger seat provides the opposite frame. In the center of the frame sits Yoon Sung-chan, occupying the rear center seat and facing forward toward the camera. In the background, flanking him on the far left and far right, are the empty rear left and rear right seats.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 주먹을 꽉 쥔 채 앞좌석을 향해 상체를 바짝 내민 윤성찬의 격양된 얼굴.\n\nLOCATION (lock): Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination within the car, retaining readable facial tension and restrained contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 윤성찬: He is seated inside the car and retains his handkerchief.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, looking straight back into the rear cabin. On the near left edge of the screen, the inner side of the empty driver's seat frames the view. On the near right edge, the inner side of the empty front passenger seat provides the opposite frame. In the center of the frame sits Yoon Sung-chan, occupying the rear center seat and facing forward toward the camera. In the background, flanking him on the far left and far right, are the empty rear left and rear right seats.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 주먹을 꽉 쥔 채 앞좌석을 향해 상체를 바짝 내민 윤성찬의 격양된 얼굴.\n\nLOCATION (lock): Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination within the car, retaining readable facial tension and restrained contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 윤성찬: He is seated inside the car and retains his handkerchief.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S19sh3_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "윤성찬",
     "path": "<bytes:886743>",
     "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S19sh3_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "윤성찬",
     "path": "<bytes:886743>",
     "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 앞좌석 사이의 카메라를 향하고 있다.",
    "built_space": "차량 뒷좌석이며 두 앞좌석이 화면 양옆을 프레이밍하고 있다.",
    "entities": "인물의 외모와 의상은 일치하지만 손수건이 빈 좌석에 방치되어 있다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조: 양손 모두 엄지손가락이 손의 바깥쪽(새끼손가락 방향)에 위치함",
     "[gpt-high] 앞좌석의 뒷면이어야 할 전경에 착석면과 측면 볼스터가 드러나, 두 앞좌석이 뒷좌석 승객 쪽을 향한 것처럼 배치되어 있다."
    ],
    "physics": "상체를 숙인 채 양손을 허공에 띄우고 코어 근육으로 자세를 지탱하고 있다."
   },
   {
    "label": "B",
    "direction": "시선은 카메라를 향하고 상체를 앞좌석 사이로 깊게 내밀고 있다.",
    "built_space": "차량 내부로 앞좌석 두 개가 양옆의 시야를 좁게 형성하고 있다.",
    "entities": "노인의 얼굴과 의상이 일치하며 오른손에 지정된 손수건을 쥐고 있다.",
    "hard_violations": [],
    "physics": "앞으로 쏠린 체중을 오른손으로 앞좌석 등받이를 짚어 지탱하고 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 격양된 표정과 바짝 내민 자세를 훌륭히 구현했으며, 손수건 소지 조건도 완벽하게 충족했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "양손의 해부학적 오류가 있는 치명적 위반을 범했으며 표정의 강도가 약하고 손수건을 소지하지 않았습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 앞좌석 사이의 카메라를 향하고 있다.",
        "built_space": "차량 뒷좌석이며 두 앞좌석이 화면 양옆을 프레이밍하고 있다.",
        "entities": "인물의 외모와 의상은 일치하지만 손수건이 빈 좌석에 방치되어 있다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조: 양손 모두 엄지손가락이 손의 바깥쪽(새끼손가락 방향)에 위치함"
        ],
        "physics": "상체를 숙인 채 양손을 허공에 띄우고 코어 근육으로 자세를 지탱하고 있다."
       },
       {
        "label": "B",
        "direction": "시선은 카메라를 향하고 상체를 앞좌석 사이로 깊게 내밀고 있다.",
        "built_space": "차량 내부로 앞좌석 두 개가 양옆의 시야를 좁게 형성하고 있다.",
        "entities": "노인의 얼굴과 의상이 일치하며 오른손에 지정된 손수건을 쥐고 있다.",
        "hard_violations": [],
        "physics": "앞으로 쏠린 체중을 오른손으로 앞좌석 등받이를 짚어 지탱하고 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 격양된 표정과 바짝 내민 자세를 훌륭히 구현했으며, 손수건 소지 조건도 완벽하게 충족했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "양손의 해부학적 오류가 있는 치명적 위반을 범했으며 표정의 강도가 약하고 손수건을 소지하지 않았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 앞좌석 사이의 카메라를 향하고 있다.",
        "built_space": "차량 뒷좌석이며 두 앞좌석이 화면 양옆을 프레이밍하고 있다.",
        "entities": "인물의 외모와 의상은 일치하지만 손수건이 빈 좌석에 방치되어 있다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조: 양손 모두 엄지손가락이 손의 바깥쪽(새끼손가락 방향)에 위치함"
        ],
        "physics": "상체를 숙인 채 양손을 허공에 띄우고 코어 근육으로 자세를 지탱하고 있다."
       },
       {
        "label": "B",
        "direction": "시선은 카메라를 향하고 상체를 앞좌석 사이로 깊게 내밀고 있다.",
        "built_space": "차량 내부로 앞좌석 두 개가 양옆의 시야를 좁게 형성하고 있다.",
        "entities": "노인의 얼굴과 의상이 일치하며 오른손에 지정된 손수건을 쥐고 있다.",
        "hard_violations": [],
        "physics": "앞으로 쏠린 체중을 오른손으로 앞좌석 등받이를 짚어 지탱하고 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "앞좌석 사이로 얼굴을 바짝 내밀고 손수건을 움켜쥔 격양된 순간을 더 가까이 포착했으며, 앞좌석 뒷면과 신체 지지도 자연스럽지만 참조의 안경은 빠졌다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "무릎까지 보이는 넓은 구도로 얼굴 클로즈업과 격양된 순간이 약해졌고, 앞좌석의 착석면이 뒷좌석 쪽을 향하는 공간 오류가 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 상체가 뒷좌석 중앙에서 앞좌석 사이로 전진해 있으며, 시선은 카메라 부근의 차량 앞쪽을 향한다. 별도의 상대 인물은 보이지 않는다. 두 손은 주먹을 쥐고 있고, 화면 왼쪽 주먹은 앞좌석 등받이 위로 나와 있어 요구된 전방 지향 동작이 분명하다.",
        "built_space": "전경 좌우에 앞좌석 두 개의 머리받침과 등받이 뒷면이 보이며, 뒷면 수납부도 확인된다. 두 좌석의 안쪽 가장자리가 얼굴과 주먹을 가리지 않고 통로를 만든다. 뒤에는 벤치형 뒷좌석과 양쪽 머리받침 일부, 후면창과 측면창, 천장 양쪽 손잡이가 보인다. 인물은 뒷좌석 중앙에서 앞으로 숙이고 있으며 카메라의 후방 관찰 방향과 좌석 방향이 맞는다.",
        "entities": "고령의 한국인 남성 설정에 부합하는 인물 한 명이 보인다. 회색으로 센 옆가르마 머리, 깊은 이마와 눈가 주름, 짙은 남색 외투와 정장, 흰 셔츠, 어두운 넥타이는 참조와 대체로 일치한다. 참조의 안경은 없다. 화면 왼쪽 주먹에 구겨진 손수건이 잡혀 있고 반대쪽 손목에는 시계가 보인다. 입을 벌리고 미간을 찌푸린 표정은 격양된 상태를 뚜렷하게 전달한다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 전완과 주먹이 앞좌석 등받이 상단에 걸쳐져 전방으로 기운 상체를 지지한다. 손수건은 손가락에 붙잡혀 있으며 다른 주먹도 손목과 팔에 자연스럽게 연결된다. 골반과 발은 프레임 밖이므로 착석 접촉 자체는 확인되지 않지만, 보이는 자세는 뒷좌석에 앉아 몸을 앞으로 숙이는 동작으로 가능하다. 공중에 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "뒷좌석 중앙의 남성이 상체를 앞으로 숙이고 카메라 부근의 차량 앞쪽을 바라본다. 두 주먹은 복부 앞에 나란히 모여 있으며 특정 대상을 가격하려는 방향은 아니다. 전방으로 숙이는 관계는 맞지만, 얼굴이 앞좌석 사이로 바짝 돌출된 정도와 분노의 표현은 A보다 약하다.",
        "built_space": "전경에 앞좌석 두 개와 각각의 머리받침이 있고, 뒤에는 벤치형 뒷좌석과 바깥쪽 머리받침 두 개가 보인다. 중앙 머리받침 위치는 인물에게 가려져 있다. 측면창 두 곳, 후면창, 천장 손잡이 두 개와 중앙 실내등 하나가 보인다. 인물은 뒷좌석 중앙에 앉아 있지만 전경 좌석들은 타공된 중앙 쿠션과 신체를 감싸는 측면 볼스터 등 착석면의 형태를 카메라에 드러낸다. 요구된 앞좌석 뒷면 방향과 맞지 않는다. 얼굴뿐 아니라 허벅지와 무릎까지 보여 클로즈업보다 넓다.",
        "entities": "고령의 한국인 남성 설정에 부합하는 인물 한 명이며, 센 옆가르마 머리와 깊은 주름, 남색 외투와 정장, 흰 셔츠, 넥타이, 줄무늬 바지가 참조와 대체로 일치한다. 참조의 안경은 없다. 손목시계와 가슴 주머니의 접힌 천이 보이고, 화면 오른쪽 좌석 위에도 구겨진 회색 천이 놓여 있으나 어느 것이 유지해야 할 손수건인지는 확실하지 않다. 두 주먹은 명확하지만 표정은 크게 격앙되기보다 굳고 억눌린 인상이다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "앞좌석의 뒷면이어야 할 전경에 착석면과 측면 볼스터가 드러나, 두 앞좌석이 뒷좌석 승객 쪽을 향한 것처럼 배치되어 있다."
        ],
        "physics": "골반과 허벅지가 뒷좌석 방석에 지지되고 무릎이 앞쪽으로 내려가 있어 착석은 자연스럽다. 허리를 굽혀 상체를 기울인 자세도 가능하며 두 주먹은 굽힌 팔로 지지된다. 옆의 회색 천은 좌석 위에 놓여 있다. 지지 없이 떠 있는 물체나 신체는 없고, 문제는 신체 역학이 아니라 앞좌석의 방향이다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "앞좌석 사이로 얼굴을 바짝 내밀고 손수건을 움켜쥔 격양된 순간을 더 가까이 포착했으며, 앞좌석 뒷면과 신체 지지도 자연스럽지만 참조의 안경은 빠졌다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "무릎까지 보이는 넓은 구도로 얼굴 클로즈업과 격양된 순간이 약해졌고, 앞좌석의 착석면이 뒷좌석 쪽을 향하는 공간 오류가 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 상체가 뒷좌석 중앙에서 앞좌석 사이로 전진해 있으며, 시선은 카메라 부근의 차량 앞쪽을 향한다. 별도의 상대 인물은 보이지 않는다. 두 손은 주먹을 쥐고 있고, 화면 왼쪽 주먹은 앞좌석 등받이 위로 나와 있어 요구된 전방 지향 동작이 분명하다.",
        "built_space": "전경 좌우에 앞좌석 두 개의 머리받침과 등받이 뒷면이 보이며, 뒷면 수납부도 확인된다. 두 좌석의 안쪽 가장자리가 얼굴과 주먹을 가리지 않고 통로를 만든다. 뒤에는 벤치형 뒷좌석과 양쪽 머리받침 일부, 후면창과 측면창, 천장 양쪽 손잡이가 보인다. 인물은 뒷좌석 중앙에서 앞으로 숙이고 있으며 카메라의 후방 관찰 방향과 좌석 방향이 맞는다.",
        "entities": "고령의 한국인 남성 설정에 부합하는 인물 한 명이 보인다. 회색으로 센 옆가르마 머리, 깊은 이마와 눈가 주름, 짙은 남색 외투와 정장, 흰 셔츠, 어두운 넥타이는 참조와 대체로 일치한다. 참조의 안경은 없다. 화면 왼쪽 주먹에 구겨진 손수건이 잡혀 있고 반대쪽 손목에는 시계가 보인다. 입을 벌리고 미간을 찌푸린 표정은 격양된 상태를 뚜렷하게 전달한다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 전완과 주먹이 앞좌석 등받이 상단에 걸쳐져 전방으로 기운 상체를 지지한다. 손수건은 손가락에 붙잡혀 있으며 다른 주먹도 손목과 팔에 자연스럽게 연결된다. 골반과 발은 프레임 밖이므로 착석 접촉 자체는 확인되지 않지만, 보이는 자세는 뒷좌석에 앉아 몸을 앞으로 숙이는 동작으로 가능하다. 공중에 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "뒷좌석 중앙의 남성이 상체를 앞으로 숙이고 카메라 부근의 차량 앞쪽을 바라본다. 두 주먹은 복부 앞에 나란히 모여 있으며 특정 대상을 가격하려는 방향은 아니다. 전방으로 숙이는 관계는 맞지만, 얼굴이 앞좌석 사이로 바짝 돌출된 정도와 분노의 표현은 A보다 약하다.",
        "built_space": "전경에 앞좌석 두 개와 각각의 머리받침이 있고, 뒤에는 벤치형 뒷좌석과 바깥쪽 머리받침 두 개가 보인다. 중앙 머리받침 위치는 인물에게 가려져 있다. 측면창 두 곳, 후면창, 천장 손잡이 두 개와 중앙 실내등 하나가 보인다. 인물은 뒷좌석 중앙에 앉아 있지만 전경 좌석들은 타공된 중앙 쿠션과 신체를 감싸는 측면 볼스터 등 착석면의 형태를 카메라에 드러낸다. 요구된 앞좌석 뒷면 방향과 맞지 않는다. 얼굴뿐 아니라 허벅지와 무릎까지 보여 클로즈업보다 넓다.",
        "entities": "고령의 한국인 남성 설정에 부합하는 인물 한 명이며, 센 옆가르마 머리와 깊은 주름, 남색 외투와 정장, 흰 셔츠, 넥타이, 줄무늬 바지가 참조와 대체로 일치한다. 참조의 안경은 없다. 손목시계와 가슴 주머니의 접힌 천이 보이고, 화면 오른쪽 좌석 위에도 구겨진 회색 천이 놓여 있으나 어느 것이 유지해야 할 손수건인지는 확실하지 않다. 두 주먹은 명확하지만 표정은 크게 격앙되기보다 굳고 억눌린 인상이다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "앞좌석의 뒷면이어야 할 전경에 착석면과 측면 볼스터가 드러나, 두 앞좌석이 뒷좌석 승객 쪽을 향한 것처럼 배치되어 있다."
        ],
        "physics": "골반과 허벅지가 뒷좌석 방석에 지지되고 무릎이 앞쪽으로 내려가 있어 착석은 자연스럽다. 허리를 굽혀 상체를 기울인 자세도 가능하며 두 주먹은 굽힌 팔로 지지된다. 옆의 회색 천은 좌석 위에 놓여 있다. 지지 없이 떠 있는 물체나 신체는 없고, 문제는 신체 역학이 아니라 앞좌석의 방향이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.804,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.554,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조: 양손 모두 엄지손가락이 손의 바깥쪽(새끼손가락 방향)에 위치함",
     "[gpt-high] 앞좌석의 뒷면이어야 할 전경에 착석면과 측면 볼스터가 드러나, 두 앞좌석이 뒷좌석 승객 쪽을 향한 것처럼 배치되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 554
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 격양된 표정과 바짝 내민 자세를 훌륭히 구현했으며, 손수건 소지 조건도 완벽하게 충족했습니다."
   },
   {
    "label": "A",
    "score": 554,
    "verdict_ko": "양손의 해부학적 오류가 있는 치명적 위반을 범했으며 표정의 강도가 약하고 손수건을 소지하지 않았습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학적 구조: 양손 모두 엄지손가락이 손의 바깥쪽(새끼손가락 방향)에 위치함 / [gpt-high] 앞좌석의 뒷면이어야 할 전경에 착석면과 측면 볼스터가 드러나, 두 앞좌석이 뒷좌석 승객 쪽을 향한 것처럼 배치되어 있다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S19sh3_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "윤성찬",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-e6a7-7408-90ab-a93d4bb2c9c0",
  "confined_fp": {
   "base_key": "confinedfp::3a46c2692d54",
   "apt_reason": "이 숏은 차량 내부라는 제한된 공간에서 진행되며, 뒷좌석에 앉은 인물이 앞좌석을 향해 몸을 기울이는 구체적인 행동을 묘사하고 있습니다. 따라서 앞뒤 좌석의 공간적 배치와 인물의 위치 관계를 정확하게 표현하지 않으면 상황을 이해하기 어려우므로 평면도 레이아웃 보조가 필요합니다.",
   "fixed": true,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S19sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:44:22.796917+00:00",
  "fingerprint": "3a7f8bc9e62aaae53021ed083e9084f1a5573631563ed6e5d21788be57dee0b1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S19sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S19sh3_sel.png",
  "source_sha256": "6c7ff2c7a295e806db88763b46880ede5f0797233dbfbe1659959ff61748ecea",
  "file": "S19sh3_cine.png",
  "staged_sha256": "ba3edd03fe30b54e4ca98df44842d35a5df60c15e56353a6a0c486d354dd0274",
  "latency_ms": 10314
 },
 "S20sh3::signage": {
  "fp": "f1d679db43b52ec9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S20sh3": {
  "input_fingerprint": "cb2da4ddd8ac9daa",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 모퉁이 뒤에 바짝 숨은 채 눈이 커진 현우와 페드로의 얼굴.\n\nLOCATION (lock): Behind a container corner near the searched family home, along an outdoor lane in the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container corner (Concealing both men from the search beyond it) — The concealed-side face ends at a narrow vertical edge beside their sightline; used as Make the hiding position immediately readable while leaving both faces unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and controlled contrast, keeping both startled faces legible without inventing atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container is the home marked 7-31. 현우: He is hiding near his container, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed. 페드로: He is hiding near the container.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 모퉁이 뒤에 바짝 숨은 채 눈이 커진 현우와 페드로의 얼굴.\n\nLOCATION (lock): Behind a container corner near the searched family home, along an outdoor lane in the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container corner (Concealing both men from the search beyond it) — The concealed-side face ends at a narrow vertical edge beside their sightline; used as Make the hiding position immediately readable while leaving both faces unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and controlled contrast, keeping both startled faces legible without inventing atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container is the home marked 7-31. 현우: He is hiding near his container, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed. 페드로: He is hiding near the container.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 모퉁이 뒤에 바짝 숨은 채 눈이 커진 현우와 페드로의 얼굴.\n\nLOCATION (lock): Behind a container corner near the searched family home, along an outdoor lane in the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container corner (Concealing both men from the search beyond it) — The concealed-side face ends at a narrow vertical edge beside their sightline; used as Make the hiding position immediately readable while leaving both faces unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and controlled contrast, keeping both startled faces legible without inventing atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container is the home marked 7-31. 현우: He is hiding near his container, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed. 페드로: He is hiding near the container.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 사람 모두 컨테이너 모서리 너머 우측 바깥쪽을 주시하고 있음.",
    "built_space": "좌측에 파란색 컨테이너 모서리가 수직으로 배치되었고, 우측 뒤편으로 야외 통로와 다른 컨테이너들이 원근감 있게 보임.",
    "entities": "얼굴 상처와 회색 셔츠를 착용한 현우, 비니와 짙은 후드를 입은 페드로 모두 레퍼런스의 외형과 일치함. 지시되지 않은 다른 인물은 없음.",
    "hard_violations": [],
    "physics": "두 사람 모두 땅을 딛고 서서 컨테이너에 몸을 밀착시키고 있으며, 현우의 손이 모서리를 자연스럽게 짚고 지탱함."
   },
   {
    "label": "B",
    "direction": "두 사람 모두 프레임 우측 배경 쪽으로 시선을 향하고 있음.",
    "built_space": "파란색 컨테이너 벽면 앞에 배치되었으나 모서리에 숨어 있는 형태가 아니며, 우측으로 통로가 보임.",
    "entities": "현우와 페드로의 상반신 외형은 레퍼런스와 일치하나, 배경에 지시문에서 허용하지 않은 무장 군인들이 다수 등장함.",
    "hard_violations": [
     "[gemini-pro] 지시문에 없는 인물(배경의 무장 군인들) 추가",
     "[gemini-pro] 클로즈업 프레이밍 위반 및 모서리에 숨는 설정 불일치",
     "[gpt-high] 현우와 페드로만 등장해야 하는 장면에 무장한 수색 인물 최소 네 명의 몸을 추가했다."
    ],
    "physics": "두 사람은 바닥에 앉아 체중을 땅과 벽에 싣고 있으며, 배경의 인물들은 바닥을 딛고 걷고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 프레이밍과 컨테이너 모서리 뒤에 숨는 연출을 정확히 구현하였고, 추가된 인물 없이 요구사항을 충실히 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프레이밍이 클로즈업이 아니며, 지시문에 없는 무장 인물들을 배경에 임의로 추가하여 엄격한 규정을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 사람 모두 컨테이너 모서리 너머 우측 바깥쪽을 주시하고 있음.",
        "built_space": "좌측에 파란색 컨테이너 모서리가 수직으로 배치되었고, 우측 뒤편으로 야외 통로와 다른 컨테이너들이 원근감 있게 보임.",
        "entities": "얼굴 상처와 회색 셔츠를 착용한 현우, 비니와 짙은 후드를 입은 페드로 모두 레퍼런스의 외형과 일치함. 지시되지 않은 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "두 사람 모두 땅을 딛고 서서 컨테이너에 몸을 밀착시키고 있으며, 현우의 손이 모서리를 자연스럽게 짚고 지탱함."
       },
       {
        "label": "B",
        "direction": "두 사람 모두 프레임 우측 배경 쪽으로 시선을 향하고 있음.",
        "built_space": "파란색 컨테이너 벽면 앞에 배치되었으나 모서리에 숨어 있는 형태가 아니며, 우측으로 통로가 보임.",
        "entities": "현우와 페드로의 상반신 외형은 레퍼런스와 일치하나, 배경에 지시문에서 허용하지 않은 무장 군인들이 다수 등장함.",
        "hard_violations": [
         "지시문에 없는 인물(배경의 무장 군인들) 추가",
         "클로즈업 프레이밍 위반 및 모서리에 숨는 설정 불일치"
        ],
        "physics": "두 사람은 바닥에 앉아 체중을 땅과 벽에 싣고 있으며, 배경의 인물들은 바닥을 딛고 걷고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 프레이밍과 컨테이너 모서리 뒤에 숨는 연출을 정확히 구현하였고, 추가된 인물 없이 요구사항을 충실히 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프레이밍이 클로즈업이 아니며, 지시문에 없는 무장 인물들을 배경에 임의로 추가하여 엄격한 규정을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 사람 모두 컨테이너 모서리 너머 우측 바깥쪽을 주시하고 있음.",
        "built_space": "좌측에 파란색 컨테이너 모서리가 수직으로 배치되었고, 우측 뒤편으로 야외 통로와 다른 컨테이너들이 원근감 있게 보임.",
        "entities": "얼굴 상처와 회색 셔츠를 착용한 현우, 비니와 짙은 후드를 입은 페드로 모두 레퍼런스의 외형과 일치함. 지시되지 않은 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "두 사람 모두 땅을 딛고 서서 컨테이너에 몸을 밀착시키고 있으며, 현우의 손이 모서리를 자연스럽게 짚고 지탱함."
       },
       {
        "label": "B",
        "direction": "두 사람 모두 프레임 우측 배경 쪽으로 시선을 향하고 있음.",
        "built_space": "파란색 컨테이너 벽면 앞에 배치되었으나 모서리에 숨어 있는 형태가 아니며, 우측으로 통로가 보임.",
        "entities": "현우와 페드로의 상반신 외형은 레퍼런스와 일치하나, 배경에 지시문에서 허용하지 않은 무장 군인들이 다수 등장함.",
        "hard_violations": [
         "지시문에 없는 인물(배경의 무장 군인들) 추가",
         "클로즈업 프레이밍 위반 및 모서리에 숨는 설정 불일치"
        ],
        "physics": "두 사람은 바닥에 앉아 체중을 땅과 벽에 싣고 있으며, 배경의 인물들은 바닥을 딛고 걷고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 수색 인물들을 추가했으며, 얼굴 클로즈업 대신 상체와 무릎까지 넓혀 보여 주어 핵심 구도도 어겼다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "모퉁이에 밀착한 두 사람의 놀란 얼굴과 모퉁이 너머를 향한 시선을 클로즈업으로 구현했지만, 배경 골목 바닥은 참조의 포장도로와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 페드로 모두 화면 오른쪽 골목의 수색 인물들 쪽을 바라본다. 배경 인물들은 대체로 카메라에서 멀어지는 방향으로 걷고 있으며, 흐릿한 총기들은 아래쪽을 향하고 두 주인공을 겨냥하지 않는다.",
        "built_space": "녹슨 청회색 컨테이너 벽 하나가 왼쪽과 중앙을 차지하고, 오른쪽의 수직 모서리 하나 너머로 포장 골목이 보인다. 두 사람은 벽에 등을 붙이고 낮게 앉아 있다. 다만 현우의 굽힌 무릎이 모서리 바깥 골목 쪽으로 돌출되어 완전히 숨은 배치가 약해진다. 문이나 창문 등 고정 설비는 이 구도에서 확인되지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성 외모, 검은 머리, 회색 겉셔츠와 얼굴의 생채기를 갖추었다. 페드로는 참조와 유사한 젊은 남성 얼굴에 검은 비니와 짙은 후드티를 착용했다. 현우의 드러난 다리에는 붕대와 상처가 보인다. 그러나 골목에는 두 주인공 외에 최소 네 명의 무장 수색 인물이 부분적으로 나타난다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우와 페드로만 등장해야 하는 장면에 무장한 수색 인물 최소 네 명의 몸을 추가했다."
        ],
        "physics": "두 사람은 몸을 낮춰 벽에 기대고 있으며 하체가 화면 아래로 이어진다. 엉덩이의 바닥 접촉은 잘렸지만 공중에 떠 있는 모습은 아니다. 현우의 무릎 굽힘도 가능한 자세다. 배경 보행자들은 지면에 디딘 발과 들어 올린 발이 보여 보행의 지지가 성립한다."
       },
       {
        "label": "B",
        "direction": "현우와 페드로의 눈은 모두 화면 왼쪽, 컨테이너 모서리 너머의 화면 밖 공간을 향한다. 수색 대상이나 수색자는 직접 보이지 않지만 같은 방향을 경계하는 관계가 분명하며, 렌즈를 응시하지 않는다.",
        "built_space": "왼쪽의 녹슨 청회색 컨테이너 벽이 좁은 수직 모서리 하나에서 끝나고, 그 오른쪽에 현우가 앞쪽, 페드로가 바로 뒤쪽으로 밀착해 있다. 두 얼굴은 식별 가능하게 드러나며 모서리가 은폐물로 읽힌다. 오른쪽 배경에는 컨테이너 외벽과 출입구가 이어지지만, 골목이 참조의 아스팔트 대신 흙길처럼 보여 장소 연속성이 약하다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "등장인물은 현우와 페드로 두 명뿐이다. 현우의 앳된 동아시아계 얼굴, 헝클어진 검은 머리, 회색 겉셔츠와 아직 아물지 않은 얼굴 상처가 보인다. 페드로는 참조와 유사한 젊은 라틴계 혼혈 남성 외모에 검은 비니와 짙은 후드티를 유지한다. 두 사람의 눈은 정상적인 인체 비례 안에서 긴장과 놀람을 표현한다. 다리와 하의는 클로즈업 밖이며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 머리와 어깨를 모서리에 바짝 붙이고, 화면 아래의 손으로 금속 가장자리를 잡고 있다. 페드로는 그 뒤에서 상체를 조금 기울인다. 두 몸은 화면 아래로 자연스럽게 이어지며, 발의 접지는 구도 밖이지만 도약하거나 떠 있는 형상은 아니다. 손과 모서리의 접촉도 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 수색 인물들을 추가했으며, 얼굴 클로즈업 대신 상체와 무릎까지 넓혀 보여 주어 핵심 구도도 어겼다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "모퉁이에 밀착한 두 사람의 놀란 얼굴과 모퉁이 너머를 향한 시선을 클로즈업으로 구현했지만, 배경 골목 바닥은 참조의 포장도로와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 페드로 모두 화면 오른쪽 골목의 수색 인물들 쪽을 바라본다. 배경 인물들은 대체로 카메라에서 멀어지는 방향으로 걷고 있으며, 흐릿한 총기들은 아래쪽을 향하고 두 주인공을 겨냥하지 않는다.",
        "built_space": "녹슨 청회색 컨테이너 벽 하나가 왼쪽과 중앙을 차지하고, 오른쪽의 수직 모서리 하나 너머로 포장 골목이 보인다. 두 사람은 벽에 등을 붙이고 낮게 앉아 있다. 다만 현우의 굽힌 무릎이 모서리 바깥 골목 쪽으로 돌출되어 완전히 숨은 배치가 약해진다. 문이나 창문 등 고정 설비는 이 구도에서 확인되지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성 외모, 검은 머리, 회색 겉셔츠와 얼굴의 생채기를 갖추었다. 페드로는 참조와 유사한 젊은 남성 얼굴에 검은 비니와 짙은 후드티를 착용했다. 현우의 드러난 다리에는 붕대와 상처가 보인다. 그러나 골목에는 두 주인공 외에 최소 네 명의 무장 수색 인물이 부분적으로 나타난다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우와 페드로만 등장해야 하는 장면에 무장한 수색 인물 최소 네 명의 몸을 추가했다."
        ],
        "physics": "두 사람은 몸을 낮춰 벽에 기대고 있으며 하체가 화면 아래로 이어진다. 엉덩이의 바닥 접촉은 잘렸지만 공중에 떠 있는 모습은 아니다. 현우의 무릎 굽힘도 가능한 자세다. 배경 보행자들은 지면에 디딘 발과 들어 올린 발이 보여 보행의 지지가 성립한다."
       },
       {
        "label": "A",
        "direction": "현우와 페드로의 눈은 모두 화면 왼쪽, 컨테이너 모서리 너머의 화면 밖 공간을 향한다. 수색 대상이나 수색자는 직접 보이지 않지만 같은 방향을 경계하는 관계가 분명하며, 렌즈를 응시하지 않는다.",
        "built_space": "왼쪽의 녹슨 청회색 컨테이너 벽이 좁은 수직 모서리 하나에서 끝나고, 그 오른쪽에 현우가 앞쪽, 페드로가 바로 뒤쪽으로 밀착해 있다. 두 얼굴은 식별 가능하게 드러나며 모서리가 은폐물로 읽힌다. 오른쪽 배경에는 컨테이너 외벽과 출입구가 이어지지만, 골목이 참조의 아스팔트 대신 흙길처럼 보여 장소 연속성이 약하다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "등장인물은 현우와 페드로 두 명뿐이다. 현우의 앳된 동아시아계 얼굴, 헝클어진 검은 머리, 회색 겉셔츠와 아직 아물지 않은 얼굴 상처가 보인다. 페드로는 참조와 유사한 젊은 라틴계 혼혈 남성 외모에 검은 비니와 짙은 후드티를 유지한다. 두 사람의 눈은 정상적인 인체 비례 안에서 긴장과 놀람을 표현한다. 다리와 하의는 클로즈업 밖이며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 머리와 어깨를 모서리에 바짝 붙이고, 화면 아래의 손으로 금속 가장자리를 잡고 있다. 페드로는 그 뒤에서 상체를 조금 기울인다. 두 몸은 화면 아래로 자연스럽게 이어지며, 발의 접지는 구도 밖이지만 도약하거나 떠 있는 형상은 아니다. 손과 모서리의 접촉도 자연스럽다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 지시문에 없는 인물(배경의 무장 군인들) 추가",
     "[gemini-pro] 클로즈업 프레이밍 위반 및 모서리에 숨는 설정 불일치",
     "[gpt-high] 현우와 페드로만 등장해야 하는 장면에 무장한 수색 인물 최소 네 명의 몸을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 프레이밍과 컨테이너 모서리 뒤에 숨는 연출을 정확히 구현하였고, 추가된 인물 없이 요구사항을 충실히 따름."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "프레이밍이 클로즈업이 아니며, 지시문에 없는 무장 인물들을 배경에 임의로 추가하여 엄격한 규정을 위반함.  ★위반: [gemini-pro] 지시문에 없는 인물(배경의 무장 군인들) 추가 / [gemini-pro] 클로즈업 프레이밍 위반 및 모서리에 숨는 설정 불일치 / [gpt-high] 현우와 페드로만 등장해야 하는 장면에 무장한 수색 인물 최소 네 명의 몸을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4_sel.png",
    "asset_id": "c8038c6b-9a89-4367-b728-42df6425b41a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-e9f0-7270-9072-191b9c5d628b",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S11sh4"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S20sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:45:21.193587+00:00",
  "fingerprint": "25c699c15a2fda3f442fd7bfba4d35baa0a9ab5c9f7e6cb6d2d9e6ff0c5aedaa",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S20sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S20sh3_sel.png",
  "source_sha256": "ddbcf553c2f37ed4d04f8c4972311ef1cde4d6c05cc77bb21fad487d6b59d76c",
  "file": "S20sh3_cine.png",
  "staged_sha256": "e69db4c5cfb51a9f5196b94b2f29a3bd2d63f611ff5114f302479a57a3291cdb",
  "latency_ms": 14985
 },
 "S20sh6::signage": {
  "fp": "9c8a18a1fb831a09",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S20sh6": {
  "input_fingerprint": "8eacabd84e42a4de",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손이 경찰(현우 체포·호송)의 가슴에 닿아 거칠게 뒤로 밀어내는 힘이 실린 페드로의 미드액션 자세.\n\nLOCATION (lock): On the outdoor container-lined lane near the searched home, where police block the route toward the settlement market. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refugee-settlement passage (The officer blocks the route as the shove begins); used as Leave a modest visible strip around the bodies so the backward displacement reads within a real passage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral daytime illumination and restrained tonal contrast as the concealment beat, with no dramatic lighting shift at contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 remains the site of the militia search. 페드로: He is at the blocked escape route, with his arms extended in a forceful push.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손이 경찰(현우 체포·호송)의 가슴에 닿아 거칠게 뒤로 밀어내는 힘이 실린 페드로의 미드액션 자세.\n\nLOCATION (lock): On the outdoor container-lined lane near the searched home, where police block the route toward the settlement market. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refugee-settlement passage (The officer blocks the route as the shove begins); used as Leave a modest visible strip around the bodies so the backward displacement reads within a real passage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral daytime illumination and restrained tonal contrast as the concealment beat, with no dramatic lighting shift at contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 remains the site of the militia search. 페드로: He is at the blocked escape route, with his arms extended in a forceful push.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손이 경찰(현우 체포·호송)의 가슴에 닿아 거칠게 뒤로 밀어내는 힘이 실린 페드로의 미드액션 자세.\n\nLOCATION (lock): On the outdoor container-lined lane near the searched home, where police block the route toward the settlement market. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refugee-settlement passage (The officer blocks the route as the shove begins); used as Leave a modest visible strip around the bodies so the backward displacement reads within a real passage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral daytime illumination and restrained tonal contrast as the concealment beat, with no dramatic lighting shift at contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 remains the site of the militia search. 페드로: He is at the blocked escape route, with his arms extended in a forceful push.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "페드로의 시선과 양손이 목표 대상인 경찰관의 가슴을 향해 정확히 조준되어 밀쳐내고 있음.",
    "built_space": "야외 컨테이너 통로. 파란색과 흰색 컨테이너가 양옆에 늘어서 있고 인물들이 중앙에 적절히 위치함.",
    "entities": "페드로와 경찰관 1명이 등장함. 단, 경찰관이 참조 이미지 속 다른 인물(현우)의 얼굴과 분장을 그대로 모방함.",
    "hard_violations": [
     "[gpt-high] 이전 컷에서 페드로 외 인물의 얼굴·외형을 이어 쓰지 말라는 명시적 금지에도, 제외 대상 청년의 얼굴과 상처를 경찰에게 옮겼다."
    ],
    "physics": "페드로가 런지 자세로 지면을 단단히 딛고 밀고 있으며, 경찰관은 뒤로 체중이 쏠린 채 밀려나는 물리적 균형이 안정적임."
   },
   {
    "label": "B",
    "direction": "페드로의 양손과 시선이 경찰관을 향해 밀어내고 있음.",
    "built_space": "컨테이너가 늘어선 통로 공간으로 배경 구조물들이 늘어서 있음.",
    "entities": "전면에 페드로와 경찰관이 있으나, 지시되지 않은 다수의 군중(경찰 3명, 민간인 2명)이 배경에 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 지시되지 않은 다수의 인물(경찰, 민간인)이 배경에 임의로 추가됨 (invented people)",
     "[gpt-high] 샷 텍스트가 보여 주지 않는 배경 인물 5명을 추가했다."
    ],
    "physics": "경찰관이 뒤로 기울어진 각도에 비해 하체가 지나치게 뻣뻣하게 멈춰 있으나, 양측 모두 지면에 발을 딛고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "요구된 미디엄 샷과 밀치는 액션은 자연스럽게 구현했으나, 경찰관의 얼굴에 제외되어야 할 참조 이미지 속 인물(현우)의 외모와 상처를 그대로 적용한 치명적인 신원 오류가 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 다수의 배경 인물(경찰, 민간인)을 임의로 추가하여 하드 위반에 해당하며, 페드로의 소품(체인) 방향이 뒤바뀌었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 시선과 양손이 목표 대상인 경찰관의 가슴을 향해 정확히 조준되어 밀쳐내고 있음.",
        "built_space": "야외 컨테이너 통로. 파란색과 흰색 컨테이너가 양옆에 늘어서 있고 인물들이 중앙에 적절히 위치함.",
        "entities": "페드로와 경찰관 1명이 등장함. 단, 경찰관이 참조 이미지 속 다른 인물(현우)의 얼굴과 분장을 그대로 모방함.",
        "hard_violations": [],
        "physics": "페드로가 런지 자세로 지면을 단단히 딛고 밀고 있으며, 경찰관은 뒤로 체중이 쏠린 채 밀려나는 물리적 균형이 안정적임."
       },
       {
        "label": "B",
        "direction": "페드로의 양손과 시선이 경찰관을 향해 밀어내고 있음.",
        "built_space": "컨테이너가 늘어선 통로 공간으로 배경 구조물들이 늘어서 있음.",
        "entities": "전면에 페드로와 경찰관이 있으나, 지시되지 않은 다수의 군중(경찰 3명, 민간인 2명)이 배경에 존재함.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 다수의 인물(경찰, 민간인)이 배경에 임의로 추가됨 (invented people)"
        ],
        "physics": "경찰관이 뒤로 기울어진 각도에 비해 하체가 지나치게 뻣뻣하게 멈춰 있으나, 양측 모두 지면에 발을 딛고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "요구된 미디엄 샷과 밀치는 액션은 자연스럽게 구현했으나, 경찰관의 얼굴에 제외되어야 할 참조 이미지 속 인물(현우)의 외모와 상처를 그대로 적용한 치명적인 신원 오류가 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 다수의 배경 인물(경찰, 민간인)을 임의로 추가하여 하드 위반에 해당하며, 페드로의 소품(체인) 방향이 뒤바뀌었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 시선과 양손이 목표 대상인 경찰관의 가슴을 향해 정확히 조준되어 밀쳐내고 있음.",
        "built_space": "야외 컨테이너 통로. 파란색과 흰색 컨테이너가 양옆에 늘어서 있고 인물들이 중앙에 적절히 위치함.",
        "entities": "페드로와 경찰관 1명이 등장함. 단, 경찰관이 참조 이미지 속 다른 인물(현우)의 얼굴과 분장을 그대로 모방함.",
        "hard_violations": [],
        "physics": "페드로가 런지 자세로 지면을 단단히 딛고 밀고 있으며, 경찰관은 뒤로 체중이 쏠린 채 밀려나는 물리적 균형이 안정적임."
       },
       {
        "label": "B",
        "direction": "페드로의 양손과 시선이 경찰관을 향해 밀어내고 있음.",
        "built_space": "컨테이너가 늘어선 통로 공간으로 배경 구조물들이 늘어서 있음.",
        "entities": "전면에 페드로와 경찰관이 있으나, 지시되지 않은 다수의 군중(경찰 3명, 민간인 2명)이 배경에 존재함.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 다수의 인물(경찰, 민간인)이 배경에 임의로 추가됨 (invented people)"
        ],
        "physics": "경찰관이 뒤로 기울어진 각도에 비해 하체가 지나치게 뻣뻣하게 멈춰 있으나, 양측 모두 지면에 발을 딛고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "양손으로 경찰의 가슴을 미는 동작은 명확하지만, 허용되지 않은 배경 인물 5명을 추가했고 미디엄 숏을 전신 숏으로 넓혔다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두 사람의 접촉과 후퇴는 잘 읽히고 군중도 없지만, 경찰에게 이전 컷의 제외 대상 인물 외형을 옮겼으며 구도도 미디엄 숏보다 지나치게 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 페드로는 왼쪽 경찰의 얼굴을 보며 양팔을 뻗는다. 두 손바닥은 경찰의 상부 가슴과 방탄조끼 앞면에 닿고, 경찰의 몸통은 화면 왼쪽 뒤로 기울어 밀리는 방향이 맞는다. 경찰도 페드로 쪽으로 얼굴을 돌린다. 겨누는 무기는 없다.",
        "built_space": "녹슨 청색 컨테이너가 왼쪽, 흰색·연녹색 컨테이너가 오른쪽에 놓인 흙길이다. 왼쪽 가까운 벽에는 흰 창틀 창 하나가 뚜렷하고, 오른쪽 녹색 벽에는 창 두 개가 보인다. 두 사람은 벽 사이 통로에 서 있으며 뒤로 물러날 공간이 있다. 재질과 낮의 조명은 참조와 대체로 맞지만, 참조보다 양옆 건물과 바닥을 훨씬 많이 보여 준다. 두 사람의 머리부터 신발까지 담아 지정된 미디엄 숏이 아니다.",
        "entities": "주동작 인물은 페드로와 한국인 성인 남성으로 보이는 경찰이다. 페드로의 앳된 얼굴, 비니 아래 짙은 머리, 검은 비니, 낡은 짙은 후드와 청바지, 손목 장식 및 허리 체인은 참조에 대체로 부합한다. 경찰은 제복과 방탄조끼를 착용한다. 그러나 뒤에 제복 인물 두 명과 사복 인물 세 명이 추가되어 총 7명이 보인다. 읽을 수 있는 컨테이너 번호는 없다.",
        "hard_violations": [
         "샷 텍스트가 보여 주지 않는 배경 인물 5명을 추가했다."
        ],
        "physics": "페드로는 앞발을 땅에 붙이고 뒷발 앞부분으로 버티면서 상체의 힘을 뻗은 양팔에 싣는다. 경찰은 한쪽 발로 지면을 지지하고 다른 발을 들어 균형을 회복하려는 모습이다. 뒤로 기운 몸통은 가슴에 닿은 두 손의 밀기와 연결되며, 공중에 지지 없이 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "페드로는 왼쪽 경찰의 얼굴을 응시하고 두 손을 경찰의 윗가슴에 대고 민다. 경찰은 페드로를 향해 얼굴을 돌린 채 상체가 왼쪽 뒤로 밀린다. 팔의 힘과 몸통의 후퇴 방향이 서로 일치한다. 겨누는 무기는 없다.",
        "built_space": "왼쪽의 녹슨 청색 컨테이너, 오른쪽의 밝은 컨테이너 열, 뒤쪽을 가로지르는 청색 컨테이너와 흙·자갈 통로가 보인다. 왼쪽 전경에는 세로 창살이 있는 창 하나, 뒤쪽에는 가로등 기둥 하나가 명확하다. 두 사람은 통로 안에서 서로 마주 서고 경찰의 후퇴 공간도 확보된다. 재질과 중립적인 주간 조명은 참조와 잘 이어지지만, 발 부근까지 보여 주는 거의 전신 구도여서 요구한 미디엄 숏보다 넓다.",
        "entities": "페드로의 젊은 얼굴, 짙은 머리, 검은 비니와 낡은 짙은 후드, 청바지 및 손목 장식은 참조와 대체로 일치한다. 상대는 제복과 방탄조끼를 입은 한국인 청년 남성으로 보인다. 다만 상대의 검은 머리 형태와 얼굴, 뺨의 상처가 이전 컷에서 페드로 앞에 있던 제외 대상 청년의 외형을 재사용한 것으로 보인다. 그 외 뚜렷한 추가 인물은 없다. 읽을 수 있는 장소 번호는 보이지 않는다.",
        "hard_violations": [
         "이전 컷에서 페드로 외 인물의 얼굴·외형을 이어 쓰지 말라는 명시적 금지에도, 제외 대상 청년의 얼굴과 상처를 경찰에게 옮겼다."
        ],
        "physics": "페드로의 앞다리는 하단 프레임 밖 지면으로 이어지고 뒷발은 발끝으로 지면을 밀며 상체를 앞으로 기울인다. 경찰은 뒤쪽 발로 바닥을 짚고 다리를 벌려 밀림을 받아낸다. 양손과 가슴의 접촉, 팔의 신전, 경찰의 뒤로 기운 몸이 하나의 밀기 동작으로 연결된다. 발 일부의 크롭을 제외하면 지지 관계가 자연스럽고, 지지 없는 부유는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "양손으로 경찰의 가슴을 미는 동작은 명확하지만, 허용되지 않은 배경 인물 5명을 추가했고 미디엄 숏을 전신 숏으로 넓혔다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 사람의 접촉과 후퇴는 잘 읽히고 군중도 없지만, 경찰에게 이전 컷의 제외 대상 인물 외형을 옮겼으며 구도도 미디엄 숏보다 지나치게 넓다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 페드로는 왼쪽 경찰의 얼굴을 보며 양팔을 뻗는다. 두 손바닥은 경찰의 상부 가슴과 방탄조끼 앞면에 닿고, 경찰의 몸통은 화면 왼쪽 뒤로 기울어 밀리는 방향이 맞는다. 경찰도 페드로 쪽으로 얼굴을 돌린다. 겨누는 무기는 없다.",
        "built_space": "녹슨 청색 컨테이너가 왼쪽, 흰색·연녹색 컨테이너가 오른쪽에 놓인 흙길이다. 왼쪽 가까운 벽에는 흰 창틀 창 하나가 뚜렷하고, 오른쪽 녹색 벽에는 창 두 개가 보인다. 두 사람은 벽 사이 통로에 서 있으며 뒤로 물러날 공간이 있다. 재질과 낮의 조명은 참조와 대체로 맞지만, 참조보다 양옆 건물과 바닥을 훨씬 많이 보여 준다. 두 사람의 머리부터 신발까지 담아 지정된 미디엄 숏이 아니다.",
        "entities": "주동작 인물은 페드로와 한국인 성인 남성으로 보이는 경찰이다. 페드로의 앳된 얼굴, 비니 아래 짙은 머리, 검은 비니, 낡은 짙은 후드와 청바지, 손목 장식 및 허리 체인은 참조에 대체로 부합한다. 경찰은 제복과 방탄조끼를 착용한다. 그러나 뒤에 제복 인물 두 명과 사복 인물 세 명이 추가되어 총 7명이 보인다. 읽을 수 있는 컨테이너 번호는 없다.",
        "hard_violations": [
         "샷 텍스트가 보여 주지 않는 배경 인물 5명을 추가했다."
        ],
        "physics": "페드로는 앞발을 땅에 붙이고 뒷발 앞부분으로 버티면서 상체의 힘을 뻗은 양팔에 싣는다. 경찰은 한쪽 발로 지면을 지지하고 다른 발을 들어 균형을 회복하려는 모습이다. 뒤로 기운 몸통은 가슴에 닿은 두 손의 밀기와 연결되며, 공중에 지지 없이 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "페드로는 왼쪽 경찰의 얼굴을 응시하고 두 손을 경찰의 윗가슴에 대고 민다. 경찰은 페드로를 향해 얼굴을 돌린 채 상체가 왼쪽 뒤로 밀린다. 팔의 힘과 몸통의 후퇴 방향이 서로 일치한다. 겨누는 무기는 없다.",
        "built_space": "왼쪽의 녹슨 청색 컨테이너, 오른쪽의 밝은 컨테이너 열, 뒤쪽을 가로지르는 청색 컨테이너와 흙·자갈 통로가 보인다. 왼쪽 전경에는 세로 창살이 있는 창 하나, 뒤쪽에는 가로등 기둥 하나가 명확하다. 두 사람은 통로 안에서 서로 마주 서고 경찰의 후퇴 공간도 확보된다. 재질과 중립적인 주간 조명은 참조와 잘 이어지지만, 발 부근까지 보여 주는 거의 전신 구도여서 요구한 미디엄 숏보다 넓다.",
        "entities": "페드로의 젊은 얼굴, 짙은 머리, 검은 비니와 낡은 짙은 후드, 청바지 및 손목 장식은 참조와 대체로 일치한다. 상대는 제복과 방탄조끼를 입은 한국인 청년 남성으로 보인다. 다만 상대의 검은 머리 형태와 얼굴, 뺨의 상처가 이전 컷에서 페드로 앞에 있던 제외 대상 청년의 외형을 재사용한 것으로 보인다. 그 외 뚜렷한 추가 인물은 없다. 읽을 수 있는 장소 번호는 보이지 않는다.",
        "hard_violations": [
         "이전 컷에서 페드로 외 인물의 얼굴·외형을 이어 쓰지 말라는 명시적 금지에도, 제외 대상 청년의 얼굴과 상처를 경찰에게 옮겼다."
        ],
        "physics": "페드로의 앞다리는 하단 프레임 밖 지면으로 이어지고 뒷발은 발끝으로 지면을 밀며 상체를 앞으로 기울인다. 경찰은 뒤쪽 발로 바닥을 짚고 다리를 벌려 밀림을 받아낸다. 양손과 가슴의 접촉, 팔의 신전, 경찰의 뒤로 기운 몸이 하나의 밀기 동작으로 연결된다. 발 일부의 크롭을 제외하면 지지 관계가 자연스럽고, 지지 없는 부유는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.267
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.017
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 지시되지 않은 다수의 인물(경찰, 민간인)이 배경에 임의로 추가됨 (invented people)",
     "[gpt-high] 샷 텍스트가 보여 주지 않는 배경 인물 5명을 추가했다."
    ],
    "A": [
     "[gpt-high] 이전 컷에서 페드로 외 인물의 얼굴·외형을 이어 쓰지 말라는 명시적 금지에도, 제외 대상 청년의 얼굴과 상처를 경찰에게 옮겼다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1017
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "요구된 미디엄 샷과 밀치는 액션은 자연스럽게 구현했으나, 경찰관의 얼굴에 제외되어야 할 참조 이미지 속 인물(현우)의 외모와 상처를 그대로 적용한 치명적인 신원 오류가 있습니다.  ★위반: [gpt-high] 이전 컷에서 페드로 외 인물의 얼굴·외형을 이어 쓰지 말라는 명시적 금지에도, 제외 대상 청년의 얼굴과 상처를 경찰에게 옮겼다."
   },
   {
    "label": "B",
    "score": 1017,
    "verdict_ko": "프롬프트에 없는 다수의 배경 인물(경찰, 민간인)을 임의로 추가하여 하드 위반에 해당하며, 페드로의 소품(체인) 방향이 뒤바뀌었습니다.  ★위반: [gemini-pro] 프롬프트에 지시되지 않은 다수의 인물(경찰, 민간인)이 배경에 임의로 추가됨 (invented people) / [gpt-high] 샷 텍스트가 보여 주지 않는 배경 인물 5명을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh3_sel.png",
    "asset_id": "668f4afe-dea0-422a-bab5-5cdc76b1640c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-ebac-73eb-a081-155769b428ea",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S20sh3"
  }
 },
 "S20sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:46:45.124340+00:00",
  "fingerprint": "5ff872d8bac6d6661d980bdf3abbca159535d52b3ddba1be585b86d3ef29484b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S20sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S20sh6_sel.png",
  "source_sha256": "c172a85989425f3bd212cfd185f0b24a473807c5d0dbb325809cad539609ff31",
  "file": "S20sh6_cine.png",
  "staged_sha256": "c211685d503f41449d59a4d79699dbea4d736931d2ef09394640895ebfc0904c",
  "latency_ms": 9339
 },
 "S20sh16::signage": {
  "fp": "aaf6824d4f8ea964",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::07e119b2d4378bdf": {
  "subjects": [],
  "subject_text": "인천 난민촌 시장과 상점 골목\n컨테이너 거리 사이로 작은 상점과 상품 진열대가 이어지는 시장 골목. 바나나를 파는 가게와 상자 더미, 더욱 좁아지는 샛길이 있다.",
  "identity": "canonical",
  "scope_id": "L167",
  "scope_role": "location_exterior",
  "scope_sha": "eced2ac782880df0"
 },
 "groupbg::market_capture_alley": {
  "input_fingerprint": "1603bdda72a77d08",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "market_capture_alley",
    "tags": [
     "S20sh16"
    ]
   },
   "context_sig": "502f598537ff5204"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 더욱 좁은 골목길로 도망가려던 찰나-! 퍽!!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 더욱 좁은 골목길로 도망가려던 찰나-! 퍽!!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_market_capture_alley_13d54d.png",
  "asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47",
  "input_asset_ids": [
   "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd"
  ],
  "origin_tag": "S20sh16",
  "place_text": "On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.",
  "origin_inputs": {
   "place_text": "On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.",
   "time_of_day_en": "day",
   "conti_asset_id": "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd"
  }
 },
 "S20sh16::bgfirst_bg": {
  "input_fingerprint": "74ed9dc23d6bc9bc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh16__bgfirst_bg.png",
  "asset_id": "43c81765-5840-45d5-8f30-f37e76cc5841",
  "input_asset_ids": [
   "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd",
   "5e0698a7-8a58-4159-968f-d06cd12bcc47"
  ]
 },
 "S20sh16": {
  "input_fingerprint": "e7533cf439f6895b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is pinned face-down on the ground beneath a plainclothes officer's knee pressing into his back. His head's direction and the positions of his arms and legs are not specified.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boxes knocked over during the escape remain scattered through the market alley. 현우: He has been brought down by a rifle-butt strike, with his outer shirt still on. His earlier facial injuries and dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is pinned face-down on the ground beneath a plainclothes officer's knee pressing into his back. His head's direction and the positions of his arms and legs are not specified.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boxes knocked over during the escape remain scattered through the market alley. 현우: He has been brought down by a rifle-butt strike, with his outer shirt still on. His earlier facial injuries and dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is pinned face-down on the ground beneath a plainclothes officer's knee pressing into his back. His head's direction and the positions of his arms and legs are not specified.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boxes knocked over during the escape remain scattered through the market alley. 현우: He has been brought down by a rifle-butt strike, with his outer shirt still on. His earlier facial injuries and dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh16__bgfirst_bg.png",
     "asset_id": "43c81765-5840-45d5-8f30-f37e76cc5841",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S20sh16.png",
     "asset_id": "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_market_capture_alley_13d54d.png",
     "asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "경찰은 바닥에 엎드린 현우를 시선으로 향하며 내려다보고 있음.",
    "built_space": "참조된 시장 골목길 배경이 정확히 구현되었으며 화면 좌측에 쓰러진 상자들이 배치됨.",
    "entities": "현우는 참조 이미지의 얼굴, 체형, 의상을 잘 유지하고 있으며 얼굴에 상처가 있음. 사복 차림의 경찰이 묘사됨.",
    "hard_violations": [],
    "physics": "현우는 지면에 완전히 엎드려 지탱받고 있으며, 경찰은 왼쪽 무릎을 바닥에 대고 오른쪽 무릎으로 현우의 등을 누르며 체중을 싣는 안정적인 자세를 취함."
   },
   {
    "label": "B",
    "direction": "경찰의 시선이 바닥에 제압된 현우를 향하고 있음.",
    "built_space": "참조 이미지와 일치하는 골목길 구조이며, 좌측에 상자들이 흩어져 있음.",
    "entities": "현우의 인상착의와 의상이 참조 이미지와 일치하며 사복 경찰도 명확히 식별됨.",
    "hard_violations": [],
    "physics": "현우는 바닥에 지탱되어 있으나, 경찰은 두 무릎을 모두 지면에 대고 양손으로 현우의 등을 누르고 있어 '무릎으로 짓누르는' 지정된 물리적 동작이 나타나지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '무릎으로 등 위를 짓누르는' 핵심 제압 자세를 정확히 연출했으며, 지정된 배경과 인물의 외형 조건을 훌륭하게 준수함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경과 인물의 착장은 적절하나, 경찰이 무릎이 아닌 양손으로 등을 누르고 있어 프롬프트가 지정한 가장 중요한 물리적 접촉 자세를 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "경찰은 바닥에 엎드린 현우를 시선으로 향하며 내려다보고 있음.",
        "built_space": "참조된 시장 골목길 배경이 정확히 구현되었으며 화면 좌측에 쓰러진 상자들이 배치됨.",
        "entities": "현우는 참조 이미지의 얼굴, 체형, 의상을 잘 유지하고 있으며 얼굴에 상처가 있음. 사복 차림의 경찰이 묘사됨.",
        "hard_violations": [],
        "physics": "현우는 지면에 완전히 엎드려 지탱받고 있으며, 경찰은 왼쪽 무릎을 바닥에 대고 오른쪽 무릎으로 현우의 등을 누르며 체중을 싣는 안정적인 자세를 취함."
       },
       {
        "label": "B",
        "direction": "경찰의 시선이 바닥에 제압된 현우를 향하고 있음.",
        "built_space": "참조 이미지와 일치하는 골목길 구조이며, 좌측에 상자들이 흩어져 있음.",
        "entities": "현우의 인상착의와 의상이 참조 이미지와 일치하며 사복 경찰도 명확히 식별됨.",
        "hard_violations": [],
        "physics": "현우는 바닥에 지탱되어 있으나, 경찰은 두 무릎을 모두 지면에 대고 양손으로 현우의 등을 누르고 있어 '무릎으로 짓누르는' 지정된 물리적 동작이 나타나지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '무릎으로 등 위를 짓누르는' 핵심 제압 자세를 정확히 연출했으며, 지정된 배경과 인물의 외형 조건을 훌륭하게 준수함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경과 인물의 착장은 적절하나, 경찰이 무릎이 아닌 양손으로 등을 누르고 있어 프롬프트가 지정한 가장 중요한 물리적 접촉 자세를 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "경찰은 바닥에 엎드린 현우를 시선으로 향하며 내려다보고 있음.",
        "built_space": "참조된 시장 골목길 배경이 정확히 구현되었으며 화면 좌측에 쓰러진 상자들이 배치됨.",
        "entities": "현우는 참조 이미지의 얼굴, 체형, 의상을 잘 유지하고 있으며 얼굴에 상처가 있음. 사복 차림의 경찰이 묘사됨.",
        "hard_violations": [],
        "physics": "현우는 지면에 완전히 엎드려 지탱받고 있으며, 경찰은 왼쪽 무릎을 바닥에 대고 오른쪽 무릎으로 현우의 등을 누르며 체중을 싣는 안정적인 자세를 취함."
       },
       {
        "label": "B",
        "direction": "경찰의 시선이 바닥에 제압된 현우를 향하고 있음.",
        "built_space": "참조 이미지와 일치하는 골목길 구조이며, 좌측에 상자들이 흩어져 있음.",
        "entities": "현우의 인상착의와 의상이 참조 이미지와 일치하며 사복 경찰도 명확히 식별됨.",
        "hard_violations": [],
        "physics": "현우는 바닥에 지탱되어 있으나, 경찰은 두 무릎을 모두 지면에 대고 양손으로 현우의 등을 누르고 있어 '무릎으로 짓누르는' 지정된 물리적 동작이 나타나지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "장소와 현우의 외형은 잘 맞지만, 전신과 골목을 넓게 담아 경찰의 굳은 상체 중심 미디엄 숏에서 벗어나고 무릎이 등을 누르는 접점도 불명확하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "경찰의 굳은 상체를 크게 담으면서 등에 실린 무릎, 지면에 받쳐진 현우, 접점 주변 여백을 명확히 보여 주어 핵심 연출에 가장 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "경찰은 아래쪽 현우의 어깨와 팔 부위를 바라보며 상체를 숙인다. 현우의 얼굴은 화면 왼쪽 앞쪽으로 돌아가 있고 시선은 바닥 가까이 향한다. 무기나 겨냥하는 물체는 보이지 않는다. 경찰의 무릎은 현우의 허리 뒤쪽에 위치하지만, 등 위를 직접 아래로 누르는 방향과 접점은 팔에 가려 명확하지 않다.",
        "built_space": "왼쪽에는 붉은 차양과 여러 지지대, 과일 진열대 및 넘어진 상자들이 있고, 오른쪽에는 낡은 벽, 회색 문 하나, 두드러진 수직 배관 두 줄, 파란 통 하나와 목재 보호대 하나가 보인다. 전경 오른쪽의 사각 맨홀 하나와 갈라진 포장도 장소 참조에 부합한다. 두 사람은 골목 바닥에 놓이며 접점 주변 지면은 충분하지만, 현우의 전신과 넓은 배경을 포함해 요구한 경찰 상체 중심보다 넓게 잡혔다.",
        "entities": "보이는 사람은 현우와 체포 경찰 두 명이다. 현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 회색 겉셔츠, 올리브색 카고 바지, 갈색 신발이 참조와 대체로 맞는다. 얼굴 찰과상과 바지 무릎 부근의 손상도 보이지만 개 물림 자국인지는 확정할 수 없다. 경찰은 검은 사복을 입은 성인 동아시아계 남성이다. 흩어진 상자와 과일이 있으며 뚜렷하게 읽히는 문구나 화면 표식은 없다.",
        "hard_violations": [],
        "physics": "현우의 몸통과 다리는 포장면에 누워 있고 얼굴도 바닥 가까이 닿아 있다. 등 뒤로 꺾인 팔은 자기 몸과 경찰의 손에 의해 지지되는 것으로 보인다. 경찰은 굽힌 다리로 낮은 자세를 유지하고 손으로 현우를 고정한다. 지지 없이 떠 있는 신체는 확인되지 않지만, 체중이 무릎을 통해 현우의 등에 전달되는지는 가림 때문에 분명하지 않다."
       },
       {
        "label": "B",
        "direction": "경찰의 얼굴과 시선은 바로 아래 현우의 등과 머리 쪽을 향한다. 한쪽 무릎은 현우의 등 중앙을 아래로 누르고, 한 손은 등 쪽 셔츠를 잡으며 다른 손은 목 뒤와 어깨 가까이에 놓인다. 현우는 얼굴을 화면 오른쪽으로 돌린 채 뺨을 바닥에 대고 눈을 감고 있다. 무기나 다른 지향성 소품은 없다.",
        "built_space": "왼쪽 붉은 차양과 지지대, 과일 진열대, 넘어진 상자들이 참조의 시장 골목 배치를 유지한다. 오른쪽에는 낡은 벽과 세로 골판 외장, 굵고 가는 수직 배관, 파란 통 하나, 녹색 상자 하나와 비스듬한 붉은 판 하나가 보인다. 문 구역 일부는 경찰에게 가려져 있다. 경찰 상체가 화면의 큰 부분을 차지하며, 무릎과 등 및 현우의 얼굴·팔 주변에 바닥 여백이 있어 제압 구조를 읽기 쉽다.",
        "entities": "현우와 사복 경찰 두 명만 보인다. 현우의 젊은 동아시아계 남성 외형, 검은 머리, 회색 겉셔츠와 카고 바지는 참조와 대체로 일치하며, 얼굴에는 기존 찰과상으로 읽히는 상처가 있다. 얼굴은 눌리고 일부 가려져 세밀한 동일인 비교에는 제한이 있다. 다리의 개 물림 상처는 이 시점에서 분명하게 드러나지 않는다. 경찰은 어두운 재킷과 체크 셔츠를 입은 성인 동아시아계 남성이다. 흩어진 상자들이 유지되며 읽을 수 있는 글이나 그래픽 표식은 없다.",
        "hard_violations": [],
        "physics": "경찰은 한쪽 무릎을 현우의 등에 밀착하고 다른 무릎과 정강이를 지면에 내려 체중을 지지한다. 숙인 상체의 하중이 등 위 무릎으로 전달되는 자세가 명확하다. 현우의 뺨, 가슴, 팔뚝과 펼쳐진 손은 바닥에 놓이고, 하체와 신발도 지면에 지지된다. 손은 경찰의 셔츠 파지와 어깨 부근 접촉을 실제로 이루며, 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "장소와 현우의 외형은 잘 맞지만, 전신과 골목을 넓게 담아 경찰의 굳은 상체 중심 미디엄 숏에서 벗어나고 무릎이 등을 누르는 접점도 불명확하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "경찰의 굳은 상체를 크게 담으면서 등에 실린 무릎, 지면에 받쳐진 현우, 접점 주변 여백을 명확히 보여 주어 핵심 연출에 가장 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "경찰은 아래쪽 현우의 어깨와 팔 부위를 바라보며 상체를 숙인다. 현우의 얼굴은 화면 왼쪽 앞쪽으로 돌아가 있고 시선은 바닥 가까이 향한다. 무기나 겨냥하는 물체는 보이지 않는다. 경찰의 무릎은 현우의 허리 뒤쪽에 위치하지만, 등 위를 직접 아래로 누르는 방향과 접점은 팔에 가려 명확하지 않다.",
        "built_space": "왼쪽에는 붉은 차양과 여러 지지대, 과일 진열대 및 넘어진 상자들이 있고, 오른쪽에는 낡은 벽, 회색 문 하나, 두드러진 수직 배관 두 줄, 파란 통 하나와 목재 보호대 하나가 보인다. 전경 오른쪽의 사각 맨홀 하나와 갈라진 포장도 장소 참조에 부합한다. 두 사람은 골목 바닥에 놓이며 접점 주변 지면은 충분하지만, 현우의 전신과 넓은 배경을 포함해 요구한 경찰 상체 중심보다 넓게 잡혔다.",
        "entities": "보이는 사람은 현우와 체포 경찰 두 명이다. 현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 회색 겉셔츠, 올리브색 카고 바지, 갈색 신발이 참조와 대체로 맞는다. 얼굴 찰과상과 바지 무릎 부근의 손상도 보이지만 개 물림 자국인지는 확정할 수 없다. 경찰은 검은 사복을 입은 성인 동아시아계 남성이다. 흩어진 상자와 과일이 있으며 뚜렷하게 읽히는 문구나 화면 표식은 없다.",
        "hard_violations": [],
        "physics": "현우의 몸통과 다리는 포장면에 누워 있고 얼굴도 바닥 가까이 닿아 있다. 등 뒤로 꺾인 팔은 자기 몸과 경찰의 손에 의해 지지되는 것으로 보인다. 경찰은 굽힌 다리로 낮은 자세를 유지하고 손으로 현우를 고정한다. 지지 없이 떠 있는 신체는 확인되지 않지만, 체중이 무릎을 통해 현우의 등에 전달되는지는 가림 때문에 분명하지 않다."
       },
       {
        "label": "A",
        "direction": "경찰의 얼굴과 시선은 바로 아래 현우의 등과 머리 쪽을 향한다. 한쪽 무릎은 현우의 등 중앙을 아래로 누르고, 한 손은 등 쪽 셔츠를 잡으며 다른 손은 목 뒤와 어깨 가까이에 놓인다. 현우는 얼굴을 화면 오른쪽으로 돌린 채 뺨을 바닥에 대고 눈을 감고 있다. 무기나 다른 지향성 소품은 없다.",
        "built_space": "왼쪽 붉은 차양과 지지대, 과일 진열대, 넘어진 상자들이 참조의 시장 골목 배치를 유지한다. 오른쪽에는 낡은 벽과 세로 골판 외장, 굵고 가는 수직 배관, 파란 통 하나, 녹색 상자 하나와 비스듬한 붉은 판 하나가 보인다. 문 구역 일부는 경찰에게 가려져 있다. 경찰 상체가 화면의 큰 부분을 차지하며, 무릎과 등 및 현우의 얼굴·팔 주변에 바닥 여백이 있어 제압 구조를 읽기 쉽다.",
        "entities": "현우와 사복 경찰 두 명만 보인다. 현우의 젊은 동아시아계 남성 외형, 검은 머리, 회색 겉셔츠와 카고 바지는 참조와 대체로 일치하며, 얼굴에는 기존 찰과상으로 읽히는 상처가 있다. 얼굴은 눌리고 일부 가려져 세밀한 동일인 비교에는 제한이 있다. 다리의 개 물림 상처는 이 시점에서 분명하게 드러나지 않는다. 경찰은 어두운 재킷과 체크 셔츠를 입은 성인 동아시아계 남성이다. 흩어진 상자들이 유지되며 읽을 수 있는 글이나 그래픽 표식은 없다.",
        "hard_violations": [],
        "physics": "경찰은 한쪽 무릎을 현우의 등에 밀착하고 다른 무릎과 정강이를 지면에 내려 체중을 지지한다. 숙인 상체의 하중이 등 위 무릎으로 전달되는 자세가 명확하다. 현우의 뺨, 가슴, 팔뚝과 펼쳐진 손은 바닥에 놓이고, 하체와 신발도 지면에 지지된다. 손은 경찰의 셔츠 파지와 어깨 부근 접촉을 실제로 이루며, 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.238
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.238
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1238
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 '무릎으로 등 위를 짓누르는' 핵심 제압 자세를 정확히 연출했으며, 지정된 배경과 인물의 외형 조건을 훌륭하게 준수함."
   },
   {
    "label": "B",
    "score": 1238,
    "verdict_ko": "배경과 인물의 착장은 적절하나, 경찰이 무릎이 아닌 양손으로 등을 누르고 있어 프롬프트가 지정한 가장 중요한 물리적 접촉 자세를 위반함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_market_capture_alley_13d54d.png",
    "asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-ed8e-7160-9bb6-c4eee56217cb",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh16__bgfirst_bg.png",
   "bg_asset_id": "43c81765-5840-45d5-8f30-f37e76cc5841",
   "bg_record_key": "S20sh16::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "market_capture_alley",
   "groupbg_asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S20sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:47:49.984455+00:00",
  "fingerprint": "62743e2a0e77aacae4ef309c832d60190ab5739a2b709f1e719490356611b857",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S20sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S20sh16_sel.png",
  "source_sha256": "b313fd9a2f8c9b20697f0f8909121e025f7bea33e86ea71a9b2010da38796168",
  "file": "S20sh16_cine.png",
  "staged_sha256": "481d97a6a9578bf0dedb12b2a1fb9ea36cf6eb2674e24de6a75cff548cbbf46e",
  "latency_ms": 10549
 },
 "S21sh2::signage": {
  "fp": "6f54fd162b569634",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::885ecb7c61f2a28a": {
  "subjects": [],
  "subject_text": "라울의 컨테이너 내부\n잠자리를 둘 수 있는 소박한 컨테이너 주거 공간. 좁은 직사각형 실내에 출입문과 간단한 침구가 있다.",
  "identity": "canonical",
  "scope_id": "L170",
  "scope_role": "location_interior",
  "scope_sha": "bfeabefdb3bedf56"
 },
 "S21sh2::bgfirst_bg": {
  "input_fingerprint": "0d65fd073a5c8b80",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2__bgfirst_bg.png",
  "asset_id": "c6f5da55-c68d-4a4c-892e-04c73c186cf0",
  "input_asset_ids": [
   "f81f4ae2-7f77-4950-92a2-abc01a2199a4",
   "948d45e7-41c1-4297-974a-8863400e8720"
  ]
 },
 "S21sh2": {
  "input_fingerprint": "9c97f8c1d001d158",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door of Raul's container has been thrown open. 라울: He has just entered the container through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door of Raul's container has been thrown open. 라울: He has just entered the container through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door of Raul's container has been thrown open. 라울: He has just entered the container through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2__bgfirst_bg.png",
     "asset_id": "c6f5da55-c68d-4a4c-892e-04c73c186cf0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S21sh2.png",
     "asset_id": "f81f4ae2-7f77-4950-92a2-abc01a2199a4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L170B01.png",
     "asset_id": "948d45e7-41c1-4297-974a-8863400e8720",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "복도에 위치한 소년이 방 안쪽(카메라 방향)을 향해 시선을 두고 걸어오고 있음.",
    "built_space": "침대와 벽면 사진이 왼쪽에, 문이 오른쪽에 위치하는 등 실내 및 복도의 전체 공간 배치가 레퍼런스와 좌우로 완전히 반전되어 있음.",
    "entities": "라울의 외모, 꽁지머리, 의상 등은 레퍼런스와 정확히 일치함.",
    "hard_violations": [
     "[gpt-high] 침대와 방문뿐 아니라 복도 냉장고와 신발장의 위치까지 좌우 반전하여, 지정된 실제 장소와 양립하지 않는 고정 구조를 만들었다."
    ],
    "physics": "바닥을 딛고 걷는 자세가 자연스러우며 신체나 사물이 공중에 떠 있는 부분은 없음."
   },
   {
    "label": "B",
    "direction": "소년이 방 안쪽을 향해 시선을 두고 몸을 들이밀며 진입하고 있음.",
    "built_space": "전경의 침대와 사진 위치는 레퍼런스와 일치하나, 왼쪽 벽면이 골판 철재로 임의 변경되었고 복도의 주방이 완전히 다른 빈 컨테이너 복도로 묘사됨.",
    "entities": "라울의 인상착의가 일치하며, 눈을 크게 뜬 다급한 표정 연기가 잘 표현됨.",
    "hard_violations": [
     "[gpt-high] 참조에 없는 별도의 외부 출입구를 복도 왼쪽에 추가하여 고정된 장소 구조를 변경했다."
    ],
    "physics": "오른손으로 문짝을 짚고 오른발로 바닥을 단단히 디디고 있어 신체의 지지 상태와 체중 이동이 명확함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "배경 복도의 건축 구조를 다르게 묘사한 점은 감점 요소이나, 다급하게 문을 열고 들이닥치는 라울의 행동과 표정을 완벽하게 포착하여 샷 텍스트의 연출 의도를 잘 살렸습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "공간의 좌우 배치가 완전히 반전되는 오류가 발생했으며, 다급히 들이닥치는 순간보다는 단순히 복도를 걷는 모습에 가깝게 연출되어 역동성이 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "복도에 위치한 소년이 방 안쪽(카메라 방향)을 향해 시선을 두고 걸어오고 있음.",
        "built_space": "침대와 벽면 사진이 왼쪽에, 문이 오른쪽에 위치하는 등 실내 및 복도의 전체 공간 배치가 레퍼런스와 좌우로 완전히 반전되어 있음.",
        "entities": "라울의 외모, 꽁지머리, 의상 등은 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "바닥을 딛고 걷는 자세가 자연스러우며 신체나 사물이 공중에 떠 있는 부분은 없음."
       },
       {
        "label": "B",
        "direction": "소년이 방 안쪽을 향해 시선을 두고 몸을 들이밀며 진입하고 있음.",
        "built_space": "전경의 침대와 사진 위치는 레퍼런스와 일치하나, 왼쪽 벽면이 골판 철재로 임의 변경되었고 복도의 주방이 완전히 다른 빈 컨테이너 복도로 묘사됨.",
        "entities": "라울의 인상착의가 일치하며, 눈을 크게 뜬 다급한 표정 연기가 잘 표현됨.",
        "hard_violations": [],
        "physics": "오른손으로 문짝을 짚고 오른발로 바닥을 단단히 디디고 있어 신체의 지지 상태와 체중 이동이 명확함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "배경 복도의 건축 구조를 다르게 묘사한 점은 감점 요소이나, 다급하게 문을 열고 들이닥치는 라울의 행동과 표정을 완벽하게 포착하여 샷 텍스트의 연출 의도를 잘 살렸습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "공간의 좌우 배치가 완전히 반전되는 오류가 발생했으며, 다급히 들이닥치는 순간보다는 단순히 복도를 걷는 모습에 가깝게 연출되어 역동성이 떨어집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "복도에 위치한 소년이 방 안쪽(카메라 방향)을 향해 시선을 두고 걸어오고 있음.",
        "built_space": "침대와 벽면 사진이 왼쪽에, 문이 오른쪽에 위치하는 등 실내 및 복도의 전체 공간 배치가 레퍼런스와 좌우로 완전히 반전되어 있음.",
        "entities": "라울의 외모, 꽁지머리, 의상 등은 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "바닥을 딛고 걷는 자세가 자연스러우며 신체나 사물이 공중에 떠 있는 부분은 없음."
       },
       {
        "label": "B",
        "direction": "소년이 방 안쪽을 향해 시선을 두고 몸을 들이밀며 진입하고 있음.",
        "built_space": "전경의 침대와 사진 위치는 레퍼런스와 일치하나, 왼쪽 벽면이 골판 철재로 임의 변경되었고 복도의 주방이 완전히 다른 빈 컨테이너 복도로 묘사됨.",
        "entities": "라울의 인상착의가 일치하며, 눈을 크게 뜬 다급한 표정 연기가 잘 표현됨.",
        "hard_violations": [],
        "physics": "오른손으로 문짝을 짚고 오른발로 바닥을 단단히 디디고 있어 신체의 지지 상태와 체중 이동이 명확함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "문을 밀며 문턱 안으로 급히 들어오는 전신과 다급한 표정은 가장 충실하지만, 복도 왼쪽에 추가된 외부 출입구가 장소의 고정 구조를 바꾼다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전신과 침대 쪽 시선은 적절하지만, 장소 구조가 좌우로 뒤집혔고 앞발도 아직 문턱 밖에 있어 막 실내로 들이닥친 순간이 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "라울의 몸은 복도에서 카메라가 있는 방 안으로 전진하며, 시선은 화면 왼쪽의 실내를 향한다. 화면 오른쪽 침대를 직접 바라보지는 않지만, 지문은 시선의 특정 도착점을 지정하지 않는다. 한 손은 열린 문짝을 밀고 있다.",
        "built_space": "실내에서 비스듬히 문턱을 보는 전신 와이드숏이다. 전경의 방문과 문틀이 하나씩, 오른쪽에 목제 침대 하나와 체크 베개 하나, 벽 사진 네 장과 스위치 하나가 보인다. 이 배치는 참조의 침실과 대체로 맞는다. 그러나 복도 끝의 밝은 개구부 외에 복도 왼쪽 가까이에 바깥 지면까지 보이는 별도 출입구가 있어, 참조의 방문과 복도 끝 외부 출입구로 이어지는 구조에 개구부가 추가됐다. 문과 아이의 크기 관계는 자연스럽다.",
        "entities": "인물은 라울로 보이는 어린 남자아이 한 명뿐이다. 어린 얼굴, 갈색 피부, 뒤로 묶은 곱슬머리와 마른 체격이 인물 참조에 대체로 부합한다. 혼합 민족 배경 자체는 외형만으로 확정할 수 없다. 낡은 회갈색 반팔, 해진 카고 반바지, 어두운 운동화가 참조와 맞는다. 잠자는 아이를 별도로 추가하지 않았고, 읽을 수 있는 글자나 그래픽 표시는 보이지 않는다.",
        "hard_violations": [
         "참조에 없는 별도의 외부 출입구를 복도 왼쪽에 추가하여 고정된 장소 구조를 변경했다."
        ],
        "physics": "앞발은 문턱을 넘어 방 바닥에 닿아 체중을 받고, 뒤쪽 다리는 다음 걸음을 위해 따라오는 자세다. 문짝에 댄 손과 앞으로 기울인 상체가 급히 문을 열고 진입한 동작을 뒷받침한다. 문은 문틀에 연결되어 있고, 침대는 목제 다리로 바닥에 지지된다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "라울은 복도에서 방 쪽으로 몸을 기울이며 화면 왼쪽 침대를 바라본다. 몸과 시선이 방 안을 향한다는 점은 명확하다. 보이는 팔은 뒤로 벌어져 있고 손은 문짝이나 손잡이를 잡고 있지 않는다.",
        "built_space": "방 안에서 문을 비스듬히 보는 와이드숏이며 라울의 전신이 들어온다. 방문 하나와 뒤쪽 외부 출입구 하나, 침대 하나, 체크 베개 하나, 벽 사진 세 장, 스위치 하나가 보인다. 하지만 참조에서 오른쪽인 침대와 사진 벽이 왼쪽으로, 왼쪽인 열린 방문이 오른쪽으로 바뀌었다. 복도 냉장고와 신발장도 각각 반대편에 있어 단순한 카메라 이동이 아니라 장소 전체를 좌우 반전한 배치다. 라울의 앞발은 아직 침실 문턱의 복도 쪽에 있다.",
        "entities": "어린 남자아이 한 명만 등장하며 어린 얼굴, 갈색 피부, 뒤로 묶은 머리, 마른 체격은 라울 참조에 대체로 맞는다. 회갈색 반팔과 해진 카고 반바지, 낡은 운동화도 일치한다. 머리 뒤 묶음은 참조보다 조금 풍성하게 보인다. 별도의 잠자는 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "침대와 방문뿐 아니라 복도 냉장고와 신발장의 위치까지 좌우 반전하여, 지정된 실제 장소와 양립하지 않는 고정 구조를 만들었다."
        ],
        "physics": "앞발이 복도 바닥에 닿아 몸을 지지하고 뒤쪽 발은 들려 있어, 앞으로 내딛는 보행 자세로 성립한다. 상체를 방 안으로 내민 동작도 물리적으로 가능하다. 열린 문은 문틀에 연결되어 있으므로 손이 떨어져 있어도 지지 문제가 없다. 침대와 침구 역시 정상적으로 지지된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "문을 밀며 문턱 안으로 급히 들어오는 전신과 다급한 표정은 가장 충실하지만, 복도 왼쪽에 추가된 외부 출입구가 장소의 고정 구조를 바꾼다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전신과 침대 쪽 시선은 적절하지만, 장소 구조가 좌우로 뒤집혔고 앞발도 아직 문턱 밖에 있어 막 실내로 들이닥친 순간이 약하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "라울의 몸은 복도에서 카메라가 있는 방 안으로 전진하며, 시선은 화면 왼쪽의 실내를 향한다. 화면 오른쪽 침대를 직접 바라보지는 않지만, 지문은 시선의 특정 도착점을 지정하지 않는다. 한 손은 열린 문짝을 밀고 있다.",
        "built_space": "실내에서 비스듬히 문턱을 보는 전신 와이드숏이다. 전경의 방문과 문틀이 하나씩, 오른쪽에 목제 침대 하나와 체크 베개 하나, 벽 사진 네 장과 스위치 하나가 보인다. 이 배치는 참조의 침실과 대체로 맞는다. 그러나 복도 끝의 밝은 개구부 외에 복도 왼쪽 가까이에 바깥 지면까지 보이는 별도 출입구가 있어, 참조의 방문과 복도 끝 외부 출입구로 이어지는 구조에 개구부가 추가됐다. 문과 아이의 크기 관계는 자연스럽다.",
        "entities": "인물은 라울로 보이는 어린 남자아이 한 명뿐이다. 어린 얼굴, 갈색 피부, 뒤로 묶은 곱슬머리와 마른 체격이 인물 참조에 대체로 부합한다. 혼합 민족 배경 자체는 외형만으로 확정할 수 없다. 낡은 회갈색 반팔, 해진 카고 반바지, 어두운 운동화가 참조와 맞는다. 잠자는 아이를 별도로 추가하지 않았고, 읽을 수 있는 글자나 그래픽 표시는 보이지 않는다.",
        "hard_violations": [
         "참조에 없는 별도의 외부 출입구를 복도 왼쪽에 추가하여 고정된 장소 구조를 변경했다."
        ],
        "physics": "앞발은 문턱을 넘어 방 바닥에 닿아 체중을 받고, 뒤쪽 다리는 다음 걸음을 위해 따라오는 자세다. 문짝에 댄 손과 앞으로 기울인 상체가 급히 문을 열고 진입한 동작을 뒷받침한다. 문은 문틀에 연결되어 있고, 침대는 목제 다리로 바닥에 지지된다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "라울은 복도에서 방 쪽으로 몸을 기울이며 화면 왼쪽 침대를 바라본다. 몸과 시선이 방 안을 향한다는 점은 명확하다. 보이는 팔은 뒤로 벌어져 있고 손은 문짝이나 손잡이를 잡고 있지 않는다.",
        "built_space": "방 안에서 문을 비스듬히 보는 와이드숏이며 라울의 전신이 들어온다. 방문 하나와 뒤쪽 외부 출입구 하나, 침대 하나, 체크 베개 하나, 벽 사진 세 장, 스위치 하나가 보인다. 하지만 참조에서 오른쪽인 침대와 사진 벽이 왼쪽으로, 왼쪽인 열린 방문이 오른쪽으로 바뀌었다. 복도 냉장고와 신발장도 각각 반대편에 있어 단순한 카메라 이동이 아니라 장소 전체를 좌우 반전한 배치다. 라울의 앞발은 아직 침실 문턱의 복도 쪽에 있다.",
        "entities": "어린 남자아이 한 명만 등장하며 어린 얼굴, 갈색 피부, 뒤로 묶은 머리, 마른 체격은 라울 참조에 대체로 맞는다. 회갈색 반팔과 해진 카고 반바지, 낡은 운동화도 일치한다. 머리 뒤 묶음은 참조보다 조금 풍성하게 보인다. 별도의 잠자는 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "침대와 방문뿐 아니라 복도 냉장고와 신발장의 위치까지 좌우 반전하여, 지정된 실제 장소와 양립하지 않는 고정 구조를 만들었다."
        ],
        "physics": "앞발이 복도 바닥에 닿아 몸을 지지하고 뒤쪽 발은 들려 있어, 앞으로 내딛는 보행 자세로 성립한다. 상체를 방 안으로 내민 동작도 물리적으로 가능하다. 열린 문은 문틀에 연결되어 있으므로 손이 떨어져 있어도 지지 문제가 없다. 침대와 침구 역시 정상적으로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.171,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.921,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gpt-high] 참조에 없는 별도의 외부 출입구를 복도 왼쪽에 추가하여 고정된 장소 구조를 변경했다."
    ],
    "A": [
     "[gpt-high] 침대와 방문뿐 아니라 복도 냉장고와 신발장의 위치까지 좌우 반전하여, 지정된 실제 장소와 양립하지 않는 고정 구조를 만들었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 921
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "배경 복도의 건축 구조를 다르게 묘사한 점은 감점 요소이나, 다급하게 문을 열고 들이닥치는 라울의 행동과 표정을 완벽하게 포착하여 샷 텍스트의 연출 의도를 잘 살렸습니다.  ★위반: [gpt-high] 참조에 없는 별도의 외부 출입구를 복도 왼쪽에 추가하여 고정된 장소 구조를 변경했다."
   },
   {
    "label": "A",
    "score": 921,
    "verdict_ko": "공간의 좌우 배치가 완전히 반전되는 오류가 발생했으며, 다급히 들이닥치는 순간보다는 단순히 복도를 걷는 모습에 가깝게 연출되어 역동성이 떨어집니다.  ★위반: [gpt-high] 침대와 방문뿐 아니라 복도 냉장고와 신발장의 위치까지 좌우 반전하여, 지정된 실제 장소와 양립하지 않는 고정 구조를 만들었다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L170B01.png",
    "asset_id": "948d45e7-41c1-4297-974a-8863400e8720",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-f23e-7ccc-9118-062a4f60ba2b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2__bgfirst_bg.png",
   "bg_asset_id": "c6f5da55-c68d-4a4c-892e-04c73c186cf0",
   "bg_record_key": "S21sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S21sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:49:30.438079+00:00",
  "fingerprint": "7f46743abe46fe7249ec94cc135123f7ca95bd6575773f079670b64a3550782e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S21sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S21sh2_sel.png",
  "source_sha256": "9148f57510d4c9116f8111bb9ca2459744496bf3c1d9453762cddb6ee238306d",
  "file": "S21sh2_cine.png",
  "staged_sha256": "97563a4a7b42b8a2489bd622064c334bbe0a5dc9f327fe78ce15489416ccd8c7",
  "latency_ms": 9406
 },
 "S21sh5::signage": {
  "fp": "764871cb15ce99ad",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S21sh5": {
  "input_fingerprint": "b274c7101ecb83fa",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 라울의 손가락 끝이 가리키는 텅 빈 공간을 응시하며 눈동자가 커진 앰버의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the sleeping area inside the container home, facing an empty patch of room in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior beside 앰버 (Only a narrow, unoccupied portion is visible at the edge of the close framing); used as Retain breathing room on the side of her attention without inventing an object to explain 찰리's absence.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the interior's neutral ambient illumination and soft tonal separation, presenting the empty-space reaction as direct reality without a visual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 라울 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container door remains open after the abrupt entrance. 앰버: She has roused from sleep and is sitting up. Her mask and waist tool pouch remain established belongings without a stated removal. 라울: He remains inside the container after entering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 라울의 손가락 끝이 가리키는 텅 빈 공간을 응시하며 눈동자가 커진 앰버의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the sleeping area inside the container home, facing an empty patch of room in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior beside 앰버 (Only a narrow, unoccupied portion is visible at the edge of the close framing); used as Retain breathing room on the side of her attention without inventing an object to explain 찰리's absence.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the interior's neutral ambient illumination and soft tonal separation, presenting the empty-space reaction as direct reality without a visual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 라울 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container door remains open after the abrupt entrance. 앰버: She has roused from sleep and is sitting up. Her mask and waist tool pouch remain established belongings without a stated removal. 라울: He remains inside the container after entering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 라울의 손가락 끝이 가리키는 텅 빈 공간을 응시하며 눈동자가 커진 앰버의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the sleeping area inside the container home, facing an empty patch of room in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior beside 앰버 (Only a narrow, unoccupied portion is visible at the edge of the close framing); used as Retain breathing room on the side of her attention without inventing an object to explain 찰리's absence.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the interior's neutral ambient illumination and soft tonal separation, presenting the empty-space reaction as direct reality without a visual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 라울 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container door remains open after the abrupt entrance. 앰버: She has roused from sleep and is sitting up. Her mask and waist tool pouch remain established belongings without a stated removal. 라울: He remains inside the container after entering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "오른쪽에서 들어온 라울의 검지가 화면 왼쪽을 가리키며, 그 방향에는 사람이 없는 벽면과 화면 밖 공간이 있다. 앰버의 눈도 왼쪽으로 향해 손가락 방향과 대체로 일치한다. 다만 화면 밖의 정확한 응시 지점까지 확인되지는 않는다.",
    "built_space": "밝은 세로 패널 벽과 골이 깊은 금속 벽, 목재 침대 한 개, 짙은 침구, 체크무늬 베개 두 개가 보인다. 이전 장면의 재료와 침구 계열은 이어지지만, 이전에는 베개 하나만 확인된다. 앰버는 침대 위에서 상체를 세우고 있다. 문은 화면 밖이므로 개방 상태를 판단할 수 없다. 배경과 침대가 넓게 보여 가장자리에 좁은 빈 공간만 남기는 얼굴 클로즈업과는 다르다.",
    "entities": "앰버는 금발의 어린 여자아이로, 큰 눈과 둥근 얼굴, 남색 상의와 갈색 작업복을 갖췄다. 참조보다 얼굴이 조금 길고 서구적으로 보이지만 외모만으로 혼혈 배경을 확정할 수는 없다. 마스크는 참조의 머리 위가 아니라 목과 가슴 앞에 있으며, 허리 도구와 주머니도 보인다. 라울은 갈색 피부의 팔과 낡은 회갈색 소매만 보여 이전 장면의 복장은 이어진다. 얼굴과 머리는 평가할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "앰버의 하체 접촉점은 일부 잘렸지만 침대에 앉아 상체를 세운 자세로 자연스럽게 이어진다. 라울의 손은 손목과 팔, 화면 오른쪽의 소매로 연결되어 있으며 공중에 분리되어 있지 않다. 마스크는 목의 끈에 걸려 가슴에 닿고, 베개와 침구는 침대가 받친다. 눈을 크게 뜬 표정도 정상적인 인체 범위다."
   },
   {
    "label": "A",
    "direction": "라울의 검지는 화면 오른쪽의 비어 있는 벽 쪽을 가리킨다. 반면 앰버의 두 눈은 왼쪽 전경의 라울 얼굴 쪽을 향한다. 따라서 앰버가 손끝이 가리키는 빈 공간을 응시한다는 핵심 행동과 반대다. 라울의 얼굴은 오른쪽을 향하지만 정확한 시선 종착점은 옆모습으로 명확하지 않다.",
    "built_space": "밝은 세로 패널 벽, 침대 한 개의 짙은 침구와 체크무늬 베개 한 개가 보인다. 앰버는 침대에 앉아 있고 라울은 왼쪽 전경을 크게 차지한다. 문과 다른 고정 설비는 화면 밖이라 연속성을 확인할 수 없다. 침실 재료는 참조와 양립하지만, 넓은 벽면과 두 인물의 상체를 담아 요청한 얼굴 클로즈업 및 좁은 주변 여백을 벗어난다.",
    "entities": "앰버는 금발의 어린 여자아이이며 남색 상의, 갈색 작업복, 머리 위의 방독 마스크가 참조와 대체로 맞는다. 큰 눈과 벌어진 입으로 놀람을 표현한다. 라울은 갈색 피부의 어린아이로 뒤로 묶은 곱슬머리와 이전 장면의 낡은 회갈색 상의를 유지한다. 두 사람의 구체적인 혼혈 배경은 외모만으로 확인할 수 없다. 허리 도구 주머니는 잘려 있어 유무를 판정하지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "라울의 뻗은 팔은 보이는 어깨와 몸통에 연결되고 검지를 펴는 동작도 가능하다. 앰버는 침대 위에 앉은 몸통 자세이며, 정확한 엉덩이 접촉점은 하단 밖에 있다. 마스크는 머리에 얹혀 지지되고 베개와 침구는 침대 위에 놓여 있다. 지지 없이 떠 있는 물체나 불가능한 관절은 보이지 않는다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "손가락과 앰버의 시선이 왼쪽 빈 공간으로 이어져 B보다 충실하지만, 허리까지 드러낸 구도는 요청한 얼굴 클로즈업을 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "라울은 오른쪽을 가리키는데 앰버는 왼쪽의 라울을 바라보므로 핵심 시선 관계가 틀렸으며, 얼굴 클로즈업 대신 넓은 두 인물 구도가 되었다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽에서 들어온 라울의 검지가 화면 왼쪽을 가리키며, 그 방향에는 사람이 없는 벽면과 화면 밖 공간이 있다. 앰버의 눈도 왼쪽으로 향해 손가락 방향과 대체로 일치한다. 다만 화면 밖의 정확한 응시 지점까지 확인되지는 않는다.",
        "built_space": "밝은 세로 패널 벽과 골이 깊은 금속 벽, 목재 침대 한 개, 짙은 침구, 체크무늬 베개 두 개가 보인다. 이전 장면의 재료와 침구 계열은 이어지지만, 이전에는 베개 하나만 확인된다. 앰버는 침대 위에서 상체를 세우고 있다. 문은 화면 밖이므로 개방 상태를 판단할 수 없다. 배경과 침대가 넓게 보여 가장자리에 좁은 빈 공간만 남기는 얼굴 클로즈업과는 다르다.",
        "entities": "앰버는 금발의 어린 여자아이로, 큰 눈과 둥근 얼굴, 남색 상의와 갈색 작업복을 갖췄다. 참조보다 얼굴이 조금 길고 서구적으로 보이지만 외모만으로 혼혈 배경을 확정할 수는 없다. 마스크는 참조의 머리 위가 아니라 목과 가슴 앞에 있으며, 허리 도구와 주머니도 보인다. 라울은 갈색 피부의 팔과 낡은 회갈색 소매만 보여 이전 장면의 복장은 이어진다. 얼굴과 머리는 평가할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 하체 접촉점은 일부 잘렸지만 침대에 앉아 상체를 세운 자세로 자연스럽게 이어진다. 라울의 손은 손목과 팔, 화면 오른쪽의 소매로 연결되어 있으며 공중에 분리되어 있지 않다. 마스크는 목의 끈에 걸려 가슴에 닿고, 베개와 침구는 침대가 받친다. 눈을 크게 뜬 표정도 정상적인 인체 범위다."
       },
       {
        "label": "B",
        "direction": "라울의 검지는 화면 오른쪽의 비어 있는 벽 쪽을 가리킨다. 반면 앰버의 두 눈은 왼쪽 전경의 라울 얼굴 쪽을 향한다. 따라서 앰버가 손끝이 가리키는 빈 공간을 응시한다는 핵심 행동과 반대다. 라울의 얼굴은 오른쪽을 향하지만 정확한 시선 종착점은 옆모습으로 명확하지 않다.",
        "built_space": "밝은 세로 패널 벽, 침대 한 개의 짙은 침구와 체크무늬 베개 한 개가 보인다. 앰버는 침대에 앉아 있고 라울은 왼쪽 전경을 크게 차지한다. 문과 다른 고정 설비는 화면 밖이라 연속성을 확인할 수 없다. 침실 재료는 참조와 양립하지만, 넓은 벽면과 두 인물의 상체를 담아 요청한 얼굴 클로즈업 및 좁은 주변 여백을 벗어난다.",
        "entities": "앰버는 금발의 어린 여자아이이며 남색 상의, 갈색 작업복, 머리 위의 방독 마스크가 참조와 대체로 맞는다. 큰 눈과 벌어진 입으로 놀람을 표현한다. 라울은 갈색 피부의 어린아이로 뒤로 묶은 곱슬머리와 이전 장면의 낡은 회갈색 상의를 유지한다. 두 사람의 구체적인 혼혈 배경은 외모만으로 확인할 수 없다. 허리 도구 주머니는 잘려 있어 유무를 판정하지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "라울의 뻗은 팔은 보이는 어깨와 몸통에 연결되고 검지를 펴는 동작도 가능하다. 앰버는 침대 위에 앉은 몸통 자세이며, 정확한 엉덩이 접촉점은 하단 밖에 있다. 마스크는 머리에 얹혀 지지되고 베개와 침구는 침대 위에 놓여 있다. 지지 없이 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "손가락과 앰버의 시선이 왼쪽 빈 공간으로 이어져 B보다 충실하지만, 허리까지 드러낸 구도는 요청한 얼굴 클로즈업을 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "라울은 오른쪽을 가리키는데 앰버는 왼쪽의 라울을 바라보므로 핵심 시선 관계가 틀렸으며, 얼굴 클로즈업 대신 넓은 두 인물 구도가 되었다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽에서 들어온 라울의 검지가 화면 왼쪽을 가리키며, 그 방향에는 사람이 없는 벽면과 화면 밖 공간이 있다. 앰버의 눈도 왼쪽으로 향해 손가락 방향과 대체로 일치한다. 다만 화면 밖의 정확한 응시 지점까지 확인되지는 않는다.",
        "built_space": "밝은 세로 패널 벽과 골이 깊은 금속 벽, 목재 침대 한 개, 짙은 침구, 체크무늬 베개 두 개가 보인다. 이전 장면의 재료와 침구 계열은 이어지지만, 이전에는 베개 하나만 확인된다. 앰버는 침대 위에서 상체를 세우고 있다. 문은 화면 밖이므로 개방 상태를 판단할 수 없다. 배경과 침대가 넓게 보여 가장자리에 좁은 빈 공간만 남기는 얼굴 클로즈업과는 다르다.",
        "entities": "앰버는 금발의 어린 여자아이로, 큰 눈과 둥근 얼굴, 남색 상의와 갈색 작업복을 갖췄다. 참조보다 얼굴이 조금 길고 서구적으로 보이지만 외모만으로 혼혈 배경을 확정할 수는 없다. 마스크는 참조의 머리 위가 아니라 목과 가슴 앞에 있으며, 허리 도구와 주머니도 보인다. 라울은 갈색 피부의 팔과 낡은 회갈색 소매만 보여 이전 장면의 복장은 이어진다. 얼굴과 머리는 평가할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 하체 접촉점은 일부 잘렸지만 침대에 앉아 상체를 세운 자세로 자연스럽게 이어진다. 라울의 손은 손목과 팔, 화면 오른쪽의 소매로 연결되어 있으며 공중에 분리되어 있지 않다. 마스크는 목의 끈에 걸려 가슴에 닿고, 베개와 침구는 침대가 받친다. 눈을 크게 뜬 표정도 정상적인 인체 범위다."
       },
       {
        "label": "A",
        "direction": "라울의 검지는 화면 오른쪽의 비어 있는 벽 쪽을 가리킨다. 반면 앰버의 두 눈은 왼쪽 전경의 라울 얼굴 쪽을 향한다. 따라서 앰버가 손끝이 가리키는 빈 공간을 응시한다는 핵심 행동과 반대다. 라울의 얼굴은 오른쪽을 향하지만 정확한 시선 종착점은 옆모습으로 명확하지 않다.",
        "built_space": "밝은 세로 패널 벽, 침대 한 개의 짙은 침구와 체크무늬 베개 한 개가 보인다. 앰버는 침대에 앉아 있고 라울은 왼쪽 전경을 크게 차지한다. 문과 다른 고정 설비는 화면 밖이라 연속성을 확인할 수 없다. 침실 재료는 참조와 양립하지만, 넓은 벽면과 두 인물의 상체를 담아 요청한 얼굴 클로즈업 및 좁은 주변 여백을 벗어난다.",
        "entities": "앰버는 금발의 어린 여자아이이며 남색 상의, 갈색 작업복, 머리 위의 방독 마스크가 참조와 대체로 맞는다. 큰 눈과 벌어진 입으로 놀람을 표현한다. 라울은 갈색 피부의 어린아이로 뒤로 묶은 곱슬머리와 이전 장면의 낡은 회갈색 상의를 유지한다. 두 사람의 구체적인 혼혈 배경은 외모만으로 확인할 수 없다. 허리 도구 주머니는 잘려 있어 유무를 판정하지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "라울의 뻗은 팔은 보이는 어깨와 몸통에 연결되고 검지를 펴는 동작도 가능하다. 앰버는 침대 위에 앉은 몸통 자세이며, 정확한 엉덩이 접촉점은 하단 밖에 있다. 마스크는 머리에 얹혀 지지되고 베개와 침구는 침대 위에 놓여 있다. 지지 없이 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 6,
   "A": 3
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 6,
    "verdict_ko": "손가락과 앰버의 시선이 왼쪽 빈 공간으로 이어져 B보다 충실하지만, 허리까지 드러낸 구도는 요청한 얼굴 클로즈업을 충족하지 못한다."
   },
   {
    "label": "A",
    "score": 3,
    "verdict_ko": "라울은 오른쪽을 가리키는데 앰버는 왼쪽의 라울을 바라보므로 핵심 시선 관계가 틀렸으며, 얼굴 클로즈업 대신 넓은 두 인물 구도가 되었다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 라울 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2_sel.png",
    "asset_id": "6a2bc852-9447-4d4b-995c-be6463de09e0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-f57e-7143-ac1b-13a0b0298fff",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S21sh2"
  },
  "staged_characters_added": [
   "C08"
  ]
 },
 "S21sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:50:28.983052+00:00",
  "fingerprint": "ee93a9ec940d3b519e5a40b25eb2ad5637cb587a4158176048e2b642cd2cb15d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S21sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S21sh5_sel.png",
  "source_sha256": "26376340b4aa155728f4091424fcdc2036dc29406431153b979868361535d19f",
  "file": "S21sh5_cine.png",
  "staged_sha256": "aefbd4b6ec0fd89918e3501da08c20f086070957a27e4d3dbc9e0b3f55c12085",
  "latency_ms": 9609
 },
 "S22sh3::signage": {
  "fp": "4ca48b7b76a20d29",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S22sh3": {
  "input_fingerprint": "9d351f8577c8e335",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bananas are displayed for sale at the market stall. Charlie's washed but worn gorilla-shaped body retains its blue-lit eyes and worn Ubik logo, with the old coat and hat still serving as his disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bananas are displayed for sale at the market stall. Charlie's washed but worn gorilla-shaped body retains its blue-lit eyes and worn Ubik logo, with the old coat and hat still serving as his disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bananas are displayed for sale at the market stall. Charlie's washed but worn gorilla-shaped body retains its blue-lit eyes and worn Ubik logo, with the old coat and hat still serving as his disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3__bgfirst_bg.png",
     "asset_id": "e6ccf295-a01a-4188-a4ca-5ce36e1609e9",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S22sh3.png",
     "asset_id": "774388e6-e167-4f76-a648-7a8dce68ade1",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_banana_market_stall_8677ef.png",
     "asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "금속 손가락이 가판대의 바나나를 향해 뻗어 있음.",
    "built_space": "야외 시장 가판대에 바나나가 진열되어 있으며, 레퍼런스와 일치하는 배경 구조물들이 보임.",
    "entities": "찰리(오렌지색 눈, 로고 없음), 바나나, 그리고 샷 텍스트에 없는 배경의 상인.",
    "hard_violations": [
     "[gemini-pro] invented people: 샷 텍스트에 없는 상인이 배경에 등장함.",
     "[gpt-high] 찰리만 허용된 장면에 저울 뒤의 상인과 중앙 통로의 행인을 추가했다."
    ],
    "physics": "몸통에 연결된 팔이 허공을 가로질러 뻗어 있음."
   },
   {
    "label": "B",
    "direction": "금속 손가락이 가판대의 바나나를 향해 뻗어 있음.",
    "built_space": "야외 시장 가판대에 바나나가 진열되어 있으며, 레퍼런스와 일치하는 배경 구조물들이 보임.",
    "entities": "찰리의 금속 팔과 코트(Ubik 로고 확인됨), 바나나. 지정되지 않은 추가 인물 없음.",
    "hard_violations": [
     "[gpt-high] 팔 장갑의 ‘Ubik’ 표기가 선명하게 판독되어, 읽을 수 있는 문자와 로고를 금지한 지시를 위반한다."
    ],
    "physics": "화면 밖 몸통에 연결된 팔이 자연스럽게 뻗어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지정된 프레이밍을 정확히 구현했으며, 요구된 Ubik 로고를 포함하고 샷 텍스트에 없는 인물을 배제하여 지침을 훌륭하게 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 명시되지 않은 상인을 배경에 추가하는 치명적인 규정 위반을 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "금속 손가락이 가판대의 바나나를 향해 뻗어 있음.",
        "built_space": "야외 시장 가판대에 바나나가 진열되어 있으며, 레퍼런스와 일치하는 배경 구조물들이 보임.",
        "entities": "찰리(오렌지색 눈, 로고 없음), 바나나, 그리고 샷 텍스트에 없는 배경의 상인.",
        "hard_violations": [
         "invented people: 샷 텍스트에 없는 상인이 배경에 등장함."
        ],
        "physics": "몸통에 연결된 팔이 허공을 가로질러 뻗어 있음."
       },
       {
        "label": "B",
        "direction": "금속 손가락이 가판대의 바나나를 향해 뻗어 있음.",
        "built_space": "야외 시장 가판대에 바나나가 진열되어 있으며, 레퍼런스와 일치하는 배경 구조물들이 보임.",
        "entities": "찰리의 금속 팔과 코트(Ubik 로고 확인됨), 바나나. 지정되지 않은 추가 인물 없음.",
        "hard_violations": [],
        "physics": "화면 밖 몸통에 연결된 팔이 자연스럽게 뻗어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지정된 프레이밍을 정확히 구현했으며, 요구된 Ubik 로고를 포함하고 샷 텍스트에 없는 인물을 배제하여 지침을 훌륭하게 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 명시되지 않은 상인을 배경에 추가하는 치명적인 규정 위반을 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "금속 손가락이 가판대의 바나나를 향해 뻗어 있음.",
        "built_space": "야외 시장 가판대에 바나나가 진열되어 있으며, 레퍼런스와 일치하는 배경 구조물들이 보임.",
        "entities": "찰리(오렌지색 눈, 로고 없음), 바나나, 그리고 샷 텍스트에 없는 배경의 상인.",
        "hard_violations": [
         "invented people: 샷 텍스트에 없는 상인이 배경에 등장함."
        ],
        "physics": "몸통에 연결된 팔이 허공을 가로질러 뻗어 있음."
       },
       {
        "label": "B",
        "direction": "금속 손가락이 가판대의 바나나를 향해 뻗어 있음.",
        "built_space": "야외 시장 가판대에 바나나가 진열되어 있으며, 레퍼런스와 일치하는 배경 구조물들이 보임.",
        "entities": "찰리의 금속 팔과 코트(Ubik 로고 확인됨), 바나나. 지정되지 않은 추가 인물 없음.",
        "hard_violations": [],
        "physics": "화면 밖 몸통에 연결된 팔이 자연스럽게 뻗어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "바나나에 닿기 직전의 금속 손 클로즈업과 찰리의 장갑·외투는 더 충실하지만, 팔의 선명한 ‘Ubik’ 글자가 판독 가능한 문자 금지를 위반한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "손가락은 바나나를 향하지만 금지된 상인과 행인을 추가했고, 얼굴까지 포함한 구도와 주황색 눈빛도 지정된 손 클로즈업 및 파란 눈 설정에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "팔이 왼쪽에서 오른쪽 진열대로 뻗어 있고, 펼친 금속 손가락들은 오른쪽 아래 나무 상자의 바나나를 향한다. 손끝과 과일 사이에는 간격이 있어 아직 만지지 않은 순간으로 읽힌다. 얼굴과 시선은 화면 밖이다.",
        "built_space": "오른쪽 아래에 바나나를 받치는 나무 진열 상자, 그 뒤에 감귤류 상자와 녹색 운반 상자들이 보인다. 오른쪽 위에는 매달린 바나나 한 송이, 뒤쪽 판매대에는 저울 한 대가 보인다. 금속 지지대와 천막, 시장 통로는 장소 참조와 대응하지만 중앙의 커다란 빨강·파랑 파라솔은 참조보다 두드러진다. 팔 옆에서 진열대를 비스듬히 보는 위치는 맞으며, 손 외에 몸통과 시장 배경도 상당 부분 포함한다.",
        "entities": "보이는 인물 부분은 찰리의 팔과 외투 입은 몸통뿐이다. 닳은 샌드 베이지 장갑판, 검은 기계 관절, 육중한 금속 손과 낡은 올리브색 외투는 참조에 부합한다. 얼굴·눈·모자는 프레임 밖이므로 평가 대상이 아니다. 판매용 바나나는 분명하며 전체 화면의 3분의 1 미만을 차지한다. 팔 장갑에는 ‘Ubik’이라는 글자가 읽힌다.",
        "hard_violations": [
         "팔 장갑의 ‘Ubik’ 표기가 선명하게 판독되어, 읽을 수 있는 문자와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "손은 손목 관절과 전완을 통해 화면 왼쪽의 몸통으로 연결되어 있으며, 팔을 뻗는 자세로 지지된다. 손가락 관절의 굽힘과 접촉 전 간격은 가능한 동작이다. 진열 바나나는 나무 상자에 놓여 있고, 위쪽 바나나는 줄에 매달려 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "길게 편 금속 손가락 하나가 오른쪽 아래 진열 바나나를 향하며 끝은 과일에 닿지 않는다. 찰리의 고개도 손과 바나나 쪽으로 숙여져 있다. 뒤쪽 상인은 자기 앞 저울과 판매대 쪽을 내려다본다.",
        "built_space": "오른쪽의 나무 바나나 진열 상자, 뒤의 감귤류 상자와 녹색 상자들, 위쪽에 매달린 바나나 한 송이, 중경의 저울 한 대가 보인다. 금속 기둥과 붉은 천막, 투명 가림막 및 통로는 장소 참조와 대체로 일치한다. 다만 저울 뒤에 상인이 서 있고 중앙 통로에도 행인이 있어 지정되지 않은 인물 배치가 생겼다. 손은 크게 보이지만 얼굴과 상체, 넓은 통로까지 포함하여 손 중심의 제한된 클로즈업에서 벗어난다.",
        "entities": "찰리의 샌드 베이지 금속 팔, 흰 마스크형 얼굴, 낡은 외투와 모자가 보인다. 눈빛은 파란색이 아니라 주황색이다. 오른쪽에는 모자와 녹색 앞치마를 착용한 성인 남성으로 보이는 상인이 있고, 중앙 통로에는 어두운 옷의 행인이 추가되어 있다. 먼 인물의 나이와 민족적 정체성은 판별하기 어렵다. 바나나는 판매용 진열 과일로 명확하며 화면의 3분의 1 미만을 차지한다.",
        "hard_violations": [
         "찰리만 허용된 장면에 저울 뒤의 상인과 중앙 통로의 행인을 추가했다."
        ],
        "physics": "전경의 금속 손은 손목과 팔을 거쳐 찰리의 몸통에 연결되어 있다. 한 손가락을 펴고 나머지를 굽힌 자세는 바나나에 손을 뻗는 동작으로 가능하며, 과일과 접촉하지 않는다. 바나나는 상자 또는 위쪽 줄로 지지된다. 상인의 하체는 판매대에 가려져 있으나 공중에 떠 있다는 징후는 없고, 통로의 행인도 지면 위에 서 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "바나나에 닿기 직전의 금속 손 클로즈업과 찰리의 장갑·외투는 더 충실하지만, 팔의 선명한 ‘Ubik’ 글자가 판독 가능한 문자 금지를 위반한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "손가락은 바나나를 향하지만 금지된 상인과 행인을 추가했고, 얼굴까지 포함한 구도와 주황색 눈빛도 지정된 손 클로즈업 및 파란 눈 설정에서 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "팔이 왼쪽에서 오른쪽 진열대로 뻗어 있고, 펼친 금속 손가락들은 오른쪽 아래 나무 상자의 바나나를 향한다. 손끝과 과일 사이에는 간격이 있어 아직 만지지 않은 순간으로 읽힌다. 얼굴과 시선은 화면 밖이다.",
        "built_space": "오른쪽 아래에 바나나를 받치는 나무 진열 상자, 그 뒤에 감귤류 상자와 녹색 운반 상자들이 보인다. 오른쪽 위에는 매달린 바나나 한 송이, 뒤쪽 판매대에는 저울 한 대가 보인다. 금속 지지대와 천막, 시장 통로는 장소 참조와 대응하지만 중앙의 커다란 빨강·파랑 파라솔은 참조보다 두드러진다. 팔 옆에서 진열대를 비스듬히 보는 위치는 맞으며, 손 외에 몸통과 시장 배경도 상당 부분 포함한다.",
        "entities": "보이는 인물 부분은 찰리의 팔과 외투 입은 몸통뿐이다. 닳은 샌드 베이지 장갑판, 검은 기계 관절, 육중한 금속 손과 낡은 올리브색 외투는 참조에 부합한다. 얼굴·눈·모자는 프레임 밖이므로 평가 대상이 아니다. 판매용 바나나는 분명하며 전체 화면의 3분의 1 미만을 차지한다. 팔 장갑에는 ‘Ubik’이라는 글자가 읽힌다.",
        "hard_violations": [
         "팔 장갑의 ‘Ubik’ 표기가 선명하게 판독되어, 읽을 수 있는 문자와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "손은 손목 관절과 전완을 통해 화면 왼쪽의 몸통으로 연결되어 있으며, 팔을 뻗는 자세로 지지된다. 손가락 관절의 굽힘과 접촉 전 간격은 가능한 동작이다. 진열 바나나는 나무 상자에 놓여 있고, 위쪽 바나나는 줄에 매달려 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "길게 편 금속 손가락 하나가 오른쪽 아래 진열 바나나를 향하며 끝은 과일에 닿지 않는다. 찰리의 고개도 손과 바나나 쪽으로 숙여져 있다. 뒤쪽 상인은 자기 앞 저울과 판매대 쪽을 내려다본다.",
        "built_space": "오른쪽의 나무 바나나 진열 상자, 뒤의 감귤류 상자와 녹색 상자들, 위쪽에 매달린 바나나 한 송이, 중경의 저울 한 대가 보인다. 금속 기둥과 붉은 천막, 투명 가림막 및 통로는 장소 참조와 대체로 일치한다. 다만 저울 뒤에 상인이 서 있고 중앙 통로에도 행인이 있어 지정되지 않은 인물 배치가 생겼다. 손은 크게 보이지만 얼굴과 상체, 넓은 통로까지 포함하여 손 중심의 제한된 클로즈업에서 벗어난다.",
        "entities": "찰리의 샌드 베이지 금속 팔, 흰 마스크형 얼굴, 낡은 외투와 모자가 보인다. 눈빛은 파란색이 아니라 주황색이다. 오른쪽에는 모자와 녹색 앞치마를 착용한 성인 남성으로 보이는 상인이 있고, 중앙 통로에는 어두운 옷의 행인이 추가되어 있다. 먼 인물의 나이와 민족적 정체성은 판별하기 어렵다. 바나나는 판매용 진열 과일로 명확하며 화면의 3분의 1 미만을 차지한다.",
        "hard_violations": [
         "찰리만 허용된 장면에 저울 뒤의 상인과 중앙 통로의 행인을 추가했다."
        ],
        "physics": "전경의 금속 손은 손목과 팔을 거쳐 찰리의 몸통에 연결되어 있다. 한 손가락을 펴고 나머지를 굽힌 자세는 바나나에 손을 뻗는 동작으로 가능하며, 과일과 접촉하지 않는다. 바나나는 상자 또는 위쪽 줄로 지지된다. 상인의 하체는 판매대에 가려져 있으나 공중에 떠 있다는 징후는 없고, 통로의 행인도 지면 위에 서 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.833,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.583,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people: 샷 텍스트에 없는 상인이 배경에 등장함.",
     "[gpt-high] 찰리만 허용된 장면에 저울 뒤의 상인과 중앙 통로의 행인을 추가했다."
    ],
    "B": [
     "[gpt-high] 팔 장갑의 ‘Ubik’ 표기가 선명하게 판독되어, 읽을 수 있는 문자와 로고를 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 583
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지정된 프레이밍을 정확히 구현했으며, 요구된 Ubik 로고를 포함하고 샷 텍스트에 없는 인물을 배제하여 지침을 훌륭하게 따랐습니다.  ★위반: [gpt-high] 팔 장갑의 ‘Ubik’ 표기가 선명하게 판독되어, 읽을 수 있는 문자와 로고를 금지한 지시를 위반한다."
   },
   {
    "label": "A",
    "score": 583,
    "verdict_ko": "샷 텍스트에 명시되지 않은 상인을 배경에 추가하는 치명적인 규정 위반을 범했습니다.  ★위반: [gemini-pro] invented people: 샷 텍스트에 없는 상인이 배경에 등장함. / [gpt-high] 찰리만 허용된 장면에 저울 뒤의 상인과 중앙 통로의 행인을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_banana_market_stall_8677ef.png",
    "asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-f725-713b-9f0b-97e68231a0f0",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3__bgfirst_bg.png",
   "bg_asset_id": "e6ccf295-a01a-4188-a4ca-5ce36e1609e9",
   "bg_record_key": "S22sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "banana_market_stall",
   "groupbg_asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S22sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:51:25.016643+00:00",
  "fingerprint": "9ab8c37c3733f38cccd76558ccfbf7b71babfa9583e4325bb40a7415b80e4845",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S22sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S22sh3_sel.png",
  "source_sha256": "bc65f5714330f9edea19849f1fb4be48bf739a396d16524283aa14875e68c61a",
  "file": "S22sh3_cine.png",
  "staged_sha256": "21e82df9209bae7c12ff7d0963b6da92bdca6ad4b85415fd3b8562854ef24dc0",
  "latency_ms": 9229
 },
 "S22sh7::signage": {
  "fp": "f1377a0da1e411fd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S22sh7": {
  "input_fingerprint": "81935c5a09549be9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수많은 사람들의 인파 속을 향해 한 발을 앞으로 내디딘 mid-stride 자세의 찰리의 거대한 뒷모습.\n\nLOCATION (lock): In the crowded outdoor market lane of the refugee settlement, beyond the banana stall. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route into the crowd in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage through the crowd (Occupied by numerous people with naturally varying gaps and unsynchronized steps) — The visible route recedes ahead of 찰리 toward the upper center; used as Give his forward step a definite destination while allowing the crowd to progressively conceal him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue neutral daytime ambient illumination with controlled contrast, separating 찰리 from the crowd through tonal organization rather than a new light cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains stocked, with no banana taken. Charlie moves into the market crowd still wearing the old coat and hat over his washed, worn robot body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수많은 사람들의 인파 속을 향해 한 발을 앞으로 내디딘 mid-stride 자세의 찰리의 거대한 뒷모습.\n\nLOCATION (lock): In the crowded outdoor market lane of the refugee settlement, beyond the banana stall. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route into the crowd in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage through the crowd (Occupied by numerous people with naturally varying gaps and unsynchronized steps) — The visible route recedes ahead of 찰리 toward the upper center; used as Give his forward step a definite destination while allowing the crowd to progressively conceal him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue neutral daytime ambient illumination with controlled contrast, separating 찰리 from the crowd through tonal organization rather than a new light cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains stocked, with no banana taken. Charlie moves into the market crowd still wearing the old coat and hat over his washed, worn robot body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수많은 사람들의 인파 속을 향해 한 발을 앞으로 내디딘 mid-stride 자세의 찰리의 거대한 뒷모습.\n\nLOCATION (lock): In the crowded outdoor market lane of the refugee settlement, beyond the banana stall. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route into the crowd in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage through the crowd (Occupied by numerous people with naturally varying gaps and unsynchronized steps) — The visible route recedes ahead of 찰리 toward the upper center; used as Give his forward step a definite destination while allowing the crowd to progressively conceal him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue neutral daytime ambient illumination with controlled contrast, separating 찰리 from the crowd through tonal organization rather than a new light cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains stocked, with no banana taken. Charlie moves into the market crowd still wearing the old coat and hat over his washed, worn robot body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 카메라를 등지고 시장 골목 안쪽을 향해 걷고 있으며, 군중들 사이로 향하고 있습니다.",
    "built_space": "야외 시장 골목으로, 우측에 이전 샷과 일치하는 바나나 가판대와 빨간색/파란색 파라솔이 배치되어 있습니다.",
    "entities": "찰리의 뒷모습(중절모, 벨트가 있는 코트, 로봇 팔다리)이 레퍼런스와 일치합니다. 군중과 바나나 가판대 모두 지시사항에 맞게 묘사되었습니다.",
    "hard_violations": [],
    "physics": "찰리의 오른발이 들려 있고 왼발이 땅을 지탱하여 자연스러운 걷기(mid-stride) 동작을 보여줍니다."
   },
   {
    "label": "B",
    "direction": "찰리는 카메라를 등지고 골목을 향해 걷고 있으나, 주변 군중 다수는 찰리 쪽을 향해 걸어오고 있습니다.",
    "built_space": "야외 시장 골목 환경이며, 우측에 과일과 바나나 가판대가 배치되어 있습니다.",
    "entities": "다인종 군중이 묘사되었으나, 찰리의 뒷모습 코트 위에 레퍼런스에 없는 거대한 기계 장갑판이 덧붙여져 의상이 심각하게 왜곡되었습니다.",
    "hard_violations": [
     "[gemini-pro] 찰리의 코트 등 부분에 존재하지 않는 전면부 기계 장갑판이 결합된 물리적으로 불가능한 의상 구조"
    ],
    "physics": "찰리의 오른발이 공중에 떠 있어 걷는 자세 자체는 땅의 지지를 받아 유지되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 뒷모습과 코트 구조를 레퍼런스에 맞게 정확히 구현했으며, 이전 샷의 바나나 가판대와 시장 환경을 충실히 재현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리의 코트 등 부분에 정면에 있어야 할 기계 장갑판이 잘못 생성되는 치명적인 의상 왜곡이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라를 등지고 시장 골목 안쪽을 향해 걷고 있으며, 군중들 사이로 향하고 있습니다.",
        "built_space": "야외 시장 골목으로, 우측에 이전 샷과 일치하는 바나나 가판대와 빨간색/파란색 파라솔이 배치되어 있습니다.",
        "entities": "찰리의 뒷모습(중절모, 벨트가 있는 코트, 로봇 팔다리)이 레퍼런스와 일치합니다. 군중과 바나나 가판대 모두 지시사항에 맞게 묘사되었습니다.",
        "hard_violations": [],
        "physics": "찰리의 오른발이 들려 있고 왼발이 땅을 지탱하여 자연스러운 걷기(mid-stride) 동작을 보여줍니다."
       },
       {
        "label": "B",
        "direction": "찰리는 카메라를 등지고 골목을 향해 걷고 있으나, 주변 군중 다수는 찰리 쪽을 향해 걸어오고 있습니다.",
        "built_space": "야외 시장 골목 환경이며, 우측에 과일과 바나나 가판대가 배치되어 있습니다.",
        "entities": "다인종 군중이 묘사되었으나, 찰리의 뒷모습 코트 위에 레퍼런스에 없는 거대한 기계 장갑판이 덧붙여져 의상이 심각하게 왜곡되었습니다.",
        "hard_violations": [
         "찰리의 코트 등 부분에 존재하지 않는 전면부 기계 장갑판이 결합된 물리적으로 불가능한 의상 구조"
        ],
        "physics": "찰리의 오른발이 공중에 떠 있어 걷는 자세 자체는 땅의 지지를 받아 유지되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 뒷모습과 코트 구조를 레퍼런스에 맞게 정확히 구현했으며, 이전 샷의 바나나 가판대와 시장 환경을 충실히 재현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리의 코트 등 부분에 정면에 있어야 할 기계 장갑판이 잘못 생성되는 치명적인 의상 왜곡이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라를 등지고 시장 골목 안쪽을 향해 걷고 있으며, 군중들 사이로 향하고 있습니다.",
        "built_space": "야외 시장 골목으로, 우측에 이전 샷과 일치하는 바나나 가판대와 빨간색/파란색 파라솔이 배치되어 있습니다.",
        "entities": "찰리의 뒷모습(중절모, 벨트가 있는 코트, 로봇 팔다리)이 레퍼런스와 일치합니다. 군중과 바나나 가판대 모두 지시사항에 맞게 묘사되었습니다.",
        "hard_violations": [],
        "physics": "찰리의 오른발이 들려 있고 왼발이 땅을 지탱하여 자연스러운 걷기(mid-stride) 동작을 보여줍니다."
       },
       {
        "label": "B",
        "direction": "찰리는 카메라를 등지고 골목을 향해 걷고 있으나, 주변 군중 다수는 찰리 쪽을 향해 걸어오고 있습니다.",
        "built_space": "야외 시장 골목 환경이며, 우측에 과일과 바나나 가판대가 배치되어 있습니다.",
        "entities": "다인종 군중이 묘사되었으나, 찰리의 뒷모습 코트 위에 레퍼런스에 없는 거대한 기계 장갑판이 덧붙여져 의상이 심각하게 왜곡되었습니다.",
        "hard_violations": [
         "찰리의 코트 등 부분에 존재하지 않는 전면부 기계 장갑판이 결합된 물리적으로 불가능한 의상 구조"
        ],
        "physics": "찰리의 오른발이 공중에 떠 있어 걷는 자세 자체는 땅의 지지를 받아 유지되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "거대한 뒷모습과 전진하는 한 걸음, 상단 중앙의 인파 속 목적지가 명확하며, 오른쪽 바나나 좌판과 적·청색 파라솔이 이전 장소의 연속성을 더 잘 보존한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "와이드 후면 구도와 보행은 충실하지만, 찰리 주변의 넓은 빈 공간이 군중 속으로 묻혀 들어가는 관계를 약화하고 등 장갑판과 좌판 주변 배치의 연속성도 B보다 떨어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 머리와 몸통을 카메라 반대쪽, 화면 상단 중앙으로 이어지는 시장 통로에 향하고 있다. 왼발보다 앞쪽에 놓인 오른발도 그 통로를 향한다. 주변 사람들은 카메라 쪽으로 오거나 반대 방향으로 이동하며 일부는 좌판을 본다. 찰리의 진행 목적지는 분명하지만 바로 앞뒤 공간이 비교적 넓게 비어 있다.",
        "built_space": "통로 양쪽에 금속 외벽 점포, 천막과 파라솔, 상자 위 상품 진열대가 여러 개 이어진다. 오른쪽 전경에는 바나나 진열대 한 곳과 그 뒤 감귤류 진열대가 있고, 오른쪽 점포 열을 따라 여러 바나나 송이가 걸려 있다. 이전 사진의 낡은 시장 재료와 색은 이어지지만, 바나나 좌판 바로 위의 큰 적·청색 파라솔 관계는 유지되지 않는다. 바닥도 이전 사진보다 건조하게 보인다. 사람들은 좌판 사이 통로에 서 있으며 불가능한 반사나 명백한 시설 중복은 보이지 않는다.",
        "entities": "찰리 한 명과 수많은 성인 남녀가 보인다. 군중에는 동아시아계로 보이는 사람들과 흑인으로 보이는 사람들이 섞여 있어 다민족 난민 시장 설정에 부합한다. 찰리는 갈색 챙모자, 낡은 올리브색 벨트 코트, 모래색 각진 장갑, 긴 팔과 짧은 기계 다리를 갖췄다. 얼굴은 후면 구도에 맞게 숨겨져 있다. 다만 등에 노출된 커다란 장갑판 두 장은 제공된 참조에서 확인되지 않는 디자인이다. 바나나는 좌판에 충분히 남아 있고 양손은 비어 있다. 명확히 읽히는 문구나 화면 위 그래픽은 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 왼발은 지면에 닿아 체중을 지지하고 오른발은 무릎이 굽혀진 채 앞쪽으로 들려 있어 전진 보행으로 성립한다. 팔과 손은 어깨 및 팔꿈치 관절에 연결되어 자연스럽게 내려와 있다. 군중은 보이는 발로 바닥을 딛고 있으며, 오른쪽 인물의 비닐봉지는 손에 들려 있다. 상품은 상자와 진열대에 놓이고 매달린 바나나는 위쪽 줄에 지지된다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리의 머리, 등과 전진하는 오른발이 모두 화면 상단 중앙의 군중 통로를 향한다. 다수의 주변 사람도 그 방향으로 이동하고 일부는 좌판이나 옆 사람을 보고 있어 진행 방향이 하나의 목적지로 이어진다. 찰리 바로 앞부터 사람이 촘촘히 들어차 있어 다음 걸음에 군중 사이로 들어갈 관계가 명확하다.",
        "built_space": "양쪽에 낡은 금속 점포와 천막이 줄지어 있고 중앙 통로가 상단 중앙으로 좁아진다. 오른쪽에는 큰 적·청색 파라솔 한 개, 그 아래 바나나 진열대 한 곳, 매달린 바나나 한 송이와 뒤쪽 감귤류 상자가 보인다. 이 주요 시설들의 상대적 배치가 이전 사진과 잘 이어진다. 상인들은 진열대 뒤에, 보행자는 통로에 있어 공간 사용도 타당하다. 바닥은 이전 사진보다 다소 건조하다. 불가능한 반사나 명백한 시설 중복은 없다.",
        "entities": "찰리 한 명과 다수의 성인 남녀가 보이며, 군중은 주로 동아시아계 외양이라 A보다 다민족성이 덜 드러난다. 찰리의 갈색 챙모자, 해진 올리브색 벨트 코트, 넓은 어깨 장갑, 모래색 전완과 기계 손, 짧고 육중한 다리가 참조의 정체성과 부합한다. 등은 코트로 덮여 있고 얼굴은 보이지 않아 후면 연출을 지킨다. 바나나는 걸린 송이와 진열된 여러 송이 모두 남아 있으며 찰리는 아무것도 들고 있지 않다. 읽을 수 있는 글자나 오버레이는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 왼발이 바닥에 놓여 몸을 지탱하고, 오른발은 뒤꿈치가 높고 앞쪽 끝이 지면 가까이 내려간 전진 착지 순간으로 보인다. 무릎과 발목의 굽힘이 보행 동작에 연결되며 몸 전체가 떠 있지 않다. 주변 보행자들도 지면에 서 있고 옷과 가방은 몸에 걸리거나 손에 들려 있다. 파라솔은 기둥으로, 바나나 진열은 상자와 좌판으로, 매달린 송이는 끈으로 지지된다. 지지 없는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "거대한 뒷모습과 전진하는 한 걸음, 상단 중앙의 인파 속 목적지가 명확하며, 오른쪽 바나나 좌판과 적·청색 파라솔이 이전 장소의 연속성을 더 잘 보존한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "와이드 후면 구도와 보행은 충실하지만, 찰리 주변의 넓은 빈 공간이 군중 속으로 묻혀 들어가는 관계를 약화하고 등 장갑판과 좌판 주변 배치의 연속성도 B보다 떨어진다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 머리와 몸통을 카메라 반대쪽, 화면 상단 중앙으로 이어지는 시장 통로에 향하고 있다. 왼발보다 앞쪽에 놓인 오른발도 그 통로를 향한다. 주변 사람들은 카메라 쪽으로 오거나 반대 방향으로 이동하며 일부는 좌판을 본다. 찰리의 진행 목적지는 분명하지만 바로 앞뒤 공간이 비교적 넓게 비어 있다.",
        "built_space": "통로 양쪽에 금속 외벽 점포, 천막과 파라솔, 상자 위 상품 진열대가 여러 개 이어진다. 오른쪽 전경에는 바나나 진열대 한 곳과 그 뒤 감귤류 진열대가 있고, 오른쪽 점포 열을 따라 여러 바나나 송이가 걸려 있다. 이전 사진의 낡은 시장 재료와 색은 이어지지만, 바나나 좌판 바로 위의 큰 적·청색 파라솔 관계는 유지되지 않는다. 바닥도 이전 사진보다 건조하게 보인다. 사람들은 좌판 사이 통로에 서 있으며 불가능한 반사나 명백한 시설 중복은 보이지 않는다.",
        "entities": "찰리 한 명과 수많은 성인 남녀가 보인다. 군중에는 동아시아계로 보이는 사람들과 흑인으로 보이는 사람들이 섞여 있어 다민족 난민 시장 설정에 부합한다. 찰리는 갈색 챙모자, 낡은 올리브색 벨트 코트, 모래색 각진 장갑, 긴 팔과 짧은 기계 다리를 갖췄다. 얼굴은 후면 구도에 맞게 숨겨져 있다. 다만 등에 노출된 커다란 장갑판 두 장은 제공된 참조에서 확인되지 않는 디자인이다. 바나나는 좌판에 충분히 남아 있고 양손은 비어 있다. 명확히 읽히는 문구나 화면 위 그래픽은 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 왼발은 지면에 닿아 체중을 지지하고 오른발은 무릎이 굽혀진 채 앞쪽으로 들려 있어 전진 보행으로 성립한다. 팔과 손은 어깨 및 팔꿈치 관절에 연결되어 자연스럽게 내려와 있다. 군중은 보이는 발로 바닥을 딛고 있으며, 오른쪽 인물의 비닐봉지는 손에 들려 있다. 상품은 상자와 진열대에 놓이고 매달린 바나나는 위쪽 줄에 지지된다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리의 머리, 등과 전진하는 오른발이 모두 화면 상단 중앙의 군중 통로를 향한다. 다수의 주변 사람도 그 방향으로 이동하고 일부는 좌판이나 옆 사람을 보고 있어 진행 방향이 하나의 목적지로 이어진다. 찰리 바로 앞부터 사람이 촘촘히 들어차 있어 다음 걸음에 군중 사이로 들어갈 관계가 명확하다.",
        "built_space": "양쪽에 낡은 금속 점포와 천막이 줄지어 있고 중앙 통로가 상단 중앙으로 좁아진다. 오른쪽에는 큰 적·청색 파라솔 한 개, 그 아래 바나나 진열대 한 곳, 매달린 바나나 한 송이와 뒤쪽 감귤류 상자가 보인다. 이 주요 시설들의 상대적 배치가 이전 사진과 잘 이어진다. 상인들은 진열대 뒤에, 보행자는 통로에 있어 공간 사용도 타당하다. 바닥은 이전 사진보다 다소 건조하다. 불가능한 반사나 명백한 시설 중복은 없다.",
        "entities": "찰리 한 명과 다수의 성인 남녀가 보이며, 군중은 주로 동아시아계 외양이라 A보다 다민족성이 덜 드러난다. 찰리의 갈색 챙모자, 해진 올리브색 벨트 코트, 넓은 어깨 장갑, 모래색 전완과 기계 손, 짧고 육중한 다리가 참조의 정체성과 부합한다. 등은 코트로 덮여 있고 얼굴은 보이지 않아 후면 연출을 지킨다. 바나나는 걸린 송이와 진열된 여러 송이 모두 남아 있으며 찰리는 아무것도 들고 있지 않다. 읽을 수 있는 글자나 오버레이는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 왼발이 바닥에 놓여 몸을 지탱하고, 오른발은 뒤꿈치가 높고 앞쪽 끝이 지면 가까이 내려간 전진 착지 순간으로 보인다. 무릎과 발목의 굽힘이 보행 동작에 연결되며 몸 전체가 떠 있지 않다. 주변 보행자들도 지면에 서 있고 옷과 가방은 몸에 걸리거나 손에 들려 있다. 파라솔은 기둥으로, 바나나 진열은 상자와 좌판으로, 매달린 송이는 끈으로 지지된다. 지지 없는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.317
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.067
   },
   "violations": {
    "B": [
     "[gemini-pro] 찰리의 코트 등 부분에 존재하지 않는 전면부 기계 장갑판이 결합된 물리적으로 불가능한 의상 구조"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1067
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "찰리의 뒷모습과 코트 구조를 레퍼런스에 맞게 정확히 구현했으며, 이전 샷의 바나나 가판대와 시장 환경을 충실히 재현했습니다."
   },
   {
    "label": "B",
    "score": 1067,
    "verdict_ko": "찰리의 코트 등 부분에 정면에 있어야 할 기계 장갑판이 잘못 생성되는 치명적인 의상 왜곡이 발생했습니다.  ★위반: [gemini-pro] 찰리의 코트 등 부분에 존재하지 않는 전면부 기계 장갑판이 결합된 물리적으로 불가능한 의상 구조"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3_sel.png",
    "asset_id": "01a89dd9-210b-408e-8e42-1ef9872a2f23",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-fbe2-7a1b-bc3c-c5a215a89159",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S22sh3"
  }
 },
 "S22sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:52:24.887539+00:00",
  "fingerprint": "c248540d1613b14eb1ca42956a3bfb02d2e8e008e561c9cd238de21720a4e1d2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S22sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S22sh7_sel.png",
  "source_sha256": "5a939feb1b404286569e8007523410e88a35c6c496fdfb3d8654fdadf36b1ab9",
  "file": "S22sh7_cine.png",
  "staged_sha256": "1da55d088477460bedaafbf54709b6200a0f5b910098b37fb030279cd9e5be57",
  "latency_ms": 10103
 },
 "S22sh9::signage": {
  "fp": "9dbf3f3e354c2eeb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S22sh9": {
  "input_fingerprint": "39a6d65d01aabf91",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 인파가 사라진 방향을 향해 검지손가락을 길게 뻗은 앰버의 단호한 상체.\n\nLOCATION (lock): At the market-lane junction near the banana stall, looking along the direction the robot vanished into the crowd. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Passage toward 찰리's route in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage in 찰리's direction (찰리 is no longer visible from this framing) — A limited section of the route extends toward the right edge beyond 앰버's pointing hand; used as Give the gesture directional meaning without placing a new object at its endpoint.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and restrained contrast consistent with the preceding market action.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the market passage, adjacent shop structures, and daytime lighting visible in the reference. Exclude the departing robot and do not duplicate the reference's passersby.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains in place, while Charlie has moved into the crowd with his coat-and-hat disguise unchanged. 앰버: She has reached the market search location, retaining her mask and waist tool pouch.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 인파가 사라진 방향을 향해 검지손가락을 길게 뻗은 앰버의 단호한 상체.\n\nLOCATION (lock): At the market-lane junction near the banana stall, looking along the direction the robot vanished into the crowd. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Passage toward 찰리's route in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage in 찰리's direction (찰리 is no longer visible from this framing) — A limited section of the route extends toward the right edge beyond 앰버's pointing hand; used as Give the gesture directional meaning without placing a new object at its endpoint.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and restrained contrast consistent with the preceding market action.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the market passage, adjacent shop structures, and daytime lighting visible in the reference. Exclude the departing robot and do not duplicate the reference's passersby.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains in place, while Charlie has moved into the crowd with his coat-and-hat disguise unchanged. 앰버: She has reached the market search location, retaining her mask and waist tool pouch.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 인파가 사라진 방향을 향해 검지손가락을 길게 뻗은 앰버의 단호한 상체.\n\nLOCATION (lock): At the market-lane junction near the banana stall, looking along the direction the robot vanished into the crowd. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Passage toward 찰리's route in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage in 찰리's direction (찰리 is no longer visible from this framing) — A limited section of the route extends toward the right edge beyond 앰버's pointing hand; used as Give the gesture directional meaning without placing a new object at its endpoint.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and restrained contrast consistent with the preceding market action.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the market passage, adjacent shop structures, and daytime lighting visible in the reference. Exclude the departing robot and do not duplicate the reference's passersby.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains in place, while Charlie has moved into the crowd with his coat-and-hat disguise unchanged. 앰버: She has reached the market search location, retaining her mask and waist tool pouch.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버의 시선과 길게 편 검지는 화면 오른쪽을 향한다. 그러나 군중이 이어지는 통로는 화면 중앙 깊숙이 뻗고, 검지의 연장선은 오른쪽 상점 앞 남성의 상체와 금속 외벽 쪽으로 향한다. 사라진 인파의 이동 경로를 지시하는 관계가 맞지 않는다.",
    "built_space": "양쪽에 낡은 상점과 천막이 있고 중앙에 시장 통로가 있다. 오른쪽에는 바나나 매대 하나, 상단에 매달린 바나나 한 묶음, 나무 진열 상자와 그 아래 플라스틱 상자 더미가 보인다. 앰버는 왼쪽 전경에 서 있다. 이전 사진의 오른쪽 바나나 매대와 시장 재료감은 비교적 잘 이어지지만, 손 너머 오른쪽 가장자리에는 통로 대신 상점과 매대가 자리한다.",
    "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 큰 눈, 남색 반소매, 갈색 작업 멜빵옷, 이마에 올린 마스크, 허리 공구집을 갖춰 앰버의 참조와 대체로 맞는다. 외형만으로 혼혈 배경을 확정할 수는 없다. 찰리는 보이지 않고 바나나 매대는 있다. 다만 앰버 외에 수십 명의 행인이 등장하여 인물 제한을 위반한다. 명확히 읽히는 글자는 보이지 않는다.",
    "hard_violations": [
     "앰버만 허용된 장면에 수십 명의 추가 행인과 상점 앞 인물을 생성했다."
    ],
    "physics": "뻗은 팔은 어깨와 팔꿈치에 자연스럽게 연결되고 검지와 접힌 손가락의 자세도 가능한 동작이다. 하체와 발의 접지는 프레임 밖이지만 상체가 떠 있다는 징후는 없다. 마스크는 머리끈으로, 공구집과 공구는 허리띠와 주머니로 지지된다. 진열 과일은 상자 위에 놓이고 매달린 바나나는 위쪽 끈에 걸려 있다."
   },
   {
    "label": "A",
    "direction": "검지는 화면 오른쪽 위로 뻗으며, 손 너머 중우측 배경에 빈 통로가 이어진다. 손끝이 통로 바닥보다는 그 위쪽 공간을 향하지만 A보다 이동 경로를 지시하는 의미가 잘 읽힌다. 시선은 손끝 방향보다 조금 더 카메라 쪽을 향해 시선과 지시 방향이 완전히 일치하지는 않는다.",
    "built_space": "앰버는 왼쪽 전경에 있고 시장 통로는 중우측으로 깊어진다. 왼쪽에 바나나 매대 하나와 매달린 바나나 한 묶음, 진열대를 받치는 상자 더미가 있으며 오른쪽에도 상품 진열대가 있다. 손 너머로 통로가 남는 배치는 요구에 가깝다. 다만 이전 사진의 오른쪽 바나나 매대와 빨강·파랑 파라솔 대신 왼쪽 매대와 다른 처마·셔터·끝 건물이 보여, 같은 장소라는 구체적 연속성은 약하다.",
    "entities": "금발의 어린 여자아이는 참조와 유사한 얼굴과 체격이며 남색 반소매, 낡은 갈색 멜빵옷, 이마의 마스크, 허리 공구집을 유지한다. 외형만으로 혼혈 배경을 확정할 수는 없다. 바나나 매대는 있고 찰리는 없다. 그러나 배경 통로와 오른쪽 매대 주변에 여러 성인 인물이 보여 앰버만 등장해야 한다는 조건을 위반한다. 간판은 있으나 글자는 흐려 명확히 판독되지 않는다.",
    "hard_violations": [
     "앰버만 허용된 장면의 배경 통로와 매대 주변에 여러 추가 인물을 생성했다."
    ],
    "physics": "팔을 들어 검지를 펴는 자세는 어깨·팔꿈치·손목의 연결상 가능하다. 몸통은 자연스럽게 세워져 있고 발은 구도 밖이라 접지를 직접 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 마스크는 머리끈에 고정되고 공구는 허리 주머니에 꽂혀 있다. 바나나와 상품은 진열대에 놓이며 매달린 바나나는 위쪽 줄로 지지된다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "앰버의 미디엄 숏과 의상은 부합하지만, 허용되지 않은 군중이 등장하고 검지가 통로가 아닌 오른쪽 상점 쪽을 가리킨다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "손 너머 중우측에 통로를 배치해 요구 구도에 더 가깝지만, 추가 인물들로 실격이며 이전 장소의 구체적인 연속성도 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선과 길게 편 검지는 화면 오른쪽을 향한다. 그러나 군중이 이어지는 통로는 화면 중앙 깊숙이 뻗고, 검지의 연장선은 오른쪽 상점 앞 남성의 상체와 금속 외벽 쪽으로 향한다. 사라진 인파의 이동 경로를 지시하는 관계가 맞지 않는다.",
        "built_space": "양쪽에 낡은 상점과 천막이 있고 중앙에 시장 통로가 있다. 오른쪽에는 바나나 매대 하나, 상단에 매달린 바나나 한 묶음, 나무 진열 상자와 그 아래 플라스틱 상자 더미가 보인다. 앰버는 왼쪽 전경에 서 있다. 이전 사진의 오른쪽 바나나 매대와 시장 재료감은 비교적 잘 이어지지만, 손 너머 오른쪽 가장자리에는 통로 대신 상점과 매대가 자리한다.",
        "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 큰 눈, 남색 반소매, 갈색 작업 멜빵옷, 이마에 올린 마스크, 허리 공구집을 갖춰 앰버의 참조와 대체로 맞는다. 외형만으로 혼혈 배경을 확정할 수는 없다. 찰리는 보이지 않고 바나나 매대는 있다. 다만 앰버 외에 수십 명의 행인이 등장하여 인물 제한을 위반한다. 명확히 읽히는 글자는 보이지 않는다.",
        "hard_violations": [
         "앰버만 허용된 장면에 수십 명의 추가 행인과 상점 앞 인물을 생성했다."
        ],
        "physics": "뻗은 팔은 어깨와 팔꿈치에 자연스럽게 연결되고 검지와 접힌 손가락의 자세도 가능한 동작이다. 하체와 발의 접지는 프레임 밖이지만 상체가 떠 있다는 징후는 없다. 마스크는 머리끈으로, 공구집과 공구는 허리띠와 주머니로 지지된다. 진열 과일은 상자 위에 놓이고 매달린 바나나는 위쪽 끈에 걸려 있다."
       },
       {
        "label": "B",
        "direction": "검지는 화면 오른쪽 위로 뻗으며, 손 너머 중우측 배경에 빈 통로가 이어진다. 손끝이 통로 바닥보다는 그 위쪽 공간을 향하지만 A보다 이동 경로를 지시하는 의미가 잘 읽힌다. 시선은 손끝 방향보다 조금 더 카메라 쪽을 향해 시선과 지시 방향이 완전히 일치하지는 않는다.",
        "built_space": "앰버는 왼쪽 전경에 있고 시장 통로는 중우측으로 깊어진다. 왼쪽에 바나나 매대 하나와 매달린 바나나 한 묶음, 진열대를 받치는 상자 더미가 있으며 오른쪽에도 상품 진열대가 있다. 손 너머로 통로가 남는 배치는 요구에 가깝다. 다만 이전 사진의 오른쪽 바나나 매대와 빨강·파랑 파라솔 대신 왼쪽 매대와 다른 처마·셔터·끝 건물이 보여, 같은 장소라는 구체적 연속성은 약하다.",
        "entities": "금발의 어린 여자아이는 참조와 유사한 얼굴과 체격이며 남색 반소매, 낡은 갈색 멜빵옷, 이마의 마스크, 허리 공구집을 유지한다. 외형만으로 혼혈 배경을 확정할 수는 없다. 바나나 매대는 있고 찰리는 없다. 그러나 배경 통로와 오른쪽 매대 주변에 여러 성인 인물이 보여 앰버만 등장해야 한다는 조건을 위반한다. 간판은 있으나 글자는 흐려 명확히 판독되지 않는다.",
        "hard_violations": [
         "앰버만 허용된 장면의 배경 통로와 매대 주변에 여러 추가 인물을 생성했다."
        ],
        "physics": "팔을 들어 검지를 펴는 자세는 어깨·팔꿈치·손목의 연결상 가능하다. 몸통은 자연스럽게 세워져 있고 발은 구도 밖이라 접지를 직접 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 마스크는 머리끈에 고정되고 공구는 허리 주머니에 꽂혀 있다. 바나나와 상품은 진열대에 놓이며 매달린 바나나는 위쪽 줄로 지지된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "앰버의 미디엄 숏과 의상은 부합하지만, 허용되지 않은 군중이 등장하고 검지가 통로가 아닌 오른쪽 상점 쪽을 가리킨다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "손 너머 중우측에 통로를 배치해 요구 구도에 더 가깝지만, 추가 인물들로 실격이며 이전 장소의 구체적인 연속성도 약하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 시선과 길게 편 검지는 화면 오른쪽을 향한다. 그러나 군중이 이어지는 통로는 화면 중앙 깊숙이 뻗고, 검지의 연장선은 오른쪽 상점 앞 남성의 상체와 금속 외벽 쪽으로 향한다. 사라진 인파의 이동 경로를 지시하는 관계가 맞지 않는다.",
        "built_space": "양쪽에 낡은 상점과 천막이 있고 중앙에 시장 통로가 있다. 오른쪽에는 바나나 매대 하나, 상단에 매달린 바나나 한 묶음, 나무 진열 상자와 그 아래 플라스틱 상자 더미가 보인다. 앰버는 왼쪽 전경에 서 있다. 이전 사진의 오른쪽 바나나 매대와 시장 재료감은 비교적 잘 이어지지만, 손 너머 오른쪽 가장자리에는 통로 대신 상점과 매대가 자리한다.",
        "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 큰 눈, 남색 반소매, 갈색 작업 멜빵옷, 이마에 올린 마스크, 허리 공구집을 갖춰 앰버의 참조와 대체로 맞는다. 외형만으로 혼혈 배경을 확정할 수는 없다. 찰리는 보이지 않고 바나나 매대는 있다. 다만 앰버 외에 수십 명의 행인이 등장하여 인물 제한을 위반한다. 명확히 읽히는 글자는 보이지 않는다.",
        "hard_violations": [
         "앰버만 허용된 장면에 수십 명의 추가 행인과 상점 앞 인물을 생성했다."
        ],
        "physics": "뻗은 팔은 어깨와 팔꿈치에 자연스럽게 연결되고 검지와 접힌 손가락의 자세도 가능한 동작이다. 하체와 발의 접지는 프레임 밖이지만 상체가 떠 있다는 징후는 없다. 마스크는 머리끈으로, 공구집과 공구는 허리띠와 주머니로 지지된다. 진열 과일은 상자 위에 놓이고 매달린 바나나는 위쪽 끈에 걸려 있다."
       },
       {
        "label": "A",
        "direction": "검지는 화면 오른쪽 위로 뻗으며, 손 너머 중우측 배경에 빈 통로가 이어진다. 손끝이 통로 바닥보다는 그 위쪽 공간을 향하지만 A보다 이동 경로를 지시하는 의미가 잘 읽힌다. 시선은 손끝 방향보다 조금 더 카메라 쪽을 향해 시선과 지시 방향이 완전히 일치하지는 않는다.",
        "built_space": "앰버는 왼쪽 전경에 있고 시장 통로는 중우측으로 깊어진다. 왼쪽에 바나나 매대 하나와 매달린 바나나 한 묶음, 진열대를 받치는 상자 더미가 있으며 오른쪽에도 상품 진열대가 있다. 손 너머로 통로가 남는 배치는 요구에 가깝다. 다만 이전 사진의 오른쪽 바나나 매대와 빨강·파랑 파라솔 대신 왼쪽 매대와 다른 처마·셔터·끝 건물이 보여, 같은 장소라는 구체적 연속성은 약하다.",
        "entities": "금발의 어린 여자아이는 참조와 유사한 얼굴과 체격이며 남색 반소매, 낡은 갈색 멜빵옷, 이마의 마스크, 허리 공구집을 유지한다. 외형만으로 혼혈 배경을 확정할 수는 없다. 바나나 매대는 있고 찰리는 없다. 그러나 배경 통로와 오른쪽 매대 주변에 여러 성인 인물이 보여 앰버만 등장해야 한다는 조건을 위반한다. 간판은 있으나 글자는 흐려 명확히 판독되지 않는다.",
        "hard_violations": [
         "앰버만 허용된 장면의 배경 통로와 매대 주변에 여러 추가 인물을 생성했다."
        ],
        "physics": "팔을 들어 검지를 펴는 자세는 어깨·팔꿈치·손목의 연결상 가능하다. 몸통은 자연스럽게 세워져 있고 발은 구도 밖이라 접지를 직접 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 마스크는 머리끈에 고정되고 공구는 허리 주머니에 꽂혀 있다. 바나나와 상품은 진열대에 놓이며 매달린 바나나는 위쪽 줄로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 2,
   "A": 4
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2,
    "verdict_ko": "앰버의 미디엄 숏과 의상은 부합하지만, 허용되지 않은 군중이 등장하고 검지가 통로가 아닌 오른쪽 상점 쪽을 가리킨다."
   },
   {
    "label": "A",
    "score": 4,
    "verdict_ko": "손 너머 중우측에 통로를 배치해 요구 구도에 더 가깝지만, 추가 인물들로 실격이며 이전 장소의 구체적인 연속성도 약하다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh7_sel.png",
    "asset_id": "5405dea4-bfbd-4f27-8bfa-751c14888566",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-fd98-758b-8106-f6e55aa9edf4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S22sh7"
  }
 },
 "S22sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:53:17.690417+00:00",
  "fingerprint": "9e1f09a492abd29f859d2f27d5fa7cf5e1a1c8cb8552065156346c6879415280",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S22sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S22sh9_sel.png",
  "source_sha256": "a8cfc114ecda436d98078a650dcb77524b11dfa914894045f61c16db801bfc56",
  "file": "S22sh9_cine.png",
  "staged_sha256": "930eebb0983b47a074280210805547dee27b950d060b2b9fc2c422096eaa38a8",
  "latency_ms": 13230
 },
 "S23sh1::signage": {
  "fp": "50585ca94203c8f1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::d817c9595eb14cf4": {
  "subjects": [],
  "subject_text": "구치소 면회실\n넓은 실내 한가운데 테이블과 의자가 놓인 면회 공간. 중앙 좌석 주변으로 여백이 넓게 남아 있는 단순한 구성이다.",
  "identity": "canonical",
  "scope_id": "L173",
  "scope_role": "location_interior",
  "scope_sha": "cf46c6ca176ce145"
 },
 "S23sh1::bgfirst_bg": {
  "input_fingerprint": "69263a57c81a657b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1__bgfirst_bg.png",
  "asset_id": "10a8564f-8f3e-46a6-8c49-5dd05488e9d9",
  "input_asset_ids": [
   "69e13ce4-f531-431c-a45d-c3679c68fd0e",
   "0977c2dd-72ef-4442-9b57-c56745b38ac5"
  ]
 },
 "S23sh1": {
  "input_fingerprint": "268030c7d58b4f2e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A table stands in the middle of the large visitation room. 현우: He sits at the table with his wrists still handcuffed and visible bruising on his face. His dog-bitten leg remains injured.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A table stands in the middle of the large visitation room. 현우: He sits at the table with his wrists still handcuffed and visible bruising on his face. His dog-bitten leg remains injured.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A table stands in the middle of the large visitation room. 현우: He sits at the table with his wrists still handcuffed and visible bruising on his face. His dog-bitten leg remains injured.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1__bgfirst_bg.png",
     "asset_id": "10a8564f-8f3e-46a6-8c49-5dd05488e9d9",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S23sh1.png",
     "asset_id": "69e13ce4-f531-431c-a45d-c3679c68fd0e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L173B01.png",
     "asset_id": "0977c2dd-72ef-4442-9b57-c56745b38ac5",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 화면 우측 전경에 있는 다른 인물을 향하고 있습니다.",
    "built_space": "면회실 내부에 테이블과 의자, 뒤쪽 창문과 왼쪽 문이 배치되어 있으나, 우측에 낯선 인물이 자리 잡고 있습니다.",
    "entities": "피멍이 든 현우가 수갑을 차고 있으나, 지시문에 없는 정장 차림의 인물이 프레임 우측에 추가되었습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (우측 전경)",
     "[gpt-high] 현우 외에 허용되지 않은 인물을 오른쪽 전경에 추가했다."
    ],
    "physics": "의자에 앉아 팔을 테이블에 기대고 있으며 물리적 지지는 자연스럽습니다."
   },
   {
    "label": "B",
    "direction": "시선이 앞쪽 아래를 향하며 절박하고 허망한 감정을 나타냅니다.",
    "built_space": "면회실 내부의 중앙 테이블, 뒤쪽 창문과 벽면 구조가 레퍼런스와 일치하며 넓은 여백이 확보되었습니다.",
    "entities": "피멍이 든 현우가 수갑을 찬 채 양손을 테이블 위로 길게 뻗고 있으며, 다른 인물은 없습니다.",
    "hard_violations": [],
    "physics": "의자에 앉아 체중을 실은 채 양손과 팔을 테이블 위에 안정적으로 지지하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 면회실 테이블과 프레이밍을 정확히 준수했으며, 피멍 든 현우가 수갑 찬 손을 뻗은 모습을 단독으로 완벽하게 연출했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 언급되지 않은 제3의 인물을 전경에 임의로 추가하여 심각한 지시 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 화면 우측 전경에 있는 다른 인물을 향하고 있습니다.",
        "built_space": "면회실 내부에 테이블과 의자, 뒤쪽 창문과 왼쪽 문이 배치되어 있으나, 우측에 낯선 인물이 자리 잡고 있습니다.",
        "entities": "피멍이 든 현우가 수갑을 차고 있으나, 지시문에 없는 정장 차림의 인물이 프레임 우측에 추가되었습니다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (우측 전경)"
        ],
        "physics": "의자에 앉아 팔을 테이블에 기대고 있으며 물리적 지지는 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "시선이 앞쪽 아래를 향하며 절박하고 허망한 감정을 나타냅니다.",
        "built_space": "면회실 내부의 중앙 테이블, 뒤쪽 창문과 벽면 구조가 레퍼런스와 일치하며 넓은 여백이 확보되었습니다.",
        "entities": "피멍이 든 현우가 수갑을 찬 채 양손을 테이블 위로 길게 뻗고 있으며, 다른 인물은 없습니다.",
        "hard_violations": [],
        "physics": "의자에 앉아 체중을 실은 채 양손과 팔을 테이블 위에 안정적으로 지지하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 면회실 테이블과 프레이밍을 정확히 준수했으며, 피멍 든 현우가 수갑 찬 손을 뻗은 모습을 단독으로 완벽하게 연출했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 언급되지 않은 제3의 인물을 전경에 임의로 추가하여 심각한 지시 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 화면 우측 전경에 있는 다른 인물을 향하고 있습니다.",
        "built_space": "면회실 내부에 테이블과 의자, 뒤쪽 창문과 왼쪽 문이 배치되어 있으나, 우측에 낯선 인물이 자리 잡고 있습니다.",
        "entities": "피멍이 든 현우가 수갑을 차고 있으나, 지시문에 없는 정장 차림의 인물이 프레임 우측에 추가되었습니다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (우측 전경)"
        ],
        "physics": "의자에 앉아 팔을 테이블에 기대고 있으며 물리적 지지는 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "시선이 앞쪽 아래를 향하며 절박하고 허망한 감정을 나타냅니다.",
        "built_space": "면회실 내부의 중앙 테이블, 뒤쪽 창문과 벽면 구조가 레퍼런스와 일치하며 넓은 여백이 확보되었습니다.",
        "entities": "피멍이 든 현우가 수갑을 찬 채 양손을 테이블 위로 길게 뻗고 있으며, 다른 인물은 없습니다.",
        "hard_violations": [],
        "physics": "의자에 앉아 체중을 실은 채 양손과 팔을 테이블 위에 안정적으로 지지하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우만 담은 미디엄 숏에서 피멍 든 얼굴과 테이블 위로 길게 뻗은 수갑 찬 양손을 충실하게 구현했으며, 배경 게시물 위치에는 참조와 차이가 있다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "수갑 찬 손과 절박한 표정은 적절하지만, 금지된 상대 인물을 전경에 추가하여 고립된 현우의 상체 숏을 어깨너머 구도로 바꾼 결정적 위반이 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽 바깥을 바라보며 입을 조금 벌리고 있다. 시선의 상대는 보이지 않으며, 프롬프트도 화면 안의 상대를 요구하지 않는다. 양팔은 몸에서 카메라 쪽 오른쪽 전경으로 길게 뻗어 있고, 손끝은 테이블 앞쪽을 향한다. 손에서 얼굴로 이어지는 사선이 분명하다.",
        "built_space": "중앙의 금속 테이블 한 개와 현우 뒤의 검은 의자 한 개가 보인다. 뒤벽에는 창살 창문 두 개, 그 사이의 돌출 기둥, 오른쪽 뒤에는 배관과 일부 가려진 설비 한 대가 보인다. 회색 하부 벽과 밝은 상부 벽, 마모된 바닥은 장소 참조와 부합한다. 현우는 테이블 건너편 의자에 앉아 있고, 카메라는 테이블 모서리와 상판을 비스듬히 본다. 주변 빈 바닥도 남아 있다. 오른쪽 위 게시물의 위치는 참조와 다르지만, 읽을 수 있는 문구는 확인되지 않는다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성 외형, 검은 머리, 마른 체격과 회색 셔츠가 인물 참조에 대체로 맞는다. 국적 자체는 외형으로 판별할 수 없다. 이마와 볼에 여러 피멍과 상처가 있으며, 양쪽 손목에는 짧은 사슬로 연결된 금속 수갑이 각각 채워져 있다. 금속 테이블과 검은 의자도 참조의 재질과 형태에 가깝다. 다리 부상은 올바른 상체 프레이밍 밖이므로 확인 대상이 아니다.",
        "hard_violations": [],
        "physics": "등 뒤로 보이는 의자가 앉은 몸을 지지하고, 앞으로 기울인 상체에서 이어지는 양팔과 손바닥은 상판에 닿아 있다. 수갑은 손목을 둘러싸고 연결 사슬은 두 손목 사이 상판 가까이에 놓인다. 손가락과 팔의 연결이 자연스럽고, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 전경에 보이는 상대 인물의 얼굴 쪽을 올려다본다. 양팔은 테이블 건너 상대 쪽으로 뻗어 있으며 손바닥을 위로 펴 호소하는 동작이다. 상대 인물도 현우 쪽으로 얼굴을 돌리고 있다. 호소의 방향은 명확하지만, 그 대상 인물을 화면 안에 등장시킨 것은 허용된 인물 구성과 어긋난다.",
        "built_space": "금속 테이블 한 개, 현우가 앉은 검은 의자 한 개, 왼쪽 금속문 한 개, 뒤벽 창살 창문 두 개, 오른쪽 설비 한 대와 배관이 보인다. 벽의 두 가지 색과 넓은 빈 바닥은 장소 참조와 잘 맞는다. 테이블 상판과 모서리도 비스듬히 보이지만, 오른쪽 가까이에 추가 인물이 크게 걸쳐 현우 단독의 고립감을 가리고 어깨너머 구도를 만든다.",
        "entities": "현우는 젊은 동아시아계 남성으로 보이며, 헝클어진 검은 머리와 회색 셔츠, 얼굴 곳곳의 피멍이 요구에 부합한다. 양쪽 손목의 금속 수갑과 연결 사슬도 보인다. 다리 부상은 프레임 밖이다. 그러나 오른쪽 전경에 검은 머리와 어두운 옷을 가진 별도의 인물이 등장하며, 이는 현우만 허용한 인물 조건을 위반한다.",
        "hard_violations": [
         "현우 외에 허용되지 않은 인물을 오른쪽 전경에 추가했다."
        ],
        "physics": "현우의 몸은 검은 의자에 지지되고 양팔은 테이블에 걸쳐 있다. 위로 편 손바닥은 손목과 팔로 지지되며, 수갑 사슬은 두 손목 사이에 자연스럽게 늘어진다. 추가 인물은 상체 일부만 보여 하체 지지는 확인할 수 없지만, 공중에 떠 있다고 볼 근거는 없다. 현우의 자세 자체에는 뚜렷한 물리적 불가능성이 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우만 담은 미디엄 숏에서 피멍 든 얼굴과 테이블 위로 길게 뻗은 수갑 찬 양손을 충실하게 구현했으며, 배경 게시물 위치에는 참조와 차이가 있다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "수갑 찬 손과 절박한 표정은 적절하지만, 금지된 상대 인물을 전경에 추가하여 고립된 현우의 상체 숏을 어깨너머 구도로 바꾼 결정적 위반이 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽 바깥을 바라보며 입을 조금 벌리고 있다. 시선의 상대는 보이지 않으며, 프롬프트도 화면 안의 상대를 요구하지 않는다. 양팔은 몸에서 카메라 쪽 오른쪽 전경으로 길게 뻗어 있고, 손끝은 테이블 앞쪽을 향한다. 손에서 얼굴로 이어지는 사선이 분명하다.",
        "built_space": "중앙의 금속 테이블 한 개와 현우 뒤의 검은 의자 한 개가 보인다. 뒤벽에는 창살 창문 두 개, 그 사이의 돌출 기둥, 오른쪽 뒤에는 배관과 일부 가려진 설비 한 대가 보인다. 회색 하부 벽과 밝은 상부 벽, 마모된 바닥은 장소 참조와 부합한다. 현우는 테이블 건너편 의자에 앉아 있고, 카메라는 테이블 모서리와 상판을 비스듬히 본다. 주변 빈 바닥도 남아 있다. 오른쪽 위 게시물의 위치는 참조와 다르지만, 읽을 수 있는 문구는 확인되지 않는다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성 외형, 검은 머리, 마른 체격과 회색 셔츠가 인물 참조에 대체로 맞는다. 국적 자체는 외형으로 판별할 수 없다. 이마와 볼에 여러 피멍과 상처가 있으며, 양쪽 손목에는 짧은 사슬로 연결된 금속 수갑이 각각 채워져 있다. 금속 테이블과 검은 의자도 참조의 재질과 형태에 가깝다. 다리 부상은 올바른 상체 프레이밍 밖이므로 확인 대상이 아니다.",
        "hard_violations": [],
        "physics": "등 뒤로 보이는 의자가 앉은 몸을 지지하고, 앞으로 기울인 상체에서 이어지는 양팔과 손바닥은 상판에 닿아 있다. 수갑은 손목을 둘러싸고 연결 사슬은 두 손목 사이 상판 가까이에 놓인다. 손가락과 팔의 연결이 자연스럽고, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 전경에 보이는 상대 인물의 얼굴 쪽을 올려다본다. 양팔은 테이블 건너 상대 쪽으로 뻗어 있으며 손바닥을 위로 펴 호소하는 동작이다. 상대 인물도 현우 쪽으로 얼굴을 돌리고 있다. 호소의 방향은 명확하지만, 그 대상 인물을 화면 안에 등장시킨 것은 허용된 인물 구성과 어긋난다.",
        "built_space": "금속 테이블 한 개, 현우가 앉은 검은 의자 한 개, 왼쪽 금속문 한 개, 뒤벽 창살 창문 두 개, 오른쪽 설비 한 대와 배관이 보인다. 벽의 두 가지 색과 넓은 빈 바닥은 장소 참조와 잘 맞는다. 테이블 상판과 모서리도 비스듬히 보이지만, 오른쪽 가까이에 추가 인물이 크게 걸쳐 현우 단독의 고립감을 가리고 어깨너머 구도를 만든다.",
        "entities": "현우는 젊은 동아시아계 남성으로 보이며, 헝클어진 검은 머리와 회색 셔츠, 얼굴 곳곳의 피멍이 요구에 부합한다. 양쪽 손목의 금속 수갑과 연결 사슬도 보인다. 다리 부상은 프레임 밖이다. 그러나 오른쪽 전경에 검은 머리와 어두운 옷을 가진 별도의 인물이 등장하며, 이는 현우만 허용한 인물 조건을 위반한다.",
        "hard_violations": [
         "현우 외에 허용되지 않은 인물을 오른쪽 전경에 추가했다."
        ],
        "physics": "현우의 몸은 검은 의자에 지지되고 양팔은 테이블에 걸쳐 있다. 위로 편 손바닥은 손목과 팔로 지지되며, 수갑 사슬은 두 손목 사이에 자연스럽게 늘어진다. 추가 인물은 상체 일부만 보여 하체 지지는 확인할 수 없지만, 공중에 떠 있다고 볼 근거는 없다. 현우의 자세 자체에는 뚜렷한 물리적 불가능성이 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.651,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.401,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (우측 전경)",
     "[gpt-high] 현우 외에 허용되지 않은 인물을 오른쪽 전경에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 401
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 면회실 테이블과 프레이밍을 정확히 준수했으며, 피멍 든 현우가 수갑 찬 손을 뻗은 모습을 단독으로 완벽하게 연출했습니다."
   },
   {
    "label": "A",
    "score": 401,
    "verdict_ko": "프롬프트에 언급되지 않은 제3의 인물을 전경에 임의로 추가하여 심각한 지시 위반이 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (우측 전경) / [gpt-high] 현우 외에 허용되지 않은 인물을 오른쪽 전경에 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L173B01.png",
    "asset_id": "0977c2dd-72ef-4442-9b57-c56745b38ac5",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-ff43-72eb-a8b3-673ef3398586",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1__bgfirst_bg.png",
   "bg_asset_id": "10a8564f-8f3e-46a6-8c49-5dd05488e9d9",
   "bg_record_key": "S23sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S23sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:54:16.367542+00:00",
  "fingerprint": "6280a1c5b2eeb4d6a9c59d212b4f5ab36c1af2c3a4a898fd828b64f9eb0b74ac",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S23sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S23sh1_sel.png",
  "source_sha256": "f3591983c462664da33ac2313b6cac359781112c5698e1fedca4136572892249",
  "file": "S23sh1_cine.png",
  "staged_sha256": "a7c420c0bb54fec8d4971d3988269f302c2236370ee8f3d863e8c650ceb38b5f",
  "latency_ms": 10269
 },
 "S23sh6::signage": {
  "fp": "7dd4d8c93e26b70b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S23sh6": {
  "input_fingerprint": "14e8d0400f4e678c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 테이블 위에 놓인 찰리의 사진을 불안한 기색으로 내려다보는 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the central visiting table inside the spacious detention-center room, with a robot photograph visible in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리 수배 사진 (Lying on the table, partially visible at the lower frame edge) — The image-bearing face is angled upward toward the camera, showing part of 찰리's likeness without introducing additional readable text; used as Provides the concrete cause of 현우's lowered gaze while remaining subordinate to his face; 면회실 테이블 (Supporting the photograph) — Only a narrow oblique strip of the upper surface is visible; used as Maintains the table-side spatial continuity of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding restrained ambient illumination, allowing the lowered eyes and bruises to remain legible without a new lighting accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same visitation table and the surrounding room surfaces and lighting. Exclude the earlier handcuffs, which have now been removed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph depicts Charlie; it is the same photograph brought into the visitation room. 현우: His wrists are now free of handcuffs. He remains at the table with facial bruising and the persistent dog-bite leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 테이블 위에 놓인 찰리의 사진을 불안한 기색으로 내려다보는 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the central visiting table inside the spacious detention-center room, with a robot photograph visible in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리 수배 사진 (Lying on the table, partially visible at the lower frame edge) — The image-bearing face is angled upward toward the camera, showing part of 찰리's likeness without introducing additional readable text; used as Provides the concrete cause of 현우's lowered gaze while remaining subordinate to his face; 면회실 테이블 (Supporting the photograph) — Only a narrow oblique strip of the upper surface is visible; used as Maintains the table-side spatial continuity of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding restrained ambient illumination, allowing the lowered eyes and bruises to remain legible without a new lighting accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same visitation table and the surrounding room surfaces and lighting. Exclude the earlier handcuffs, which have now been removed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph depicts Charlie; it is the same photograph brought into the visitation room. 현우: His wrists are now free of handcuffs. He remains at the table with facial bruising and the persistent dog-bite leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 테이블 위에 놓인 찰리의 사진을 불안한 기색으로 내려다보는 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the central visiting table inside the spacious detention-center room, with a robot photograph visible in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리 수배 사진 (Lying on the table, partially visible at the lower frame edge) — The image-bearing face is angled upward toward the camera, showing part of 찰리's likeness without introducing additional readable text; used as Provides the concrete cause of 현우's lowered gaze while remaining subordinate to his face; 면회실 테이블 (Supporting the photograph) — Only a narrow oblique strip of the upper surface is visible; used as Maintains the table-side spatial continuity of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding restrained ambient illumination, allowing the lowered eyes and bruises to remain legible without a new lighting accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same visitation table and the surrounding room surfaces and lighting. Exclude the earlier handcuffs, which have now been removed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph depicts Charlie; it is the same photograph brought into the visitation room. 현우: His wrists are now free of handcuffs. He remains at the table with facial bruising and the persistent dog-bite leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 테이블 위의 사진을 명확히 향하고 있음.",
    "built_space": "투톤 벽면과 창문, 앞쪽의 금속 테이블이 설정과 맞게 배치됨.",
    "entities": "현우와 사진 속 찰리의 모습이 레퍼런스의 특징을 잘 반영함.",
    "hard_violations": [
     "[gemini-pro] 사진 하단에 읽을 수 있는 텍스트('WANTED', 뒤집힘)가 포함됨 (텍스트 금지 조건 위반)."
    ],
    "physics": "사진은 테이블 위에 놓여 있음."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 테이블 위에 놓인 찰리의 사진을 향하고 있음.",
    "built_space": "배경의 창문과 투톤 벽면, 앞쪽의 좁은 금속 테이블 상판이 이전 샷과 일치하게 배치됨.",
    "entities": "현우(멍든 얼굴, 헝클어진 머리, 회색 셔츠)와 사진 속 찰리(로봇 형태, 코트) 모두 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "사진은 테이블 상판 위에 안정적으로 놓여 있으며 현우의 자세도 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "읽을 수 있는 텍스트 없이 프레이밍, 캐릭터의 외양 및 소품의 방향을 지시대로 정확히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "텍스트 금지 지시를 위반하여 사진에 'WANTED' 글자가 포함됨."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선은 테이블 위에 놓인 찰리의 사진을 향하고 있음.",
        "built_space": "배경의 창문과 투톤 벽면, 앞쪽의 좁은 금속 테이블 상판이 이전 샷과 일치하게 배치됨.",
        "entities": "현우(멍든 얼굴, 헝클어진 머리, 회색 셔츠)와 사진 속 찰리(로봇 형태, 코트) 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "사진은 테이블 상판 위에 안정적으로 놓여 있으며 현우의 자세도 자연스러움."
       },
       {
        "label": "A",
        "direction": "현우의 시선은 테이블 위의 사진을 명확히 향하고 있음.",
        "built_space": "투톤 벽면과 창문, 앞쪽의 금속 테이블이 설정과 맞게 배치됨.",
        "entities": "현우와 사진 속 찰리의 모습이 레퍼런스의 특징을 잘 반영함.",
        "hard_violations": [
         "사진 하단에 읽을 수 있는 텍스트('WANTED', 뒤집힘)가 포함됨 (텍스트 금지 조건 위반)."
        ],
        "physics": "사진은 테이블 위에 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "읽을 수 있는 텍스트 없이 프레이밍, 캐릭터의 외양 및 소품의 방향을 지시대로 정확히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "텍스트 금지 지시를 위반하여 사진에 'WANTED' 글자가 포함됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선은 테이블 위에 놓인 찰리의 사진을 향하고 있음.",
        "built_space": "배경의 창문과 투톤 벽면, 앞쪽의 좁은 금속 테이블 상판이 이전 샷과 일치하게 배치됨.",
        "entities": "현우(멍든 얼굴, 헝클어진 머리, 회색 셔츠)와 사진 속 찰리(로봇 형태, 코트) 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "사진은 테이블 상판 위에 안정적으로 놓여 있으며 현우의 자세도 자연스러움."
       },
       {
        "label": "A",
        "direction": "현우의 시선은 테이블 위의 사진을 명확히 향하고 있음.",
        "built_space": "투톤 벽면과 창문, 앞쪽의 금속 테이블이 설정과 맞게 배치됨.",
        "entities": "현우와 사진 속 찰리의 모습이 레퍼런스의 특징을 잘 반영함.",
        "hard_violations": [
         "사진 하단에 읽을 수 있는 텍스트('WANTED', 뒤집힘)가 포함됨 (텍스트 금지 조건 위반)."
        ],
        "physics": "사진은 테이블 위에 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "현우의 외형과 장소는 이어지지만, 시선이 사진보다 전방으로 읽히고 사진과 상판을 너무 크게 보여 얼굴 중심의 지정 구도를 약화한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "아래로 향한 눈과 긴장한 얼굴이 사진을 보고 불안해하는 순간에 더 가깝지만, 상판과 사진의 노출량은 여전히 지정 구도보다 많다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개를 숙였지만 눈은 비교적 전방을 향한다. 화면 아래 사진을 주시한다는 연결이 약하고 사진 너머를 보는 인상도 있다. 사진은 인쇄면이 위를 향하며, 찰리의 머리가 카메라 쪽에 있어 현우에게는 바로 선 방향이다.",
        "built_space": "금속 테이블 한 개가 전경을 차지하고, 뒤에는 밝은 상부와 회색 하부로 나뉜 벽 및 두 군데의 창 영역이 보인다. 기존 면회실의 재질과 낮 조명은 대체로 이어진다. 현우는 테이블 반대편에서 몸을 낮췄다. 상판이 화면 아래 약 3분의 1을 차지해 요청한 좁고 비스듬한 띠가 아니며, 의자와 창살은 이 구도에서 확인하기 어렵다.",
        "entities": "보이는 사람은 현우 한 명이다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 회색 셔츠는 참고와 대체로 맞는다. 얼굴 멍은 유지되지만 코와 양 볼의 상처 분포는 이전 장면과 다르다. 사진 한 장에는 모자와 외투, 베이지 장갑판을 갖춘 찰리가 보이며 실물 로봇이 추가되지는 않았다. 읽을 수 있는 문자는 없다. 손목과 다리는 화면 밖이므로 수갑 제거와 다리 부상은 확인할 수 없다.",
        "hard_violations": [],
        "physics": "사진은 금속 상판에 평평하게 놓여 지지된다. 현우의 머리는 목과 상체에 자연스럽게 연결되어 있고 몸을 앞으로 낮춘 자세는 가능하다. 좌석과 하체의 지지점은 화면 밖이지만, 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 눈동자는 아래쪽 전경의 사진 방향으로 내려가 있으며, 모인 눈썹과 굳은 입이 불안을 드러낸다. 사진을 내려다보는 관계가 A보다 명확하다. 사진의 인쇄면은 위를 향하고 찰리의 머리는 카메라 쪽에 있어 현우가 볼 방향과 양립한다.",
        "built_space": "금속 테이블 한 개, 투톤 벽, 왼쪽 끝·중앙·오른쪽의 세 창 영역과 벽체 돌출부가 보인다. 낮의 확산광과 낡은 금속 표면은 이전 면회실과 유사하다. 추가 창 영역이 이전 화면 밖 공간인지 확정할 수 없다. 현우는 테이블 반대편에 있으며 얼굴을 크게 잡았지만, 상판은 여전히 화면 하단 약 3분의 1을 차지하고 거의 수평이어서 지정한 좁은 사선 띠와 다르다.",
        "entities": "현우 한 명만 보이며 젊은 동아시아계 남성의 얼굴, 검은 머리, 회색 셔츠가 참고와 대체로 일치한다. 얼굴 멍은 있으나 상처의 세부 위치는 이전 장면과 달라졌다. 테이블 위 한 장의 사진에는 모자·외투·큰 장갑 팔을 가진 찰리가 보인다. 사진은 흑백 인쇄처럼 보여 베이지색과 흰 마스크의 구분이 약하다. 맨 아래에는 잘린 굵은 문자 모양이 있으나 완전한 단어를 판독하기 어렵다. 손목과 다리는 보이지 않는다.",
        "hard_violations": [],
        "physics": "사진은 테이블 상판이 받치며 가장자리의 미세한 굴곡도 종이로서 가능하다. 현우는 목을 숙이고 상체를 앞으로 기울였으며 불가능한 관절이나 부유는 보이지 않는다. 의자와 하체의 지지점은 클로즈업 밖이다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "현우의 외형과 장소는 이어지지만, 시선이 사진보다 전방으로 읽히고 사진과 상판을 너무 크게 보여 얼굴 중심의 지정 구도를 약화한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "아래로 향한 눈과 긴장한 얼굴이 사진을 보고 불안해하는 순간에 더 가깝지만, 상판과 사진의 노출량은 여전히 지정 구도보다 많다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개를 숙였지만 눈은 비교적 전방을 향한다. 화면 아래 사진을 주시한다는 연결이 약하고 사진 너머를 보는 인상도 있다. 사진은 인쇄면이 위를 향하며, 찰리의 머리가 카메라 쪽에 있어 현우에게는 바로 선 방향이다.",
        "built_space": "금속 테이블 한 개가 전경을 차지하고, 뒤에는 밝은 상부와 회색 하부로 나뉜 벽 및 두 군데의 창 영역이 보인다. 기존 면회실의 재질과 낮 조명은 대체로 이어진다. 현우는 테이블 반대편에서 몸을 낮췄다. 상판이 화면 아래 약 3분의 1을 차지해 요청한 좁고 비스듬한 띠가 아니며, 의자와 창살은 이 구도에서 확인하기 어렵다.",
        "entities": "보이는 사람은 현우 한 명이다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 회색 셔츠는 참고와 대체로 맞는다. 얼굴 멍은 유지되지만 코와 양 볼의 상처 분포는 이전 장면과 다르다. 사진 한 장에는 모자와 외투, 베이지 장갑판을 갖춘 찰리가 보이며 실물 로봇이 추가되지는 않았다. 읽을 수 있는 문자는 없다. 손목과 다리는 화면 밖이므로 수갑 제거와 다리 부상은 확인할 수 없다.",
        "hard_violations": [],
        "physics": "사진은 금속 상판에 평평하게 놓여 지지된다. 현우의 머리는 목과 상체에 자연스럽게 연결되어 있고 몸을 앞으로 낮춘 자세는 가능하다. 좌석과 하체의 지지점은 화면 밖이지만, 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 눈동자는 아래쪽 전경의 사진 방향으로 내려가 있으며, 모인 눈썹과 굳은 입이 불안을 드러낸다. 사진을 내려다보는 관계가 A보다 명확하다. 사진의 인쇄면은 위를 향하고 찰리의 머리는 카메라 쪽에 있어 현우가 볼 방향과 양립한다.",
        "built_space": "금속 테이블 한 개, 투톤 벽, 왼쪽 끝·중앙·오른쪽의 세 창 영역과 벽체 돌출부가 보인다. 낮의 확산광과 낡은 금속 표면은 이전 면회실과 유사하다. 추가 창 영역이 이전 화면 밖 공간인지 확정할 수 없다. 현우는 테이블 반대편에 있으며 얼굴을 크게 잡았지만, 상판은 여전히 화면 하단 약 3분의 1을 차지하고 거의 수평이어서 지정한 좁은 사선 띠와 다르다.",
        "entities": "현우 한 명만 보이며 젊은 동아시아계 남성의 얼굴, 검은 머리, 회색 셔츠가 참고와 대체로 일치한다. 얼굴 멍은 있으나 상처의 세부 위치는 이전 장면과 달라졌다. 테이블 위 한 장의 사진에는 모자·외투·큰 장갑 팔을 가진 찰리가 보인다. 사진은 흑백 인쇄처럼 보여 베이지색과 흰 마스크의 구분이 약하다. 맨 아래에는 잘린 굵은 문자 모양이 있으나 완전한 단어를 판독하기 어렵다. 손목과 다리는 보이지 않는다.",
        "hard_violations": [],
        "physics": "사진은 테이블 상판이 받치며 가장자리의 미세한 굴곡도 종이로서 가능하다. 현우는 목을 숙이고 상체를 앞으로 기울였으며 불가능한 관절이나 부유는 보이지 않는다. 의자와 하체의 지지점은 클로즈업 밖이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 1.833
   },
   "adjusted": {
    "A": 1.125,
    "B": 1.833
   },
   "violations": {
    "A": [
     "[gemini-pro] 사진 하단에 읽을 수 있는 텍스트('WANTED', 뒤집힘)가 포함됨 (텍스트 금지 조건 위반)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1833,
   "A": 1125
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1833,
    "verdict_ko": "읽을 수 있는 텍스트 없이 프레이밍, 캐릭터의 외양 및 소품의 방향을 지시대로 정확히 구현함."
   },
   {
    "label": "A",
    "score": 1125,
    "verdict_ko": "텍스트 금지 지시를 위반하여 사진에 'WANTED' 글자가 포함됨.  ★위반: [gemini-pro] 사진 하단에 읽을 수 있는 텍스트('WANTED', 뒤집힘)가 포함됨 (텍스트 금지 조건 위반)."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1_sel.png",
    "asset_id": "07120031-3748-44de-a3b2-3fbdb38dd6fd",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-0288-75f4-a360-094e8cee05fb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S23sh1"
  },
  "staged_characters_added": [
   "C06"
  ]
 },
 "S23sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:55:34.750360+00:00",
  "fingerprint": "cf358e70bcbc53662c00fa61ed2a73079c6e404c3a6f02c399dc1ead57911288",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S23sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S23sh6_sel.png",
  "source_sha256": "880130d3b18425885bb343b219a74f03d58d88756e8b7c6ede84b56ac49726d3",
  "file": "S23sh6_cine.png",
  "staged_sha256": "412b358edd106a288baa4cd87053813c10401d4802388ef5788c2d6ebeaa90f3",
  "latency_ms": 10373
 },
 "S23sh11::signage": {
  "fp": "5fa6b125f39d0b83",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S23sh11": {
  "input_fingerprint": "87ec438eff256211",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 마주 보며 팽팽한 긴장감 속에 마주 앉아 있는 현우와 윤성찬의 옆모습 풀샷.\n\nLOCATION (lock): Across the central table inside the spacious detention-center visiting room, under ordinary interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Separating the two seated men) — Its side and a shallow portion of its top are visible between the opposing profiles; used as Defines the negotiation axis and the physical separation; 두 사람의 좌석 (Occupied on opposite sides of the table) — Viewed from the side, with both seated bodies and feet unobscured; used as Makes the different seated weight distributions readable; 넓은 면회실 (Visible around the central seating arrangement); used as Leaves austere breathing room around the intimate negotiation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same controlled ambient contrast across both figures and the surrounding visiting room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the visitation table, nearby room surfaces, and consistent daytime interior lighting. Exclude the handcuffs from the earlier arrival state.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph of Charlie remains available in the visitation room. 현우: He remains at the table with his handcuffs removed, his face bruised and his dog-bitten leg still injured. 윤성찬: He remains in his suit during the negotiation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 마주 보며 팽팽한 긴장감 속에 마주 앉아 있는 현우와 윤성찬의 옆모습 풀샷.\n\nLOCATION (lock): Across the central table inside the spacious detention-center visiting room, under ordinary interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Separating the two seated men) — Its side and a shallow portion of its top are visible between the opposing profiles; used as Defines the negotiation axis and the physical separation; 두 사람의 좌석 (Occupied on opposite sides of the table) — Viewed from the side, with both seated bodies and feet unobscured; used as Makes the different seated weight distributions readable; 넓은 면회실 (Visible around the central seating arrangement); used as Leaves austere breathing room around the intimate negotiation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same controlled ambient contrast across both figures and the surrounding visiting room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the visitation table, nearby room surfaces, and consistent daytime interior lighting. Exclude the handcuffs from the earlier arrival state.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph of Charlie remains available in the visitation room. 현우: He remains at the table with his handcuffs removed, his face bruised and his dog-bitten leg still injured. 윤성찬: He remains in his suit during the negotiation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 마주 보며 팽팽한 긴장감 속에 마주 앉아 있는 현우와 윤성찬의 옆모습 풀샷.\n\nLOCATION (lock): Across the central table inside the spacious detention-center visiting room, under ordinary interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Separating the two seated men) — Its side and a shallow portion of its top are visible between the opposing profiles; used as Defines the negotiation axis and the physical separation; 두 사람의 좌석 (Occupied on opposite sides of the table) — Viewed from the side, with both seated bodies and feet unobscured; used as Makes the different seated weight distributions readable; 넓은 면회실 (Visible around the central seating arrangement); used as Leaves austere breathing room around the intimate negotiation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same controlled ambient contrast across both figures and the surrounding visiting room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the visitation table, nearby room surfaces, and consistent daytime interior lighting. Exclude the handcuffs from the earlier arrival state.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph of Charlie remains available in the visitation room. 현우: He remains at the table with his handcuffs removed, his face bruised and his dog-bitten leg still injured. 윤성찬: He remains in his suit during the negotiation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 테이블을 사이에 두고 서로의 얼굴을 똑바로 마주 봄.",
    "built_space": "창문이 있는 넓은 면회실 내부로, 다수의 철제 테이블과 의자가 배치되어 있으며 중앙 테이블에 인물들이 위치함.",
    "entities": "얼굴에 상처가 있는 현우와 정장 차림의 윤성찬이 참조와 일치하며, 테이블 위 사진이 존재하고 수갑은 없음.",
    "hard_violations": [],
    "physics": "두 인물 모두 의자에 안정적으로 착석하여 발이 바닥에 닿아 있으며, 손은 테이블 위에 자연스럽게 놓임."
   },
   {
    "label": "B",
    "direction": "두 인물이 중앙 테이블을 사이에 두고 서로를 주시함.",
    "built_space": "면회실 내부의 중앙 테이블에 두 인물이 마주 앉아 있으며 주변으로 여백의 공간이 보임.",
    "entities": "두 인물의 외모와 복장은 참조와 일치하나, 제거되어야 할 수갑이 현우의 손목에 채워져 있음.",
    "hard_violations": [
     "[gemini-pro] 수갑을 제거하라는 지시를 어기고 현우의 손목에 수갑이 묘사됨."
    ],
    "physics": "인물들은 의자에 앉아 체중을 지탱하고 있으며 발은 바닥을 딛고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "수갑이 제거된 상태와 두 인물의 측면 풀샷 구도를 비롯한 프롬프트의 지시사항과 참조 이미지를 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "명시적인 지시에도 불구하고 현우의 손목에 수갑을 그려 넣어 핵심 요구사항을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 테이블을 사이에 두고 서로의 얼굴을 똑바로 마주 봄.",
        "built_space": "창문이 있는 넓은 면회실 내부로, 다수의 철제 테이블과 의자가 배치되어 있으며 중앙 테이블에 인물들이 위치함.",
        "entities": "얼굴에 상처가 있는 현우와 정장 차림의 윤성찬이 참조와 일치하며, 테이블 위 사진이 존재하고 수갑은 없음.",
        "hard_violations": [],
        "physics": "두 인물 모두 의자에 안정적으로 착석하여 발이 바닥에 닿아 있으며, 손은 테이블 위에 자연스럽게 놓임."
       },
       {
        "label": "B",
        "direction": "두 인물이 중앙 테이블을 사이에 두고 서로를 주시함.",
        "built_space": "면회실 내부의 중앙 테이블에 두 인물이 마주 앉아 있으며 주변으로 여백의 공간이 보임.",
        "entities": "두 인물의 외모와 복장은 참조와 일치하나, 제거되어야 할 수갑이 현우의 손목에 채워져 있음.",
        "hard_violations": [
         "수갑을 제거하라는 지시를 어기고 현우의 손목에 수갑이 묘사됨."
        ],
        "physics": "인물들은 의자에 앉아 체중을 지탱하고 있으며 발은 바닥을 딛고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "수갑이 제거된 상태와 두 인물의 측면 풀샷 구도를 비롯한 프롬프트의 지시사항과 참조 이미지를 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "명시적인 지시에도 불구하고 현우의 손목에 수갑을 그려 넣어 핵심 요구사항을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 테이블을 사이에 두고 서로의 얼굴을 똑바로 마주 봄.",
        "built_space": "창문이 있는 넓은 면회실 내부로, 다수의 철제 테이블과 의자가 배치되어 있으며 중앙 테이블에 인물들이 위치함.",
        "entities": "얼굴에 상처가 있는 현우와 정장 차림의 윤성찬이 참조와 일치하며, 테이블 위 사진이 존재하고 수갑은 없음.",
        "hard_violations": [],
        "physics": "두 인물 모두 의자에 안정적으로 착석하여 발이 바닥에 닿아 있으며, 손은 테이블 위에 자연스럽게 놓임."
       },
       {
        "label": "B",
        "direction": "두 인물이 중앙 테이블을 사이에 두고 서로를 주시함.",
        "built_space": "면회실 내부의 중앙 테이블에 두 인물이 마주 앉아 있으며 주변으로 여백의 공간이 보임.",
        "entities": "두 인물의 외모와 복장은 참조와 일치하나, 제거되어야 할 수갑이 현우의 손목에 채워져 있음.",
        "hard_violations": [
         "수갑을 제거하라는 지시를 어기고 현우의 손목에 수갑이 묘사됨."
        ],
        "physics": "인물들은 의자에 앉아 체중을 지탱하고 있으며 발은 바닥을 딛고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "마주 보는 옆모습 풀샷과 인물 외형은 맞지만, 테이블 가로대가 하체를 더 가리고 두 사람의 무게 배분 차이가 약하며 현우의 손목에 금속 구속구가 남아 있다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 전경 자세와 윤성찬의 곧게 앉은 자세를 대비시킨 측면 풀샷이 지시에 더 충실하지만, 현우 손목의 금속 구속구는 수갑 제거 조건과 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 윤성찬의 얼굴을, 오른쪽 윤성찬은 왼쪽 현우의 얼굴을 바라본다. 두 사람의 몸과 의자도 테이블을 사이에 두고 서로 향한다. 카메라를 바라보는 인물은 없다.",
        "built_space": "중앙 금속 테이블 하나와 이를 마주 보는 접이식 의자 두 개가 있고, 뒤에는 빈 테이블 하나와 의자 두 개가 보인다. 높은 창, 밝은 상부와 회색 하부로 나뉜 벽은 이전 장면의 기본 공간 특징을 따른다. 왼쪽에는 격자형 설비와 노출 배관, 오른쪽에는 벽 설비와 세면대 일부가 보이는데 이전의 좁은 화면으로는 이들의 연속성을 확인할 수 없다. 두 사람의 전신과 신발은 프레임 안에 있지만 테이블 다리와 가로대가 하체 일부를 가린다. 상판은 얕게 보인다.",
        "entities": "인물은 두 명뿐이다. 현우는 앳된 동아시아계 남성으로 검은 헝클어진 머리, 멍든 얼굴, 회색 셔츠, 올리브색 카고 바지와 낡은 신발을 갖춰 참조와 대체로 맞는다. 윤성찬은 회색 머리와 안경을 쓴 고령의 동아시아계 남성이며 어두운 코트와 정장을 입어 참조 의상에 가깝다. 현우 손목에는 수갑처럼 보이는 금속 고리와 연결부가 남아 있어 제거 조건에 어긋난다. 테이블 위 종이는 있으나 찰리 사진의 내용은 확인하기 어렵다. 다리 부상은 바지에 가려 확인할 수 없다. 오른쪽 위 표지에는 작은 문자 같은 흔적이 있으나 명확히 판독되지는 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 엉덩이가 각자의 의자 좌판에 놓이고 등받이는 등 뒤에 있다. 테이블에 올린 팔과 손은 상판의 지지를 받는다. 가까운 신발은 바닥에 닿고 반대쪽 발도 다리와 좌석으로 지지되는 자연스러운 앉은 자세다. 의자와 테이블은 다리로 바닥에 서 있으며, 종이는 상판에 놓여 있다. 지지 없는 부유나 불가능한 신체 구조는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우의 얼굴과 시선은 오른쪽 윤성찬을 향하고, 윤성찬 역시 현우의 얼굴을 바라본다. 두 인물은 거의 정측면으로 대치하며 테이블이 두 시선 사이의 협상 축을 만든다.",
        "built_space": "중앙 금속 테이블 하나와 점유된 의자 두 개가 명확하다. 뒤에는 중앙과 좌우에 빈 테이블 세 개가 보이고, 화면 양쪽 가장자리에는 추가 테이블 일부가 걸친다. 주변 빈 의자는 여러 개이며 일부는 프레임 밖으로 잘려 정확한 총수는 확정하기 어렵다. 높은 창과 밝은 상부·회색 하부 벽, 낮의 실내광은 이전 장면과 대체로 이어진다. 뒤쪽 좌우에는 격자형 설비가 하나씩 있으나 참조는 그 수를 확정하지 않는다. 두 몸과 신발이 모두 프레임에 들어오며, 테이블 구조에 의한 하체 가림은 일부 남는다. 현우의 앞으로 실린 체중과 윤성찬의 뒤로 물러난 착석 위치가 뚜렷하다.",
        "entities": "두 명 모두 지정된 인물의 연령대와 남성 외형에 부합한다. 현우는 검은 머리, 젊은 얼굴, 볼의 상처, 회색 셔츠와 카고 바지, 낡은 신발을 유지한다. 윤성찬은 회색 머리와 안경, 주름진 얼굴, 줄무늬 정장을 갖췄지만 참조의 긴 겉코트는 보이지 않는다. 현우의 손목 아래에는 금속 고리와 짧은 사슬처럼 보이는 부분이 있어 수갑이 완전히 제거된 상태로 읽히지 않는다. 중앙에는 사진이 인쇄된 종이 한 장이 있으나 찰리와의 동일성은 이 크기로 확인할 수 없다. 부상당한 다리는 옷에 가려 상태를 판별할 수 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 골반은 좌판에 놓이고 앞으로 기울인 상체는 좌석과 테이블에 댄 팔로 지지된다. 윤성찬은 좌판에 안정적으로 앉아 손을 허벅지 위에 둔다. 두 사람의 신발은 바닥에 닿는 것으로 보이며, 테이블과 의자도 바닥에 정상적으로 지지된다. 사진은 상판에 놓여 있다. 부유하는 인물이나 물체, 불가능한 착석 구조는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "마주 보는 옆모습 풀샷과 인물 외형은 맞지만, 테이블 가로대가 하체를 더 가리고 두 사람의 무게 배분 차이가 약하며 현우의 손목에 금속 구속구가 남아 있다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "현우의 전경 자세와 윤성찬의 곧게 앉은 자세를 대비시킨 측면 풀샷이 지시에 더 충실하지만, 현우 손목의 금속 구속구는 수갑 제거 조건과 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 윤성찬의 얼굴을, 오른쪽 윤성찬은 왼쪽 현우의 얼굴을 바라본다. 두 사람의 몸과 의자도 테이블을 사이에 두고 서로 향한다. 카메라를 바라보는 인물은 없다.",
        "built_space": "중앙 금속 테이블 하나와 이를 마주 보는 접이식 의자 두 개가 있고, 뒤에는 빈 테이블 하나와 의자 두 개가 보인다. 높은 창, 밝은 상부와 회색 하부로 나뉜 벽은 이전 장면의 기본 공간 특징을 따른다. 왼쪽에는 격자형 설비와 노출 배관, 오른쪽에는 벽 설비와 세면대 일부가 보이는데 이전의 좁은 화면으로는 이들의 연속성을 확인할 수 없다. 두 사람의 전신과 신발은 프레임 안에 있지만 테이블 다리와 가로대가 하체 일부를 가린다. 상판은 얕게 보인다.",
        "entities": "인물은 두 명뿐이다. 현우는 앳된 동아시아계 남성으로 검은 헝클어진 머리, 멍든 얼굴, 회색 셔츠, 올리브색 카고 바지와 낡은 신발을 갖춰 참조와 대체로 맞는다. 윤성찬은 회색 머리와 안경을 쓴 고령의 동아시아계 남성이며 어두운 코트와 정장을 입어 참조 의상에 가깝다. 현우 손목에는 수갑처럼 보이는 금속 고리와 연결부가 남아 있어 제거 조건에 어긋난다. 테이블 위 종이는 있으나 찰리 사진의 내용은 확인하기 어렵다. 다리 부상은 바지에 가려 확인할 수 없다. 오른쪽 위 표지에는 작은 문자 같은 흔적이 있으나 명확히 판독되지는 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 엉덩이가 각자의 의자 좌판에 놓이고 등받이는 등 뒤에 있다. 테이블에 올린 팔과 손은 상판의 지지를 받는다. 가까운 신발은 바닥에 닿고 반대쪽 발도 다리와 좌석으로 지지되는 자연스러운 앉은 자세다. 의자와 테이블은 다리로 바닥에 서 있으며, 종이는 상판에 놓여 있다. 지지 없는 부유나 불가능한 신체 구조는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우의 얼굴과 시선은 오른쪽 윤성찬을 향하고, 윤성찬 역시 현우의 얼굴을 바라본다. 두 인물은 거의 정측면으로 대치하며 테이블이 두 시선 사이의 협상 축을 만든다.",
        "built_space": "중앙 금속 테이블 하나와 점유된 의자 두 개가 명확하다. 뒤에는 중앙과 좌우에 빈 테이블 세 개가 보이고, 화면 양쪽 가장자리에는 추가 테이블 일부가 걸친다. 주변 빈 의자는 여러 개이며 일부는 프레임 밖으로 잘려 정확한 총수는 확정하기 어렵다. 높은 창과 밝은 상부·회색 하부 벽, 낮의 실내광은 이전 장면과 대체로 이어진다. 뒤쪽 좌우에는 격자형 설비가 하나씩 있으나 참조는 그 수를 확정하지 않는다. 두 몸과 신발이 모두 프레임에 들어오며, 테이블 구조에 의한 하체 가림은 일부 남는다. 현우의 앞으로 실린 체중과 윤성찬의 뒤로 물러난 착석 위치가 뚜렷하다.",
        "entities": "두 명 모두 지정된 인물의 연령대와 남성 외형에 부합한다. 현우는 검은 머리, 젊은 얼굴, 볼의 상처, 회색 셔츠와 카고 바지, 낡은 신발을 유지한다. 윤성찬은 회색 머리와 안경, 주름진 얼굴, 줄무늬 정장을 갖췄지만 참조의 긴 겉코트는 보이지 않는다. 현우의 손목 아래에는 금속 고리와 짧은 사슬처럼 보이는 부분이 있어 수갑이 완전히 제거된 상태로 읽히지 않는다. 중앙에는 사진이 인쇄된 종이 한 장이 있으나 찰리와의 동일성은 이 크기로 확인할 수 없다. 부상당한 다리는 옷에 가려 상태를 판별할 수 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 골반은 좌판에 놓이고 앞으로 기울인 상체는 좌석과 테이블에 댄 팔로 지지된다. 윤성찬은 좌판에 안정적으로 앉아 손을 허벅지 위에 둔다. 두 사람의 신발은 바닥에 닿는 것으로 보이며, 테이블과 의자도 바닥에 정상적으로 지지된다. 사진은 상판에 놓여 있다. 부유하는 인물이나 물체, 불가능한 착석 구조는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.232
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.982
   },
   "violations": {
    "B": [
     "[gemini-pro] 수갑을 제거하라는 지시를 어기고 현우의 손목에 수갑이 묘사됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 982
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "수갑이 제거된 상태와 두 인물의 측면 풀샷 구도를 비롯한 프롬프트의 지시사항과 참조 이미지를 정확히 구현함."
   },
   {
    "label": "B",
    "score": 982,
    "verdict_ko": "명시적인 지시에도 불구하고 현우의 손목에 수갑을 그려 넣어 핵심 요구사항을 위반함.  ★위반: [gemini-pro] 수갑을 제거하라는 지시를 어기고 현우의 손목에 수갑이 묘사됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh6_sel.png",
    "asset_id": "91df6d50-6f4e-4dee-9ea9-e89a85ef51be",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-0438-7af3-83ee-15077ce6b4db",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S23sh6"
  }
 },
 "S23sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:56:48.860536+00:00",
  "fingerprint": "1791d99caae6926ef8c5cea00b3b3d56986a738410952c2547251854cf67af69",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S23sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S23sh11_sel.png",
  "source_sha256": "4f0b5dbad634c3f21f6e7ff84d78209ec5d2dd589d0420c9123f109dc81250d8",
  "file": "S23sh11_cine.png",
  "staged_sha256": "8a8b97be1946c92ef41fe932544d003c35bfe23a1722924eebfa40688ab1c680",
  "latency_ms": 9919
 },
 "S24sh4::signage": {
  "fp": "3089b4e860d5cdd8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S24sh4": {
  "input_fingerprint": "466595c6a40a2201",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 양손으로 꽉 붙잡은 채 눈물을 글썽이며 환하게 웃는 페드로의 얼굴.\n\nLOCATION (lock): Inside the front living area of the refugee family's container home, with daytime light entering from the entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (The reunion takes place inside the home); used as A soft, narrow background margin establishes the interior without inventing furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use unobtrusive ambient illumination appropriate to the interior, preserving the tears and smile without introducing an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the container's fixed interior surfaces, household fixtures, and daylight as the location reference. Exclude the girl and robot, and do not restore an undisturbed arrangement of belongings after the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 retains the damage and disorder from the militia's search. 현우: He is back inside his container, no longer handcuffed. His facial bruises and injured leg persist. 페드로: He has hurried inside the container and is out of breath.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 양손으로 꽉 붙잡은 채 눈물을 글썽이며 환하게 웃는 페드로의 얼굴.\n\nLOCATION (lock): Inside the front living area of the refugee family's container home, with daytime light entering from the entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (The reunion takes place inside the home); used as A soft, narrow background margin establishes the interior without inventing furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use unobtrusive ambient illumination appropriate to the interior, preserving the tears and smile without introducing an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the container's fixed interior surfaces, household fixtures, and daylight as the location reference. Exclude the girl and robot, and do not restore an undisturbed arrangement of belongings after the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 retains the damage and disorder from the militia's search. 현우: He is back inside his container, no longer handcuffed. His facial bruises and injured leg persist. 페드로: He has hurried inside the container and is out of breath.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 양손으로 꽉 붙잡은 채 눈물을 글썽이며 환하게 웃는 페드로의 얼굴.\n\nLOCATION (lock): Inside the front living area of the refugee family's container home, with daytime light entering from the entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (The reunion takes place inside the home); used as A soft, narrow background margin establishes the interior without inventing furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use unobtrusive ambient illumination appropriate to the interior, preserving the tears and smile without introducing an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the container's fixed interior surfaces, household fixtures, and daylight as the location reference. Exclude the girl and robot, and do not restore an undisturbed arrangement of belongings after the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 retains the damage and disorder from the militia's search. 현우: He is back inside his container, no longer handcuffed. His facial bruises and injured leg persist. 페드로: He has hurried inside the container and is out of breath.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "페드로는 화면 우측에 등을 보인 현우의 얼굴을 향해 시선을 고정하고 있다.",
    "built_space": "컨테이너 내부 통로로, 이전 샷과 유사하게 바닥에 흩어진 물건들이 있는 배경이 얕은 심도로 보인다.",
    "entities": "페드로는 눈물을 흘리며 환하게 웃고 있으나 비니를 쓰지 않았다. 현우는 회색 셔츠를 입고 얼굴에 상처가 있다.",
    "hard_violations": [
     "[gemini-pro] 페드로의 오른쪽 어깨에서 나온 팔 끝에 엄지손가락이 바깥쪽을 향하는 왼손이 달려 있는 해부학적 불가능 구조."
    ],
    "physics": "페드로의 팔이 현우의 어깨에 얹혀 있으나, 신체 구조상 불가능한 손의 결합으로 인해 물리적 설득력이 떨어진다."
   },
   {
    "label": "B",
    "direction": "페드로와 현우가 서로 마주 서서 시선을 교환하고 있다.",
    "built_space": "컨테이너 내부의 골판 벽면과 창문이 보이며, 바닥에 물건들이 어질러져 있어 이전 샷의 배경 설정과 일치한다.",
    "entities": "페드로는 지시된 비니와 후드티를 입고 눈물을 글썽이며 웃고 있다. 현우는 얼굴에 타박상을 입은 채 서 있다.",
    "hard_violations": [],
    "physics": "두 사람이 지면에 안정적으로 서 있으며, 페드로가 양손으로 현우의 어깨를 단단히 붙잡아 지탱하고 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 클로즈업보다 넓은 미디엄 숏으로 연출되었으나, 치명적인 오류 없이 인물의 복장(비니 포함)과 감정 표현, 공간적 배경을 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "클로즈업 프레이밍은 따랐으나 페드로의 오른팔에 왼손이 달려 있는 치명적인 해부학적 오류가 있으며, 지정된 비니 복장도 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "페드로는 화면 우측에 등을 보인 현우의 얼굴을 향해 시선을 고정하고 있다.",
        "built_space": "컨테이너 내부 통로로, 이전 샷과 유사하게 바닥에 흩어진 물건들이 있는 배경이 얕은 심도로 보인다.",
        "entities": "페드로는 눈물을 흘리며 환하게 웃고 있으나 비니를 쓰지 않았다. 현우는 회색 셔츠를 입고 얼굴에 상처가 있다.",
        "hard_violations": [
         "페드로의 오른쪽 어깨에서 나온 팔 끝에 엄지손가락이 바깥쪽을 향하는 왼손이 달려 있는 해부학적 불가능 구조."
        ],
        "physics": "페드로의 팔이 현우의 어깨에 얹혀 있으나, 신체 구조상 불가능한 손의 결합으로 인해 물리적 설득력이 떨어진다."
       },
       {
        "label": "B",
        "direction": "페드로와 현우가 서로 마주 서서 시선을 교환하고 있다.",
        "built_space": "컨테이너 내부의 골판 벽면과 창문이 보이며, 바닥에 물건들이 어질러져 있어 이전 샷의 배경 설정과 일치한다.",
        "entities": "페드로는 지시된 비니와 후드티를 입고 눈물을 글썽이며 웃고 있다. 현우는 얼굴에 타박상을 입은 채 서 있다.",
        "hard_violations": [],
        "physics": "두 사람이 지면에 안정적으로 서 있으며, 페드로가 양손으로 현우의 어깨를 단단히 붙잡아 지탱하고 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 클로즈업보다 넓은 미디엄 숏으로 연출되었으나, 치명적인 오류 없이 인물의 복장(비니 포함)과 감정 표현, 공간적 배경을 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "클로즈업 프레이밍은 따랐으나 페드로의 오른팔에 왼손이 달려 있는 치명적인 해부학적 오류가 있으며, 지정된 비니 복장도 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "페드로는 화면 우측에 등을 보인 현우의 얼굴을 향해 시선을 고정하고 있다.",
        "built_space": "컨테이너 내부 통로로, 이전 샷과 유사하게 바닥에 흩어진 물건들이 있는 배경이 얕은 심도로 보인다.",
        "entities": "페드로는 눈물을 흘리며 환하게 웃고 있으나 비니를 쓰지 않았다. 현우는 회색 셔츠를 입고 얼굴에 상처가 있다.",
        "hard_violations": [
         "페드로의 오른쪽 어깨에서 나온 팔 끝에 엄지손가락이 바깥쪽을 향하는 왼손이 달려 있는 해부학적 불가능 구조."
        ],
        "physics": "페드로의 팔이 현우의 어깨에 얹혀 있으나, 신체 구조상 불가능한 손의 결합으로 인해 물리적 설득력이 떨어진다."
       },
       {
        "label": "B",
        "direction": "페드로와 현우가 서로 마주 서서 시선을 교환하고 있다.",
        "built_space": "컨테이너 내부의 골판 벽면과 창문이 보이며, 바닥에 물건들이 어질러져 있어 이전 샷의 배경 설정과 일치한다.",
        "entities": "페드로는 지시된 비니와 후드티를 입고 눈물을 글썽이며 웃고 있다. 현우는 얼굴에 타박상을 입은 채 서 있다.",
        "hard_violations": [],
        "physics": "두 사람이 지면에 안정적으로 서 있으며, 페드로가 양손으로 현우의 어깨를 단단히 붙잡아 지탱하고 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "양손으로 어깨를 잡는 동작과 눈물 어린 웃음, 비니는 맞지만 두 사람의 상반신과 실내를 넓게 보여 페드로 얼굴 클로즈업이라는 핵심 구도를 놓쳤다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우의 어깨 너머로 페드로의 눈물 어린 웃음을 가까이 담아 지정된 순간과 구도에 더 충실하지만, 참조의 비니가 빠지고 얼굴 인상도 다소 달라졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 페드로는 오른쪽 현우의 눈을 바라보고, 현우도 페드로를 마주 본다. 페드로의 두 팔은 현우 쪽으로 뻗어 각각 양쪽 어깨에 닿는다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "흰 금속 골벽과 천장, 왼쪽 창 하나, 오른쪽 출입구 가장자리 하나가 보인다. 뒤쪽에는 탁자와 상자 및 흩어진 물건들이 있고, 오른쪽 벽에는 용기 여러 개를 올린 선반 하나가 보인다. 컨테이너 재질과 낮빛은 이어지지만, 참조에서 확인되지 않는 선반까지 드러나며 배경이 요구된 좁고 흐린 여백보다 훨씬 넓다. 두 사람은 실내에서 마주 선 상반신 구도로 보인다.",
        "entities": "젊은 남성 두 명만 있으며 소녀와 로봇, 읽을 수 있는 글자는 없다. 페드로는 라틴계 혼혈 설정에 어긋나지 않는 외모이며 검은 비니와 짙은 후드가 참조에 부합한다. 현우는 동아시아계의 앳된 얼굴, 헝클어진 검은 머리, 회색 셔츠와 얼굴 상처를 유지한다. 페드로의 눈물 자국과 이를 드러낸 밝은 웃음이 뚜렷하다. 다리 부상과 수갑 유무는 이 구도로 판단할 수 없다.",
        "hard_violations": [],
        "physics": "페드로의 가까운 손은 현우의 어깨 위를 감싸고, 반대 손도 목 아래 반대편 어깨에 접촉한다. 손목과 팔의 연결은 자연스럽고 셔츠에 손이 얹히는 관계도 성립한다. 두 몸통은 화면 아래로 이어지며 발은 잘렸지만 공중에 떠 있다는 징후는 없다. 눈물은 볼 표면을 따라 아래로 흐른다."
       },
       {
        "label": "B",
        "direction": "페드로의 시선은 바로 앞 현우의 얼굴을 향한다. 현우는 카메라에 등과 옆얼굴을 보이며 페드로 쪽으로 고개를 돌리고 있다. 전경의 손과 오른쪽 가장자리의 손이 현우의 양쪽 어깨를 각각 잡는다.",
        "built_space": "현우의 머리와 어깨가 오른쪽 전경을 차지하고 페드로 얼굴이 중심 피사체가 된다. 왼쪽에는 참조와 유사한 흰 테두리의 벽면 개구부 하나와 걸린 어두운 옷, 세로 금속 벽면이 보인다. 뒤쪽 탁자 일부와 바닥의 흩어진 의류·물건도 흐리게 남아 수색 뒤의 어수선함을 유지한다. 배경 여백이 완전히 좁지는 않지만 A보다 얼굴 중심의 가까운 구도이며, 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 사람은 페드로와 현우 두 명뿐이다. 페드로는 앳된 젊은 남성으로 짙은 후드를 입었지만 참조의 검은 비니가 없고 머리와 얼굴 인상도 다소 다르다. 현우의 검은 헝클어진 머리, 회색 셔츠와 옆얼굴 상처는 이어진다. 페드로의 젖은 눈과 볼의 눈물, 밝은 미소가 선명하다. 숨이 찬 상태는 확실하지 않으며, 하체와 수갑 상태는 프레임 밖이다. 읽을 수 있는 글자나 금지된 인물은 없다.",
        "hard_violations": [],
        "physics": "가까운 손은 후드 소매에서 이어져 현우의 어깨를 감싸고, 반대 손은 오른쪽 화면 가장자리에서 다른 어깨에 닿는다. 가려진 팔의 경로도 이 마주 선 자세에서 가능하다. 손가락과 어깨의 접촉에 명백한 해부학적 모순은 없다. 하체는 화면 밖이지만 몸이 떠 있거나 지지 없이 매달린 모습은 아니며, 눈물은 실제 피부 표면을 따라 흐른다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "양손으로 어깨를 잡는 동작과 눈물 어린 웃음, 비니는 맞지만 두 사람의 상반신과 실내를 넓게 보여 페드로 얼굴 클로즈업이라는 핵심 구도를 놓쳤다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우의 어깨 너머로 페드로의 눈물 어린 웃음을 가까이 담아 지정된 순간과 구도에 더 충실하지만, 참조의 비니가 빠지고 얼굴 인상도 다소 달라졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 페드로는 오른쪽 현우의 눈을 바라보고, 현우도 페드로를 마주 본다. 페드로의 두 팔은 현우 쪽으로 뻗어 각각 양쪽 어깨에 닿는다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "흰 금속 골벽과 천장, 왼쪽 창 하나, 오른쪽 출입구 가장자리 하나가 보인다. 뒤쪽에는 탁자와 상자 및 흩어진 물건들이 있고, 오른쪽 벽에는 용기 여러 개를 올린 선반 하나가 보인다. 컨테이너 재질과 낮빛은 이어지지만, 참조에서 확인되지 않는 선반까지 드러나며 배경이 요구된 좁고 흐린 여백보다 훨씬 넓다. 두 사람은 실내에서 마주 선 상반신 구도로 보인다.",
        "entities": "젊은 남성 두 명만 있으며 소녀와 로봇, 읽을 수 있는 글자는 없다. 페드로는 라틴계 혼혈 설정에 어긋나지 않는 외모이며 검은 비니와 짙은 후드가 참조에 부합한다. 현우는 동아시아계의 앳된 얼굴, 헝클어진 검은 머리, 회색 셔츠와 얼굴 상처를 유지한다. 페드로의 눈물 자국과 이를 드러낸 밝은 웃음이 뚜렷하다. 다리 부상과 수갑 유무는 이 구도로 판단할 수 없다.",
        "hard_violations": [],
        "physics": "페드로의 가까운 손은 현우의 어깨 위를 감싸고, 반대 손도 목 아래 반대편 어깨에 접촉한다. 손목과 팔의 연결은 자연스럽고 셔츠에 손이 얹히는 관계도 성립한다. 두 몸통은 화면 아래로 이어지며 발은 잘렸지만 공중에 떠 있다는 징후는 없다. 눈물은 볼 표면을 따라 아래로 흐른다."
       },
       {
        "label": "A",
        "direction": "페드로의 시선은 바로 앞 현우의 얼굴을 향한다. 현우는 카메라에 등과 옆얼굴을 보이며 페드로 쪽으로 고개를 돌리고 있다. 전경의 손과 오른쪽 가장자리의 손이 현우의 양쪽 어깨를 각각 잡는다.",
        "built_space": "현우의 머리와 어깨가 오른쪽 전경을 차지하고 페드로 얼굴이 중심 피사체가 된다. 왼쪽에는 참조와 유사한 흰 테두리의 벽면 개구부 하나와 걸린 어두운 옷, 세로 금속 벽면이 보인다. 뒤쪽 탁자 일부와 바닥의 흩어진 의류·물건도 흐리게 남아 수색 뒤의 어수선함을 유지한다. 배경 여백이 완전히 좁지는 않지만 A보다 얼굴 중심의 가까운 구도이며, 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 사람은 페드로와 현우 두 명뿐이다. 페드로는 앳된 젊은 남성으로 짙은 후드를 입었지만 참조의 검은 비니가 없고 머리와 얼굴 인상도 다소 다르다. 현우의 검은 헝클어진 머리, 회색 셔츠와 옆얼굴 상처는 이어진다. 페드로의 젖은 눈과 볼의 눈물, 밝은 미소가 선명하다. 숨이 찬 상태는 확실하지 않으며, 하체와 수갑 상태는 프레임 밖이다. 읽을 수 있는 글자나 금지된 인물은 없다.",
        "hard_violations": [],
        "physics": "가까운 손은 후드 소매에서 이어져 현우의 어깨를 감싸고, 반대 손은 오른쪽 화면 가장자리에서 다른 어깨에 닿는다. 가려진 팔의 경로도 이 마주 선 자세에서 가능하다. 손가락과 어깨의 접촉에 명백한 해부학적 모순은 없다. 하체는 화면 밖이지만 몸이 떠 있거나 지지 없이 매달린 모습은 아니며, 눈물은 실제 피부 표면을 따라 흐른다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.333,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.083,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 페드로의 오른쪽 어깨에서 나온 팔 끝에 엄지손가락이 바깥쪽을 향하는 왼손이 달려 있는 해부학적 불가능 구조."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1750,
   "A": 1083
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "요구된 클로즈업보다 넓은 미디엄 숏으로 연출되었으나, 치명적인 오류 없이 인물의 복장(비니 포함)과 감정 표현, 공간적 배경을 충실히 구현했습니다."
   },
   {
    "label": "A",
    "score": 1083,
    "verdict_ko": "클로즈업 프레이밍은 따랐으나 페드로의 오른팔에 왼손이 달려 있는 치명적인 해부학적 오류가 있으며, 지정된 비니 복장도 누락되었습니다.  ★위반: [gemini-pro] 페드로의 오른쪽 어깨에서 나온 팔 끝에 엄지손가락이 바깥쪽을 향하는 왼손이 달려 있는 해부학적 불가능 구조."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S14sh9_sel.png",
    "asset_id": "f5838b37-db98-49bc-b879-9704dc9a14c6",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-05e7-71cd-8cce-06ff88456cc3",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S14sh9"
  }
 },
 "S24sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:58:17.849761+00:00",
  "fingerprint": "159e53bbe75397516a878fb1bc2a5499b20056e315b343a1a5ae0eedbed5755b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S24sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S24sh4_sel.png",
  "source_sha256": "b2e29c9903bd109d1c9d761f2e1e55fb39828e6a9b0d9dfee4501beb9db78e5b",
  "file": "S24sh4_cine.png",
  "staged_sha256": "a9bf03b2f57eacefde2cd9f1e7c4e69324d88a4f4268981b7739c84b5c04bc70",
  "latency_ms": 10908
 },
 "S24sh7::signage": {
  "fp": "a8103b87dcc7eeb2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S24sh7": {
  "input_fingerprint": "39a03ced96279aaf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 페드로의 말에 충격을 받은 듯 눈이 커진 채 얼어붙은 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): In the container home's shared interior near the entrance, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (Visible only as a narrow background around 현우); used as Maintains the location while leaving the reaction unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the reunion's restrained ambient lighting and controlled contrast, keeping the widened eyes clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container's fixed surfaces, household fixtures, and daytime light. Exclude the girl and robot from the reference, and do not undo the disruption caused by the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container remains damaged and disordered; its number is 7-31. 현우: He remains inside the container with free wrists, facial bruising and an unhealed leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 페드로의 말에 충격을 받은 듯 눈이 커진 채 얼어붙은 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): In the container home's shared interior near the entrance, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (Visible only as a narrow background around 현우); used as Maintains the location while leaving the reaction unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the reunion's restrained ambient lighting and controlled contrast, keeping the widened eyes clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container's fixed surfaces, household fixtures, and daytime light. Exclude the girl and robot from the reference, and do not undo the disruption caused by the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container remains damaged and disordered; its number is 7-31. 현우: He remains inside the container with free wrists, facial bruising and an unhealed leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 페드로의 말에 충격을 받은 듯 눈이 커진 채 얼어붙은 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): In the container home's shared interior near the entrance, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (Visible only as a narrow background around 현우); used as Maintains the location while leaving the reaction unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the reunion's restrained ambient lighting and controlled contrast, keeping the widened eyes clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container's fixed surfaces, household fixtures, and daytime light. Exclude the girl and robot from the reference, and do not undo the disruption caused by the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container remains damaged and disordered; its number is 7-31. 현우: He remains inside the container with free wrists, facial bruising and an unhealed leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 정면 카메라 렌즈를 똑바로 향하고 있습니다.",
    "built_space": "컨테이너 내부 통로를 바라보는 구도입니다. 왼쪽 벽에 창문들이 있고, 오른쪽 벽에 물건이 놓인 선반과 배경의 가전제품이 있어 이전 샷의 공간 설정과 정확히 일치합니다.",
    "entities": "기준 이미지에 제시된 현우의 인상, 헝클어진 검은 머리, 회색 셔츠, 얼굴의 멍과 상처 등이 모두 일치합니다. 또한 지시문대로 눈을 크게 뜬 굳은 얼굴을 보여줍니다.",
    "hard_violations": [],
    "physics": "보이지 않는 바닥에 안정적으로 체중을 싣고 서 있는 자연스러운 자세를 유지하고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 화면 우측 바깥쪽을 향하고 있습니다.",
    "built_space": "컨테이너 내부 통로를 배경으로 하고 있으며, 왼쪽의 창문과 오른쪽의 선반 등 이전 샷의 공간적 특징을 잘 유지하고 있습니다.",
    "entities": "현우의 외형, 의상, 상처 등은 기준 이미지와 잘 일치하지만, 프롬프트가 요구한 '눈이 커진 채 얼어붙은' 표정 대신 눈물을 머금은 슬픈 표정을 짓고 있습니다.",
    "hard_violations": [],
    "physics": "바닥에 단단히 발을 딛고 서 있는 자세로, 물리적인 어색함 없이 몸을 잘 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트에서 요구한 '눈이 커진 채 얼어붙은 굳은 얼굴'의 표정과 감정 상태를 클로즈업 샷으로 매우 정확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "충격을 받아 눈이 커진 표정이 아니라 눈물이 맺힌 슬픈 표정을 렌더링하여 프롬프트의 핵심적인 연기 지시를 놓쳤습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 정면 카메라 렌즈를 똑바로 향하고 있습니다.",
        "built_space": "컨테이너 내부 통로를 바라보는 구도입니다. 왼쪽 벽에 창문들이 있고, 오른쪽 벽에 물건이 놓인 선반과 배경의 가전제품이 있어 이전 샷의 공간 설정과 정확히 일치합니다.",
        "entities": "기준 이미지에 제시된 현우의 인상, 헝클어진 검은 머리, 회색 셔츠, 얼굴의 멍과 상처 등이 모두 일치합니다. 또한 지시문대로 눈을 크게 뜬 굳은 얼굴을 보여줍니다.",
        "hard_violations": [],
        "physics": "보이지 않는 바닥에 안정적으로 체중을 싣고 서 있는 자연스러운 자세를 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 우측 바깥쪽을 향하고 있습니다.",
        "built_space": "컨테이너 내부 통로를 배경으로 하고 있으며, 왼쪽의 창문과 오른쪽의 선반 등 이전 샷의 공간적 특징을 잘 유지하고 있습니다.",
        "entities": "현우의 외형, 의상, 상처 등은 기준 이미지와 잘 일치하지만, 프롬프트가 요구한 '눈이 커진 채 얼어붙은' 표정 대신 눈물을 머금은 슬픈 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "바닥에 단단히 발을 딛고 서 있는 자세로, 물리적인 어색함 없이 몸을 잘 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트에서 요구한 '눈이 커진 채 얼어붙은 굳은 얼굴'의 표정과 감정 상태를 클로즈업 샷으로 매우 정확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "충격을 받아 눈이 커진 표정이 아니라 눈물이 맺힌 슬픈 표정을 렌더링하여 프롬프트의 핵심적인 연기 지시를 놓쳤습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 정면 카메라 렌즈를 똑바로 향하고 있습니다.",
        "built_space": "컨테이너 내부 통로를 바라보는 구도입니다. 왼쪽 벽에 창문들이 있고, 오른쪽 벽에 물건이 놓인 선반과 배경의 가전제품이 있어 이전 샷의 공간 설정과 정확히 일치합니다.",
        "entities": "기준 이미지에 제시된 현우의 인상, 헝클어진 검은 머리, 회색 셔츠, 얼굴의 멍과 상처 등이 모두 일치합니다. 또한 지시문대로 눈을 크게 뜬 굳은 얼굴을 보여줍니다.",
        "hard_violations": [],
        "physics": "보이지 않는 바닥에 안정적으로 체중을 싣고 서 있는 자연스러운 자세를 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 우측 바깥쪽을 향하고 있습니다.",
        "built_space": "컨테이너 내부 통로를 배경으로 하고 있으며, 왼쪽의 창문과 오른쪽의 선반 등 이전 샷의 공간적 특징을 잘 유지하고 있습니다.",
        "entities": "현우의 외형, 의상, 상처 등은 기준 이미지와 잘 일치하지만, 프롬프트가 요구한 '눈이 커진 채 얼어붙은' 표정 대신 눈물을 머금은 슬픈 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "바닥에 단단히 발을 딛고 서 있는 자세로, 물리적인 어색함 없이 몸을 잘 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 클로즈업 안에서 화면 밖 상대를 향한 시선과 커진 눈, 멈춘 표정이 페드로의 말에 충격받은 순간을 자연스럽게 구현한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "굳은 얼굴과 커진 눈은 명확하지만, 렌즈를 거의 정면으로 응시해 페드로에게 반응하는 장면보다 정면 초상처럼 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 화면 오른쪽의 카메라 밖 가까운 지점을 향한다. 상대는 보이지 않지만 페드로를 바라보며 말을 듣는 시선으로 성립한다. 무기나 방향을 확인해야 할 휴대 물체는 없다.",
        "built_space": "현우의 머리와 어깨 뒤로 낡은 밝은색 금속 벽과 골이 있는 천장이 보인다. 좌우 벽에 창이 두 구간씩 드러나고, 오른쪽에는 용기들이 놓인 벽 선반 하나와 그 아래 회색 수납 가구 하나가 보인다. 왼쪽에는 게시판 하나와 걸린 짙은 옷 한 벌이 있으며 뒤쪽에는 어질러진 생활 물건이 남아 있다. 입구 부근에 선 인물을 실내 방향으로 촬영한 배치가 성립한다. 참조에서 가려졌던 벽 구간의 창 위치까지 정확히 일치하는지는 확인하기 어렵지만, 명백한 중복 설비나 불가능한 반사는 없다.",
        "entities": "인물은 현우로 보이는 젊은 동아시아계 남성 한 명뿐이다. 앳된 얼굴, 헝클어진 검은 머리, 회색 셔츠가 참조와 대체로 일치하며 얼굴의 멍과 긁힌 상처도 유지된다. 정확한 나이와 한국계 미국인이라는 국적 배경은 외형만으로 확인할 수 없다. 눈은 정상적인 홍채와 동공을 유지한 채 커져 있고 입술이 약간 벌어져 있다. 다른 사람이나 로봇은 없고 읽을 수 있는 글자도 없다. 손목과 다리, 컨테이너 번호는 클로즈업 밖이므로 확인 대상에서 제외한다.",
        "hard_violations": [],
        "physics": "보이는 머리는 목에, 목은 셔츠를 입은 어깨와 몸통에 자연스럽게 연결된다. 발과 지면은 프레임 밖이지만 부유를 암시하는 자세는 없다. 선반 위 용기와 가구 위 물건은 받침면에 놓여 있고, 왼쪽 옷은 벽에 걸린 상태로 보인다. 충격으로 표정과 상체가 잠시 멈춘 자세는 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "얼굴이 정면을 향하고 두 눈의 시선도 렌즈 부근에 모인다. 페드로가 카메라 바로 곁에 있다고 해석할 여지는 있으나, 실제 화면에서는 관객을 직접 바라보는 인상이 강하다. 방향을 따질 무기나 휴대 물체는 없다.",
        "built_space": "현우 뒤로 같은 계열의 낡은 금속 벽과 골진 천장이 보인다. 좌우에 창 두 구간씩, 왼쪽에 게시판 하나와 걸린 짙은 옷 한 벌, 오른쪽에 용기 선반 하나가 보인다. 오른쪽 아래에는 회색 수납 가구 하나와 그 위 직육면체 가전 하나가 있으며, 뒤쪽 바닥과 가구 주변에는 흩어진 물건이 남아 있다. 입구 부근의 인물과 실내 배경의 원근은 성립하며 불가능한 반사는 없다. 얼굴은 A보다 약간 크게 잡혔지만 주변 생활 설비도 비교적 또렷하게 드러난다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며 검은 머리, 회색 셔츠, 얼굴 윤곽이 현우의 참조와 대체로 부합한다. 얼굴의 멍과 코·입 주변 상처가 보인다. 정확한 나이와 국적 배경은 시각적으로 확정할 수 없다. 두 눈은 크게 떠져 있으나 안구 자체의 비정상적인 변형은 없다. 머리는 A보다 조금 더 정돈되어 보인다. 다른 인물이나 로봇, 읽을 수 있는 글자는 없다. 손목과 다리 부상, 컨테이너 번호는 프레임 밖이다.",
        "hard_violations": [],
        "physics": "머리와 목, 양어깨의 연결은 자연스럽고 상체는 수직으로 유지된다. 하체가 잘렸다는 이유로 지지가 없다고 볼 근거는 없다. 뒤의 가전은 수납 가구 위에, 용기들은 벽 선반 위에 놓여 있다. 정면으로 굳은 자세가 다소 초상사진처럼 보이지만 불가능한 신체 자세나 지지 없는 물체는 아니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 클로즈업 안에서 화면 밖 상대를 향한 시선과 커진 눈, 멈춘 표정이 페드로의 말에 충격받은 순간을 자연스럽게 구현한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "굳은 얼굴과 커진 눈은 명확하지만, 렌즈를 거의 정면으로 응시해 페드로에게 반응하는 장면보다 정면 초상처럼 읽힌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 화면 오른쪽의 카메라 밖 가까운 지점을 향한다. 상대는 보이지 않지만 페드로를 바라보며 말을 듣는 시선으로 성립한다. 무기나 방향을 확인해야 할 휴대 물체는 없다.",
        "built_space": "현우의 머리와 어깨 뒤로 낡은 밝은색 금속 벽과 골이 있는 천장이 보인다. 좌우 벽에 창이 두 구간씩 드러나고, 오른쪽에는 용기들이 놓인 벽 선반 하나와 그 아래 회색 수납 가구 하나가 보인다. 왼쪽에는 게시판 하나와 걸린 짙은 옷 한 벌이 있으며 뒤쪽에는 어질러진 생활 물건이 남아 있다. 입구 부근에 선 인물을 실내 방향으로 촬영한 배치가 성립한다. 참조에서 가려졌던 벽 구간의 창 위치까지 정확히 일치하는지는 확인하기 어렵지만, 명백한 중복 설비나 불가능한 반사는 없다.",
        "entities": "인물은 현우로 보이는 젊은 동아시아계 남성 한 명뿐이다. 앳된 얼굴, 헝클어진 검은 머리, 회색 셔츠가 참조와 대체로 일치하며 얼굴의 멍과 긁힌 상처도 유지된다. 정확한 나이와 한국계 미국인이라는 국적 배경은 외형만으로 확인할 수 없다. 눈은 정상적인 홍채와 동공을 유지한 채 커져 있고 입술이 약간 벌어져 있다. 다른 사람이나 로봇은 없고 읽을 수 있는 글자도 없다. 손목과 다리, 컨테이너 번호는 클로즈업 밖이므로 확인 대상에서 제외한다.",
        "hard_violations": [],
        "physics": "보이는 머리는 목에, 목은 셔츠를 입은 어깨와 몸통에 자연스럽게 연결된다. 발과 지면은 프레임 밖이지만 부유를 암시하는 자세는 없다. 선반 위 용기와 가구 위 물건은 받침면에 놓여 있고, 왼쪽 옷은 벽에 걸린 상태로 보인다. 충격으로 표정과 상체가 잠시 멈춘 자세는 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "얼굴이 정면을 향하고 두 눈의 시선도 렌즈 부근에 모인다. 페드로가 카메라 바로 곁에 있다고 해석할 여지는 있으나, 실제 화면에서는 관객을 직접 바라보는 인상이 강하다. 방향을 따질 무기나 휴대 물체는 없다.",
        "built_space": "현우 뒤로 같은 계열의 낡은 금속 벽과 골진 천장이 보인다. 좌우에 창 두 구간씩, 왼쪽에 게시판 하나와 걸린 짙은 옷 한 벌, 오른쪽에 용기 선반 하나가 보인다. 오른쪽 아래에는 회색 수납 가구 하나와 그 위 직육면체 가전 하나가 있으며, 뒤쪽 바닥과 가구 주변에는 흩어진 물건이 남아 있다. 입구 부근의 인물과 실내 배경의 원근은 성립하며 불가능한 반사는 없다. 얼굴은 A보다 약간 크게 잡혔지만 주변 생활 설비도 비교적 또렷하게 드러난다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며 검은 머리, 회색 셔츠, 얼굴 윤곽이 현우의 참조와 대체로 부합한다. 얼굴의 멍과 코·입 주변 상처가 보인다. 정확한 나이와 국적 배경은 시각적으로 확정할 수 없다. 두 눈은 크게 떠져 있으나 안구 자체의 비정상적인 변형은 없다. 머리는 A보다 조금 더 정돈되어 보인다. 다른 인물이나 로봇, 읽을 수 있는 글자는 없다. 손목과 다리 부상, 컨테이너 번호는 프레임 밖이다.",
        "hard_violations": [],
        "physics": "머리와 목, 양어깨의 연결은 자연스럽고 상체는 수직으로 유지된다. 하체가 잘렸다는 이유로 지지가 없다고 볼 근거는 없다. 뒤의 가전은 수납 가구 위에, 용기들은 벽 선반 위에 놓여 있다. 정면으로 굳은 자세가 다소 초상사진처럼 보이지만 불가능한 신체 자세나 지지 없는 물체는 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.667
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "프롬프트에서 요구한 '눈이 커진 채 얼어붙은 굳은 얼굴'의 표정과 감정 상태를 클로즈업 샷으로 매우 정확하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "충격을 받아 눈이 커진 표정이 아니라 눈물이 맺힌 슬픈 표정을 렌더링하여 프롬프트의 핵심적인 연기 지시를 놓쳤습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S24sh4_sel.png",
    "asset_id": "ceaba4c8-d476-45dc-b54a-b95f40c5aa73",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-07a8-75db-95f7-24ff54cc6ade",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S24sh4"
  }
 },
 "S24sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:59:28.607119+00:00",
  "fingerprint": "a77d5252a6438426334c4bc9da285f5b4733f5d83a518c05db97dc0b3596dc9d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S24sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S24sh7_sel.png",
  "source_sha256": "77d163f95db3559eed212fa4f7aa493391a0fb7fac11b417c41281171447298c",
  "file": "S24sh7_cine.png",
  "staged_sha256": "79918ec97e77038a0c21f59a3288d0e3595c87635632ed147587132c25002400",
  "latency_ms": 11232
 },
 "S24sh10::signage": {
  "fp": "4608fb3a9d635532",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S24sh10": {
  "input_fingerprint": "5ed56dfdba9a823a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 밖을 향해 뒷발로 지면을 강하게 밀어내며 질주하는 mid-stride 자세의 현우 전신 풀샷.\n\nLOCATION (lock): At the front doorway of the container home, opening onto the refugee settlement lane as the youth rushes outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Home entrance behind the left-to-right running figure in the middle-left of the frame, background; Clear continuation of the exit route in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 컨테이너 집 출입구 (Open for 현우's exit) — Viewed obliquely behind him on the left, with the passage into the home visible; used as Makes the direction away from the home unambiguous without overwhelming the figure; 집 밖 지면 (Visible beneath the running figure); used as Shows the rear foot's contact and provides space for the next stride.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient daytime illumination appropriate to the exterior with restrained tonal contrast and no invented directional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The damage and disorder inside container 7-31 remain uncleared. 현우: He is leaving the container at a run, with his wrists free. His facial bruising and dog-bitten leg remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 밖을 향해 뒷발로 지면을 강하게 밀어내며 질주하는 mid-stride 자세의 현우 전신 풀샷.\n\nLOCATION (lock): At the front doorway of the container home, opening onto the refugee settlement lane as the youth rushes outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Home entrance behind the left-to-right running figure in the middle-left of the frame, background; Clear continuation of the exit route in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 컨테이너 집 출입구 (Open for 현우's exit) — Viewed obliquely behind him on the left, with the passage into the home visible; used as Makes the direction away from the home unambiguous without overwhelming the figure; 집 밖 지면 (Visible beneath the running figure); used as Shows the rear foot's contact and provides space for the next stride.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient daytime illumination appropriate to the exterior with restrained tonal contrast and no invented directional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The damage and disorder inside container 7-31 remain uncleared. 현우: He is leaving the container at a run, with his wrists free. His facial bruising and dog-bitten leg remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 밖을 향해 뒷발로 지면을 강하게 밀어내며 질주하는 mid-stride 자세의 현우 전신 풀샷.\n\nLOCATION (lock): At the front doorway of the container home, opening onto the refugee settlement lane as the youth rushes outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Home entrance behind the left-to-right running figure in the middle-left of the frame, background; Clear continuation of the exit route in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 컨테이너 집 출입구 (Open for 현우's exit) — Viewed obliquely behind him on the left, with the passage into the home visible; used as Makes the direction away from the home unambiguous without overwhelming the figure; 집 밖 지면 (Visible beneath the running figure); used as Shows the rear foot's contact and provides space for the next stride.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient daytime illumination appropriate to the exterior with restrained tonal contrast and no invented directional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The damage and disorder inside container 7-31 remain uncleared. 현우: He is leaving the container at a run, with his wrists free. His facial bruising and dog-bitten leg remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선과 몸통 모두 화면 우측의 진행 방향을 향함.",
    "built_space": "좌측에 문이 열린 컨테이너 1동이 배치됨. 내부는 어둡고 정돈되어 있어 요구된 난장판 묘사가 없음. 외부 흙길은 레퍼런스와 일치함.",
    "entities": "현우의 얼굴과 복장은 레퍼런스와 일치하고 타박상이 있으나, 바지의 개 물림 자국은 보이지 않음.",
    "hard_violations": [],
    "physics": "오른쪽 앞발이 지면에 닿아 체중을 지탱하고 왼쪽 뒷발은 들려 있어, 뒷발로 지면을 강하게 밀어내라는 지시와 다름."
   },
   {
    "label": "B",
    "direction": "시선과 몸통 모두 화면 우측의 진행 방향을 향함.",
    "built_space": "좌측에 문이 열린 컨테이너 1동이 배치되며, 내부에 어질러진 집기들이 묘사됨. 컨테이너 외벽에 숫자 '7-31'이 명확히 적혀 있음.",
    "entities": "현우의 얼굴 타박상과 인상착의가 일치하며 바지에 찢어진 자국이 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 프롬프트 텍스트 유출 및 금지된 텍스트 노출 (컨테이너 외벽의 '7-31')",
     "[gpt-high] 출입구 왼쪽 외벽에 ‘7-31’이 선명하게 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "두 발이 모두 허공에 뜬 질주 체공 상태로, 지면에 닿은 발이 없어 뒷발로 밀어내는 순간의 접촉을 보여주지 못함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "텍스트 금지 조건을 준수했으나, 뒷발로 지면을 밀어내는 동작과 컨테이너 내부의 난장판 묘사가 누락됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "내부 묘사와 복장 디테일은 우수하나, 엄격히 금지된 읽을 수 있는 텍스트('7-31')가 화면에 유출되어 치명적 규정을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선과 몸통 모두 화면 우측의 진행 방향을 향함.",
        "built_space": "좌측에 문이 열린 컨테이너 1동이 배치됨. 내부는 어둡고 정돈되어 있어 요구된 난장판 묘사가 없음. 외부 흙길은 레퍼런스와 일치함.",
        "entities": "현우의 얼굴과 복장은 레퍼런스와 일치하고 타박상이 있으나, 바지의 개 물림 자국은 보이지 않음.",
        "hard_violations": [],
        "physics": "오른쪽 앞발이 지면에 닿아 체중을 지탱하고 왼쪽 뒷발은 들려 있어, 뒷발로 지면을 강하게 밀어내라는 지시와 다름."
       },
       {
        "label": "B",
        "direction": "시선과 몸통 모두 화면 우측의 진행 방향을 향함.",
        "built_space": "좌측에 문이 열린 컨테이너 1동이 배치되며, 내부에 어질러진 집기들이 묘사됨. 컨테이너 외벽에 숫자 '7-31'이 명확히 적혀 있음.",
        "entities": "현우의 얼굴 타박상과 인상착의가 일치하며 바지에 찢어진 자국이 묘사됨.",
        "hard_violations": [
         "프롬프트 텍스트 유출 및 금지된 텍스트 노출 (컨테이너 외벽의 '7-31')"
        ],
        "physics": "두 발이 모두 허공에 뜬 질주 체공 상태로, 지면에 닿은 발이 없어 뒷발로 밀어내는 순간의 접촉을 보여주지 못함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "텍스트 금지 조건을 준수했으나, 뒷발로 지면을 밀어내는 동작과 컨테이너 내부의 난장판 묘사가 누락됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "내부 묘사와 복장 디테일은 우수하나, 엄격히 금지된 읽을 수 있는 텍스트('7-31')가 화면에 유출되어 치명적 규정을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선과 몸통 모두 화면 우측의 진행 방향을 향함.",
        "built_space": "좌측에 문이 열린 컨테이너 1동이 배치됨. 내부는 어둡고 정돈되어 있어 요구된 난장판 묘사가 없음. 외부 흙길은 레퍼런스와 일치함.",
        "entities": "현우의 얼굴과 복장은 레퍼런스와 일치하고 타박상이 있으나, 바지의 개 물림 자국은 보이지 않음.",
        "hard_violations": [],
        "physics": "오른쪽 앞발이 지면에 닿아 체중을 지탱하고 왼쪽 뒷발은 들려 있어, 뒷발로 지면을 강하게 밀어내라는 지시와 다름."
       },
       {
        "label": "B",
        "direction": "시선과 몸통 모두 화면 우측의 진행 방향을 향함.",
        "built_space": "좌측에 문이 열린 컨테이너 1동이 배치되며, 내부에 어질러진 집기들이 묘사됨. 컨테이너 외벽에 숫자 '7-31'이 명확히 적혀 있음.",
        "entities": "현우의 얼굴 타박상과 인상착의가 일치하며 바지에 찢어진 자국이 묘사됨.",
        "hard_violations": [
         "프롬프트 텍스트 유출 및 금지된 텍스트 노출 (컨테이너 외벽의 '7-31')"
        ],
        "physics": "두 발이 모두 허공에 뜬 질주 체공 상태로, 지면에 닿은 발이 없어 뒷발로 밀어내는 순간의 접촉을 보여주지 못함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "집에서 오른쪽으로 달려 나오는 방향과 전신 구도는 맞지만, 판독 가능한 ‘7-31’ 표기가 금지 조건을 위반하고 뒷발의 지면 접촉도 없다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "문자 없이 현우의 전신과 자연스러운 지지를 구현했지만, 오른쪽 탈출로가 아닌 화면 왼쪽 앞을 향하며 뒷발로 지면을 미는 순간도 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 시선, 상체 및 앞으로 뻗은 다리가 화면 오른쪽 골목을 향한다. 열린 출입구는 인물 뒤 왼쪽에 있어 집에서 멀어지는 좌→우 이동이 명확하다. 무기나 손에 든 물건은 없다.",
        "built_space": "왼쪽에 열린 출입구 하나가 있고, 내부의 어질러진 바닥과 가구가 보인다. 출입구 왼쪽 가장자리에는 잘린 창 하나, 오른쪽 외벽에는 쇠창살 창 하나가 뚜렷하다. 녹슨 청회색 골판 금속벽, 받침 블록과 흙·자갈 골목은 장소 참조와 대체로 일치한다. 오른쪽에는 다음 보폭을 위한 골목이 열려 있지만 인물은 지정된 중간 왼쪽보다 중앙에 가깝다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며, 검은 머리와 마른 체격, 회색 셔츠·갈색 벨트·올리브색 카고 바지·낡은 부츠가 현우 참조에 부합한다. 옆얼굴에 상처가 있고 바지의 무릎과 종아리 부근이 찢어졌지만 개에 물린 상처 자체는 명확하지 않다. 양손목은 자유롭다. 외벽의 ‘7-31’은 명확히 읽힌다.",
        "hard_violations": [
         "출입구 왼쪽 외벽에 ‘7-31’이 선명하게 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "두 부츠 모두 지면에서 떨어져 있어 현재 지면에 닿아 몸을 지지하는 발은 없다. 다리의 전후 벌림과 팔 동작은 달리기의 공중 구간으로 가능한 자세이며, 앞발 아래에는 착지할 지면이 있다. 따라서 불가능한 부유로 단정할 수는 없지만, 요구된 뒷발 접촉과 강한 지면 밀어내기는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "얼굴과 시선은 화면 왼쪽 앞쪽을 향하고, 몸도 카메라 쪽으로 비스듬히 전진한다. 화면 오른쪽에 열린 골목은 있지만 그쪽으로 달려가는 좌→우 이동은 아니다. 손에 든 물건이나 겨누는 물체는 없다.",
        "built_space": "왼쪽에 출입구 하나와 바깥으로 열린 문짝 하나, 문손잡이 하나가 보인다. 문 안쪽 통로는 보이지만 어두워 내부의 훼손과 어질러짐은 확인하기 어렵다. 인물 뒤 외벽의 창과 먼 컨테이너들의 창, 녹슨 청회색 금속벽, 콘크리트 받침과 흙·자갈 골목은 참조 장소의 특성을 따른다. 오른쪽 이동 공간은 확보되지만 인물은 중간 왼쪽보다 중앙에 놓였다.",
        "entities": "젊은 동아시아계 남성 한 명이며, 앳된 얼굴과 헝클어진 검은 머리, 체격이 현우 참조에 가깝다. 회색 셔츠, 갈색 벨트, 올리브색 카고 바지와 부츠도 맞는다. 뺨의 멍과 자유로운 손목은 보이지만 바지에 가려진 다리의 개 물림 상처는 확인할 수 없다. 추가 인물이나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앞쪽 부츠가 흙바닥에 닿아 몸을 지지하고, 뒤쪽 다리는 무릎을 굽힌 채 들려 있다. 상체의 전진 기울기와 팔 동작은 달리기로 가능한 자세다. 다만 지면을 지지하는 것은 앞쪽 발이며, 요구된 뒷발의 강한 밀어내기 순간은 아니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "집에서 오른쪽으로 달려 나오는 방향과 전신 구도는 맞지만, 판독 가능한 ‘7-31’ 표기가 금지 조건을 위반하고 뒷발의 지면 접촉도 없다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "문자 없이 현우의 전신과 자연스러운 지지를 구현했지만, 오른쪽 탈출로가 아닌 화면 왼쪽 앞을 향하며 뒷발로 지면을 미는 순간도 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 시선, 상체 및 앞으로 뻗은 다리가 화면 오른쪽 골목을 향한다. 열린 출입구는 인물 뒤 왼쪽에 있어 집에서 멀어지는 좌→우 이동이 명확하다. 무기나 손에 든 물건은 없다.",
        "built_space": "왼쪽에 열린 출입구 하나가 있고, 내부의 어질러진 바닥과 가구가 보인다. 출입구 왼쪽 가장자리에는 잘린 창 하나, 오른쪽 외벽에는 쇠창살 창 하나가 뚜렷하다. 녹슨 청회색 골판 금속벽, 받침 블록과 흙·자갈 골목은 장소 참조와 대체로 일치한다. 오른쪽에는 다음 보폭을 위한 골목이 열려 있지만 인물은 지정된 중간 왼쪽보다 중앙에 가깝다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며, 검은 머리와 마른 체격, 회색 셔츠·갈색 벨트·올리브색 카고 바지·낡은 부츠가 현우 참조에 부합한다. 옆얼굴에 상처가 있고 바지의 무릎과 종아리 부근이 찢어졌지만 개에 물린 상처 자체는 명확하지 않다. 양손목은 자유롭다. 외벽의 ‘7-31’은 명확히 읽힌다.",
        "hard_violations": [
         "출입구 왼쪽 외벽에 ‘7-31’이 선명하게 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "두 부츠 모두 지면에서 떨어져 있어 현재 지면에 닿아 몸을 지지하는 발은 없다. 다리의 전후 벌림과 팔 동작은 달리기의 공중 구간으로 가능한 자세이며, 앞발 아래에는 착지할 지면이 있다. 따라서 불가능한 부유로 단정할 수는 없지만, 요구된 뒷발 접촉과 강한 지면 밀어내기는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "얼굴과 시선은 화면 왼쪽 앞쪽을 향하고, 몸도 카메라 쪽으로 비스듬히 전진한다. 화면 오른쪽에 열린 골목은 있지만 그쪽으로 달려가는 좌→우 이동은 아니다. 손에 든 물건이나 겨누는 물체는 없다.",
        "built_space": "왼쪽에 출입구 하나와 바깥으로 열린 문짝 하나, 문손잡이 하나가 보인다. 문 안쪽 통로는 보이지만 어두워 내부의 훼손과 어질러짐은 확인하기 어렵다. 인물 뒤 외벽의 창과 먼 컨테이너들의 창, 녹슨 청회색 금속벽, 콘크리트 받침과 흙·자갈 골목은 참조 장소의 특성을 따른다. 오른쪽 이동 공간은 확보되지만 인물은 중간 왼쪽보다 중앙에 놓였다.",
        "entities": "젊은 동아시아계 남성 한 명이며, 앳된 얼굴과 헝클어진 검은 머리, 체격이 현우 참조에 가깝다. 회색 셔츠, 갈색 벨트, 올리브색 카고 바지와 부츠도 맞는다. 뺨의 멍과 자유로운 손목은 보이지만 바지에 가려진 다리의 개 물림 상처는 확인할 수 없다. 추가 인물이나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앞쪽 부츠가 흙바닥에 닿아 몸을 지지하고, 뒤쪽 다리는 무릎을 굽힌 채 들려 있다. 상체의 전진 기울기와 팔 동작은 달리기로 가능한 자세다. 다만 지면을 지지하는 것은 앞쪽 발이며, 요구된 뒷발의 강한 밀어내기 순간은 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.9
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.65
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트 텍스트 유출 및 금지된 텍스트 노출 (컨테이너 외벽의 '7-31')",
     "[gpt-high] 출입구 왼쪽 외벽에 ‘7-31’이 선명하게 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 650
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "텍스트 금지 조건을 준수했으나, 뒷발로 지면을 밀어내는 동작과 컨테이너 내부의 난장판 묘사가 누락됨."
   },
   {
    "label": "B",
    "score": 650,
    "verdict_ko": "내부 묘사와 복장 디테일은 우수하나, 엄격히 금지된 읽을 수 있는 텍스트('7-31')가 화면에 유출되어 치명적 규정을 위반함.  ★위반: [gemini-pro] 프롬프트 텍스트 유출 및 금지된 텍스트 노출 (컨테이너 외벽의 '7-31') / [gpt-high] 출입구 왼쪽 외벽에 ‘7-31’이 선명하게 읽혀, 이미지 어디에도 판독 가능한 글자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh6_sel.png",
    "asset_id": "b7b093bd-37aa-46b1-bc2a-724de32daf4c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-095d-700f-8dd8-d6a8e13e4c96",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S20sh6"
  }
 },
 "S24sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:00:48.832617+00:00",
  "fingerprint": "ee0b422d50b45dee8bb317f776f1a7ea22715dea5b2f5684b7a76ca77a735efc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S24sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S24sh10_sel.png",
  "source_sha256": "5b4c5cb4aab9c176acfac7850fa339b6345fade975f62302b0a6d387d6be8d58",
  "file": "S24sh10_cine.png",
  "staged_sha256": "a9822130d53c0bb8ae0a3265b462b5fea45346401f103ae81a6c26bccaf9e50a",
  "latency_ms": 9404
 },
 "S25sh12::signage": {
  "fp": "edc9b51253de3050",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S25sh12": {
  "input_fingerprint": "c88577d82f8fbf77",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights around the tented reception are relighting after the blackout; the repaired jukebox remains at the party, and some lamps burst from excessive brightness. Charlie's chest ring emits an intensifying light, with his coat-and-hat disguise otherwise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights around the tented reception are relighting after the blackout; the repaired jukebox remains at the party, and some lamps burst from excessive brightness. Charlie's chest ring emits an intensifying light, with his coat-and-hat disguise otherwise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights around the tented reception are relighting after the blackout; the repaired jukebox remains at the party, and some lamps burst from excessive brightness. Charlie's chest ring emits an intensifying light, with his coat-and-hat disguise otherwise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12__bgfirst_bg.png",
     "asset_id": "44dbc893-09f5-4451-985f-95d270aac79f",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S25sh12.png",
     "asset_id": "b30b0e2b-b930-4260-9eaf-f19f33633ca4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_reception_clearing_3488b1.png",
     "asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 정면을 향하고 가슴 중앙에서 뿜어진 빛이 렌즈 쪽으로 향함.",
    "built_space": "공터 주변 양쪽 천막과 왼쪽 주크박스 및 배경의 가로등들이 기준 이미지와 동일하게 배치됨.",
    "entities": "코트와 모자를 착용한 로봇 찰리의 외형이 기준과 정확히 일치함.",
    "hard_violations": [],
    "physics": "두 발이 흙바닥 지면에 안정적으로 닿아 무게를 지탱함."
   },
   {
    "label": "B",
    "direction": "찰리가 카메라를 등지고 서 있으며 빛이 가슴이 아닌 등 부위에서 뿜어짐.",
    "built_space": "공터와 천막 및 가로등의 배치는 기준 이미지의 공간과 일치함.",
    "entities": "찰리의 뒷모습이 보이며 우측 천막 아래 지시되지 않은 다수의 인물이 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물들 임의 추가",
     "[gpt-high] 찰리만 등장하도록 제한한 장면에 왼쪽 천막 아래와 오른쪽 가장자리의 여러 손님을 추가했다."
    ],
    "physics": "찰리와 추가된 인물들 모두 지면 위에 온전히 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴에서 뿜어지는 빛과 지정된 와이드 샷 프레이밍을 정확하게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시되지 않은 구경꾼들이 임의로 추가된 치명적 위반이 있으며 가슴이 아닌 등에서 빛이 나옴."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 정면을 향하고 가슴 중앙에서 뿜어진 빛이 렌즈 쪽으로 향함.",
        "built_space": "공터 주변 양쪽 천막과 왼쪽 주크박스 및 배경의 가로등들이 기준 이미지와 동일하게 배치됨.",
        "entities": "코트와 모자를 착용한 로봇 찰리의 외형이 기준과 정확히 일치함.",
        "hard_violations": [],
        "physics": "두 발이 흙바닥 지면에 안정적으로 닿아 무게를 지탱함."
       },
       {
        "label": "B",
        "direction": "찰리가 카메라를 등지고 서 있으며 빛이 가슴이 아닌 등 부위에서 뿜어짐.",
        "built_space": "공터와 천막 및 가로등의 배치는 기준 이미지의 공간과 일치함.",
        "entities": "찰리의 뒷모습이 보이며 우측 천막 아래 지시되지 않은 다수의 인물이 존재함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물들 임의 추가"
        ],
        "physics": "찰리와 추가된 인물들 모두 지면 위에 온전히 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴에서 뿜어지는 빛과 지정된 와이드 샷 프레이밍을 정확하게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시되지 않은 구경꾼들이 임의로 추가된 치명적 위반이 있으며 가슴이 아닌 등에서 빛이 나옴."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 정면을 향하고 가슴 중앙에서 뿜어진 빛이 렌즈 쪽으로 향함.",
        "built_space": "공터 주변 양쪽 천막과 왼쪽 주크박스 및 배경의 가로등들이 기준 이미지와 동일하게 배치됨.",
        "entities": "코트와 모자를 착용한 로봇 찰리의 외형이 기준과 정확히 일치함.",
        "hard_violations": [],
        "physics": "두 발이 흙바닥 지면에 안정적으로 닿아 무게를 지탱함."
       },
       {
        "label": "B",
        "direction": "찰리가 카메라를 등지고 서 있으며 빛이 가슴이 아닌 등 부위에서 뿜어짐.",
        "built_space": "공터와 천막 및 가로등의 배치는 기준 이미지의 공간과 일치함.",
        "entities": "찰리의 뒷모습이 보이며 우측 천막 아래 지시되지 않은 다수의 인물이 존재함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물들 임의 추가"
        ],
        "physics": "찰리와 추가된 인물들 모두 지면 위에 온전히 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 주변 인물을 여러 명 추가했고, 찰리를 등지게 세워 가슴 대신 등에서 빛이 나므로 핵심 순간과 배치를 위반한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "비스듬한 하이앵글 와이드숏에서 찰리의 노출된 가슴 고리와 후방 양쪽 가로등의 발광을 정확히 연결하지만, 일부 등이 터지는 현상은 보이지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라에 등을 보이고 머리를 화면 오른쪽 난민촌 방향으로 돌렸다. 카메라 쪽으로 발광하는 고리는 가슴이 아니라 등판에 붙어 있다. 가로등은 공터 쪽으로 빛을 내며 화면 상단 양쪽에 분포한다. 오른쪽 손님 일부는 찰리 쪽을 보거나 손으로 눈을 가린다.",
        "built_space": "중앙 흙 공터, 왼쪽 대형 천막, 오른쪽 중경 천막, 오른쪽 전경 천막의 세 천막 영역이 보인다. 주크박스는 왼쪽 전경에 하나 있고, 양쪽 가까운 사각 투광등 두 개와 중·원경의 다수 가로등, 컨테이너 주거지와 전선이 참조 장소를 따른다. 찰리는 공터의 가까운 왼쪽에 크게 서서 천막 일부를 가린다. 왼쪽 천막 아래와 오른쪽 가장자리에는 허용되지 않은 손님들이 배치되어 있다.",
        "entities": "찰리 한 개체의 모자, 긴 코트, 베이지색 어깨·팔 장갑과 짧은 다리는 참조와 대체로 맞지만 흰 얼굴은 옆면 일부만 보인다. 핵심 소품인 발광 고리는 앞가슴이 아닌 등에 나타난다. 주크박스 하나, 피로연 탁자와 의자, 천막, 가로등, 일몰의 난민촌이 보인다. 찰리 외에 성인과 작은 체구의 인물들이 여러 명 추가되었으며, 이 거리에서 각자의 연령과 민족적 외양은 확정하기 어렵다. 터지는 전구는 식별되지 않는다.",
        "hard_violations": [
         "찰리만 등장하도록 제한한 장면에 왼쪽 천막 아래와 오른쪽 가장자리의 여러 손님을 추가했다."
        ],
        "physics": "찰리의 두 발은 흙바닥에 닿아 체중을 지탱하고 두 팔은 어깨에서 자연스럽게 내려온다. 손님들도 지면에 서 있으며, 가로등은 기둥에, 천막은 기둥과 줄에 지지되어 있다. 주크박스와 가구도 바닥에 놓여 있다. 떠 있는 신체나 지지 없는 물체는 보이지 않는다. 다만 발광원의 위치가 요구된 가슴과 다르다."
       },
       {
        "label": "B",
        "direction": "찰리의 가슴은 카메라를 향해 열려 있고 얼굴은 화면 오른쪽을 조금 바라본다. 가슴 고리의 빛은 전방과 주변으로 퍼진다. 켜진 가로등들이 가슴보다 위쪽의 후방 좌우에 분포하며 공터를 비춘다. 시선이 향해야 할 특정 상대나 조준 대상은 요구되지 않았다.",
        "built_space": "비스듬히 내려다본 넓은 공터 중앙에 찰리가 서고, 가로등과 사이에 빈 지면이 충분히 드러난다. 왼쪽 대형 천막, 오른쪽 중경 천막, 오른쪽 전경 천막의 세 영역에서 지붕 윗면과 측면을 볼 수 있다. 왼쪽 전경 주크박스는 하나이며, 양쪽 가까운 사각 투광등 두 개와 중·원경의 다수 가로등이 주거지까지 이어진다. 원형 식탁, 접이식 의자, 드럼통, 컨테이너와 전선이 참조 장소의 구성을 유지한다.",
        "entities": "등장 개체는 찰리 하나뿐이다. 흰 마스크형 얼굴의 점·선 디테일, 베이지색 각진 장갑, 육중한 긴 팔과 짧은 다리, 어두운 모자와 코트가 참조에 부합한다. 가슴에는 밝은 원형 테두리가 식별되며 중심부도 강하게 빛난다. 주크박스 하나와 피로연 시설, 주변 가로등, 난민촌 및 일몰이 모두 보인다. 읽을 수 있는 문구나 추가 인물은 보이지 않는다. 일부 등이 과도한 밝기로 터지는 모습은 확인되지 않는다.",
        "hard_violations": [],
        "physics": "찰리는 벌린 두 발을 흙바닥에 디디고 약간 굽힌 다리로 몸을 지탱한다. 긴 팔과 손은 몸 양옆에 자연스럽게 내려와 있고, 모자와 코트는 몸에 걸쳐져 있다. 가로등은 기둥에 고정되고 천막은 기둥과 줄에 지지되며 주크박스와 가구는 지면에 놓여 있다. 발광은 실제 가슴 장치에서 시작하고 주변 젖은 지면에도 빛이 반응한다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 주변 인물을 여러 명 추가했고, 찰리를 등지게 세워 가슴 대신 등에서 빛이 나므로 핵심 순간과 배치를 위반한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "비스듬한 하이앵글 와이드숏에서 찰리의 노출된 가슴 고리와 후방 양쪽 가로등의 발광을 정확히 연결하지만, 일부 등이 터지는 현상은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 카메라에 등을 보이고 머리를 화면 오른쪽 난민촌 방향으로 돌렸다. 카메라 쪽으로 발광하는 고리는 가슴이 아니라 등판에 붙어 있다. 가로등은 공터 쪽으로 빛을 내며 화면 상단 양쪽에 분포한다. 오른쪽 손님 일부는 찰리 쪽을 보거나 손으로 눈을 가린다.",
        "built_space": "중앙 흙 공터, 왼쪽 대형 천막, 오른쪽 중경 천막, 오른쪽 전경 천막의 세 천막 영역이 보인다. 주크박스는 왼쪽 전경에 하나 있고, 양쪽 가까운 사각 투광등 두 개와 중·원경의 다수 가로등, 컨테이너 주거지와 전선이 참조 장소를 따른다. 찰리는 공터의 가까운 왼쪽에 크게 서서 천막 일부를 가린다. 왼쪽 천막 아래와 오른쪽 가장자리에는 허용되지 않은 손님들이 배치되어 있다.",
        "entities": "찰리 한 개체의 모자, 긴 코트, 베이지색 어깨·팔 장갑과 짧은 다리는 참조와 대체로 맞지만 흰 얼굴은 옆면 일부만 보인다. 핵심 소품인 발광 고리는 앞가슴이 아닌 등에 나타난다. 주크박스 하나, 피로연 탁자와 의자, 천막, 가로등, 일몰의 난민촌이 보인다. 찰리 외에 성인과 작은 체구의 인물들이 여러 명 추가되었으며, 이 거리에서 각자의 연령과 민족적 외양은 확정하기 어렵다. 터지는 전구는 식별되지 않는다.",
        "hard_violations": [
         "찰리만 등장하도록 제한한 장면에 왼쪽 천막 아래와 오른쪽 가장자리의 여러 손님을 추가했다."
        ],
        "physics": "찰리의 두 발은 흙바닥에 닿아 체중을 지탱하고 두 팔은 어깨에서 자연스럽게 내려온다. 손님들도 지면에 서 있으며, 가로등은 기둥에, 천막은 기둥과 줄에 지지되어 있다. 주크박스와 가구도 바닥에 놓여 있다. 떠 있는 신체나 지지 없는 물체는 보이지 않는다. 다만 발광원의 위치가 요구된 가슴과 다르다."
       },
       {
        "label": "A",
        "direction": "찰리의 가슴은 카메라를 향해 열려 있고 얼굴은 화면 오른쪽을 조금 바라본다. 가슴 고리의 빛은 전방과 주변으로 퍼진다. 켜진 가로등들이 가슴보다 위쪽의 후방 좌우에 분포하며 공터를 비춘다. 시선이 향해야 할 특정 상대나 조준 대상은 요구되지 않았다.",
        "built_space": "비스듬히 내려다본 넓은 공터 중앙에 찰리가 서고, 가로등과 사이에 빈 지면이 충분히 드러난다. 왼쪽 대형 천막, 오른쪽 중경 천막, 오른쪽 전경 천막의 세 영역에서 지붕 윗면과 측면을 볼 수 있다. 왼쪽 전경 주크박스는 하나이며, 양쪽 가까운 사각 투광등 두 개와 중·원경의 다수 가로등이 주거지까지 이어진다. 원형 식탁, 접이식 의자, 드럼통, 컨테이너와 전선이 참조 장소의 구성을 유지한다.",
        "entities": "등장 개체는 찰리 하나뿐이다. 흰 마스크형 얼굴의 점·선 디테일, 베이지색 각진 장갑, 육중한 긴 팔과 짧은 다리, 어두운 모자와 코트가 참조에 부합한다. 가슴에는 밝은 원형 테두리가 식별되며 중심부도 강하게 빛난다. 주크박스 하나와 피로연 시설, 주변 가로등, 난민촌 및 일몰이 모두 보인다. 읽을 수 있는 문구나 추가 인물은 보이지 않는다. 일부 등이 과도한 밝기로 터지는 모습은 확인되지 않는다.",
        "hard_violations": [],
        "physics": "찰리는 벌린 두 발을 흙바닥에 디디고 약간 굽힌 다리로 몸을 지탱한다. 긴 팔과 손은 몸 양옆에 자연스럽게 내려와 있고, 모자와 코트는 몸에 걸쳐져 있다. 가로등은 기둥에 고정되고 천막은 기둥과 줄에 지지되며 주크박스와 가구는 지면에 놓여 있다. 발광은 실제 가슴 장치에서 시작하고 주변 젖은 지면에도 빛이 반응한다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.651
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.401
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물들 임의 추가",
     "[gpt-high] 찰리만 등장하도록 제한한 장면에 왼쪽 천막 아래와 오른쪽 가장자리의 여러 손님을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 401
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "가슴에서 뿜어지는 빛과 지정된 와이드 샷 프레이밍을 정확하게 구현함."
   },
   {
    "label": "B",
    "score": 401,
    "verdict_ko": "지시되지 않은 구경꾼들이 임의로 추가된 치명적 위반이 있으며 가슴이 아닌 등에서 빛이 나옴.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물들 임의 추가 / [gpt-high] 찰리만 등장하도록 제한한 장면에 왼쪽 천막 아래와 오른쪽 가장자리의 여러 손님을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_reception_clearing_3488b1.png",
    "asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-0b0c-7b07-aac3-2d7056188b30",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12__bgfirst_bg.png",
   "bg_asset_id": "44dbc893-09f5-4451-985f-95d270aac79f",
   "bg_record_key": "S25sh12::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "reception_clearing",
   "groupbg_asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S25sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:01:48.172290+00:00",
  "fingerprint": "0b3e98e45aef3684e04df0dc81944cd8f41f9bd8c51dcc0f0b85100c5c1136e7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S25sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S25sh12_sel.png",
  "source_sha256": "733ee9eb1cb0c8f4e54069fd7209f736f0696894e02369773ccf6ae9efda6c9d",
  "file": "S25sh12_cine.png",
  "staged_sha256": "84d2d9744f39d37cb43777009a639165499cf4816c38321ec7103fc2b75c658b",
  "latency_ms": 11571
 },
 "S25sh19::signage": {
  "fp": "2168f4ffa565587f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S25sh19": {
  "input_fingerprint": "1456e75061070dca",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 트럭에서 내린 채 험악한 표정으로 앞을 노려보는 박철진과 그 뒤의 민병대원들 전신.\n\nLOCATION (lock): At the vehicle-access edge of the outdoor reception lot, beside a newly arrived militia truck. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 민병대 트럭 (Stopped after arriving, with the men now dismounted) — Its front-side quarter is visible behind the figures, occupying no more than the rear-right third; used as Establishes the source of the intrusion and supplies a grounded scale reference; 피로연장 입구와 공터 (The arrival interrupts the gathering in the clearing); used as Open space toward the left edge receives the men's outward-directed attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The reception remains brightly illuminated after the restoration of power, with controlled contrast preserving the severity of 박철진's expression.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the reception clearing, temporary tent structures, and streetlights that remain illuminated. Exclude the transient radiance emitted by the robot and any bulbs caught in the act of bursting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception and refugee settlement remain brightly illuminated, with some lamps burst and the jukebox playing again. A militia truck has stopped at the tented clearing, and Charlie retains his old coat and hat. 박철진: He has disembarked from the militia truck at the reception.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 트럭에서 내린 채 험악한 표정으로 앞을 노려보는 박철진과 그 뒤의 민병대원들 전신.\n\nLOCATION (lock): At the vehicle-access edge of the outdoor reception lot, beside a newly arrived militia truck. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 민병대 트럭 (Stopped after arriving, with the men now dismounted) — Its front-side quarter is visible behind the figures, occupying no more than the rear-right third; used as Establishes the source of the intrusion and supplies a grounded scale reference; 피로연장 입구와 공터 (The arrival interrupts the gathering in the clearing); used as Open space toward the left edge receives the men's outward-directed attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The reception remains brightly illuminated after the restoration of power, with controlled contrast preserving the severity of 박철진's expression.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the reception clearing, temporary tent structures, and streetlights that remain illuminated. Exclude the transient radiance emitted by the robot and any bulbs caught in the act of bursting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception and refugee settlement remain brightly illuminated, with some lamps burst and the jukebox playing again. A militia truck has stopped at the tented clearing, and Charlie retains his old coat and hat. 박철진: He has disembarked from the militia truck at the reception.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 트럭에서 내린 채 험악한 표정으로 앞을 노려보는 박철진과 그 뒤의 민병대원들 전신.\n\nLOCATION (lock): At the vehicle-access edge of the outdoor reception lot, beside a newly arrived militia truck. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 민병대 트럭 (Stopped after arriving, with the men now dismounted) — Its front-side quarter is visible behind the figures, occupying no more than the rear-right third; used as Establishes the source of the intrusion and supplies a grounded scale reference; 피로연장 입구와 공터 (The arrival interrupts the gathering in the clearing); used as Open space toward the left edge receives the men's outward-directed attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The reception remains brightly illuminated after the restoration of power, with controlled contrast preserving the severity of 박철진's expression.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the reception clearing, temporary tent structures, and streetlights that remain illuminated. Exclude the transient radiance emitted by the robot and any bulbs caught in the act of bursting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception and refugee settlement remain brightly illuminated, with some lamps burst and the jukebox playing again. A militia truck has stopped at the tented clearing, and Charlie retains his old coat and hat. 박철진: He has disembarked from the militia truck at the reception.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진과 뒤편의 민병대원들이 험악한 표정으로 정면(카메라 방향)을 노려보고 있음.",
    "built_space": "참조 이미지와 동일한 텐트와 조명이 유지되었고, 왼쪽 근경에 주크박스가 올바른 크기로 위치함. 우측 후방에 트럭이 적절히 배치됨.",
    "entities": "박철진의 얼굴과 복장(붉은 완장 포함)이 참조와 일치하며 다국적 민병대원들이 뒤에 서 있음.",
    "hard_violations": [],
    "physics": "모든 인물이 지면에 안정적으로 서 있고 무기를 쥔 손의 묘사도 자연스러움."
   },
   {
    "label": "B",
    "direction": "박철진과 대원들이 정면을 응시하며 서 있음.",
    "built_space": "기존 공간과 유사하나, 왼쪽 근경에 있어야 할 주크박스가 축소되어 후경 텐트 근처로 이동됨.",
    "entities": "박철진과 민병대원들이 존재하지만, 우측 트럭 표면에 명확한 한글 텍스트가 쓰여 있음.",
    "hard_violations": [
     "[gemini-pro] 트럭 측면에 '민병대' 등 읽을 수 있는 텍스트가 노출됨 (No readable writing 지시 위반)",
     "[gemini-pro] 참조 이미지에 명확히 존재하는 주크박스의 크기와 위치를 임의로 변경함",
     "[gpt-high] 트럭 표면에 읽을 수 있는 한글 표기가 노출되어 글자 노출 금지 조건을 위반한다."
    ],
    "physics": "인물들이 지면에 올바르게 서 있거나 트럭에 자연스럽게 기대어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "참조 이미지의 공간적 요소(주크박스 위치 등)와 인물 외형을 정확히 유지하며, 텍스트 금지 지시를 포함한 모든 조건을 훌륭하게 충족함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "트럭 측면에 금지된 텍스트가 노출되었으며, 주크박스의 위치와 크기가 임의로 왜곡되어 치명적인 오류가 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진과 뒤편의 민병대원들이 험악한 표정으로 정면(카메라 방향)을 노려보고 있음.",
        "built_space": "참조 이미지와 동일한 텐트와 조명이 유지되었고, 왼쪽 근경에 주크박스가 올바른 크기로 위치함. 우측 후방에 트럭이 적절히 배치됨.",
        "entities": "박철진의 얼굴과 복장(붉은 완장 포함)이 참조와 일치하며 다국적 민병대원들이 뒤에 서 있음.",
        "hard_violations": [],
        "physics": "모든 인물이 지면에 안정적으로 서 있고 무기를 쥔 손의 묘사도 자연스러움."
       },
       {
        "label": "B",
        "direction": "박철진과 대원들이 정면을 응시하며 서 있음.",
        "built_space": "기존 공간과 유사하나, 왼쪽 근경에 있어야 할 주크박스가 축소되어 후경 텐트 근처로 이동됨.",
        "entities": "박철진과 민병대원들이 존재하지만, 우측 트럭 표면에 명확한 한글 텍스트가 쓰여 있음.",
        "hard_violations": [
         "트럭 측면에 '민병대' 등 읽을 수 있는 텍스트가 노출됨 (No readable writing 지시 위반)",
         "참조 이미지에 명확히 존재하는 주크박스의 크기와 위치를 임의로 변경함"
        ],
        "physics": "인물들이 지면에 올바르게 서 있거나 트럭에 자연스럽게 기대어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "참조 이미지의 공간적 요소(주크박스 위치 등)와 인물 외형을 정확히 유지하며, 텍스트 금지 지시를 포함한 모든 조건을 훌륭하게 충족함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "트럭 측면에 금지된 텍스트가 노출되었으며, 주크박스의 위치와 크기가 임의로 왜곡되어 치명적인 오류가 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진과 뒤편의 민병대원들이 험악한 표정으로 정면(카메라 방향)을 노려보고 있음.",
        "built_space": "참조 이미지와 동일한 텐트와 조명이 유지되었고, 왼쪽 근경에 주크박스가 올바른 크기로 위치함. 우측 후방에 트럭이 적절히 배치됨.",
        "entities": "박철진의 얼굴과 복장(붉은 완장 포함)이 참조와 일치하며 다국적 민병대원들이 뒤에 서 있음.",
        "hard_violations": [],
        "physics": "모든 인물이 지면에 안정적으로 서 있고 무기를 쥔 손의 묘사도 자연스러움."
       },
       {
        "label": "B",
        "direction": "박철진과 대원들이 정면을 응시하며 서 있음.",
        "built_space": "기존 공간과 유사하나, 왼쪽 근경에 있어야 할 주크박스가 축소되어 후경 텐트 근처로 이동됨.",
        "entities": "박철진과 민병대원들이 존재하지만, 우측 트럭 표면에 명확한 한글 텍스트가 쓰여 있음.",
        "hard_violations": [
         "트럭 측면에 '민병대' 등 읽을 수 있는 텍스트가 노출됨 (No readable writing 지시 위반)",
         "참조 이미지에 명확히 존재하는 주크박스의 크기와 위치를 임의로 변경함"
        ],
        "physics": "인물들이 지면에 올바르게 서 있거나 트럭에 자연스럽게 기대어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전신 와이드 구도와 박철진의 복장은 맞지만, 트럭에 읽을 수 있는 한글이 노출되고 트럭이 오른쪽 전경까지 크게 차지하며 대원 한 명은 아직 하차 중이다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "험악한 박철진과 하차를 마친 후방 대원들의 전신, 밝은 일몰 피로연장을 더 충실히 구현했으나 트럭의 점유 범위와 다민족으로 보이는 대원 구성은 지시와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 왼쪽 공터보다는 카메라에 가까운 정면을 노려본다. 뒤쪽 대원 대부분도 정면을 보고, 트럭 바로 옆 두 명은 발판이나 하차하는 동료 쪽으로 시선을 내린다. 보이는 소총 총구들은 지면을 향하며 특정 인물을 겨누지 않는다. 트럭 앞부분은 화면 오른쪽 바깥을 향한다.",
        "built_space": "왼쪽에 연결된 천막 지붕과 여러 원형 식탁·접이식 의자, 전구 줄, 드럼통들이 있고 뒤로 컨테이너 정착지와 켜진 가로등들이 이어진다. 주크박스 한 대는 왼쪽 중경에 있다. 트럭 한 대는 오른쪽에 있지만 큰 앞바퀴와 운전석이 전경까지 내려와 지정된 후방 오른쪽 영역에 머물지 않는다. 박철진 뒤로 대원 여섯 명이 있으며, 가장 오른쪽 대원은 발판과 지면 사이에서 아직 내려오는 자세다.",
        "entities": "박철진은 짧은 검은 머리의 중년 한국인 남성으로 보이며 얼굴, 짙은 군복, 붉은 완장, 장비 벨트와 검은 군화가 인물 참조와 대체로 맞는다. 후방에는 무장 대원 여섯 명이 있고 일부는 한국인으로 명확하게 읽히지 않는다. 천막형 피로연장, 군용 트럭, 주크박스와 일몰은 보이며 이전 장면의 로봇과 방사광은 없다. 트럭 문에 '민병대', 보닛에 '인명대'로 읽히는 한글이 드러난다.",
        "hard_violations": [
         "트럭 표면에 읽을 수 있는 한글 표기가 노출되어 글자 노출 금지 조건을 위반한다."
        ],
        "physics": "박철진과 뒤쪽 대원들은 군화를 지면에 딛고 있다. 하차하는 오른쪽 대원은 한쪽 발이 지면에 닿고 반대쪽 다리는 발판 부근에 굽혀져 있어 체중을 지탱할 수 있다. 소총은 손과 몸에 걸린 끈으로 지지되고 트럭은 타이어로 지면에 놓인다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "박철진은 얼굴과 시선을 카메라 정면보다 약간 화면 왼쪽의 전방 공터로 향한다. 뒤쪽 대원들도 대체로 같은 전방을 경계하지만 왼쪽 가장자리로 강하게 시선이 모이는 구성은 아니다. 소총 총구는 아래쪽 지면으로 내려가 있고 특정 인물을 겨누지 않는다. 트럭은 앞면과 오른쪽 측면이 보이는 사선 방향으로 정차해 있다.",
        "built_space": "왼쪽에 연결된 피로연 천막, 원형 식탁과 접이식 의자, 전구 줄과 드럼통들이 있고, 왼쪽 전경에 주크박스 한 대가 있다. 오른쪽 전경에는 천막 일부와 식탁·의자가 남아 있으며 후방의 컨테이너와 켜진 가로등도 참조 장소와 잘 이어진다. 트럭 한 대 앞에 박철진과 후방 대원 일곱 명이 모두 지면에 서 있다. 트럭은 인물 뒤에 배치됐지만 중앙까지 뻗어 후방 오른쪽 3분의 1 제한보다 넓게 보인다.",
        "entities": "박철진의 중년 남성 얼굴, 짧은 검은 머리, 체격, 짙은 군복과 붉은 완장, 장비 벨트 및 군화는 참조와 대체로 일치한다. 뒤에는 여성으로 보이는 대원 한 명을 포함한 일곱 명이 있으며 여러 인종적 외양이 섞여 보여 한국인 현지인 기본 설정과 어긋난다. 군용 트럭, 피로연 천막, 주크박스, 켜진 조명과 일몰이 보인다. 이전 장면의 로봇이나 방사광은 없고 읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "박철진과 대원들은 모두 발을 땅에 딛고 서 있으며, 뒤 인물의 일부 가려진 다리도 서 있는 자세와 모순되지 않는다. 보이는 소총은 손으로 잡거나 멜빵으로 몸에 걸쳐 지지한다. 트럭의 타이어는 지면에 닿고 천막은 기둥과 줄로 지탱된다. 전원이 하차한 상태로 물리적으로 가능한 배치다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전신 와이드 구도와 박철진의 복장은 맞지만, 트럭에 읽을 수 있는 한글이 노출되고 트럭이 오른쪽 전경까지 크게 차지하며 대원 한 명은 아직 하차 중이다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "험악한 박철진과 하차를 마친 후방 대원들의 전신, 밝은 일몰 피로연장을 더 충실히 구현했으나 트럭의 점유 범위와 다민족으로 보이는 대원 구성은 지시와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 왼쪽 공터보다는 카메라에 가까운 정면을 노려본다. 뒤쪽 대원 대부분도 정면을 보고, 트럭 바로 옆 두 명은 발판이나 하차하는 동료 쪽으로 시선을 내린다. 보이는 소총 총구들은 지면을 향하며 특정 인물을 겨누지 않는다. 트럭 앞부분은 화면 오른쪽 바깥을 향한다.",
        "built_space": "왼쪽에 연결된 천막 지붕과 여러 원형 식탁·접이식 의자, 전구 줄, 드럼통들이 있고 뒤로 컨테이너 정착지와 켜진 가로등들이 이어진다. 주크박스 한 대는 왼쪽 중경에 있다. 트럭 한 대는 오른쪽에 있지만 큰 앞바퀴와 운전석이 전경까지 내려와 지정된 후방 오른쪽 영역에 머물지 않는다. 박철진 뒤로 대원 여섯 명이 있으며, 가장 오른쪽 대원은 발판과 지면 사이에서 아직 내려오는 자세다.",
        "entities": "박철진은 짧은 검은 머리의 중년 한국인 남성으로 보이며 얼굴, 짙은 군복, 붉은 완장, 장비 벨트와 검은 군화가 인물 참조와 대체로 맞는다. 후방에는 무장 대원 여섯 명이 있고 일부는 한국인으로 명확하게 읽히지 않는다. 천막형 피로연장, 군용 트럭, 주크박스와 일몰은 보이며 이전 장면의 로봇과 방사광은 없다. 트럭 문에 '민병대', 보닛에 '인명대'로 읽히는 한글이 드러난다.",
        "hard_violations": [
         "트럭 표면에 읽을 수 있는 한글 표기가 노출되어 글자 노출 금지 조건을 위반한다."
        ],
        "physics": "박철진과 뒤쪽 대원들은 군화를 지면에 딛고 있다. 하차하는 오른쪽 대원은 한쪽 발이 지면에 닿고 반대쪽 다리는 발판 부근에 굽혀져 있어 체중을 지탱할 수 있다. 소총은 손과 몸에 걸린 끈으로 지지되고 트럭은 타이어로 지면에 놓인다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "박철진은 얼굴과 시선을 카메라 정면보다 약간 화면 왼쪽의 전방 공터로 향한다. 뒤쪽 대원들도 대체로 같은 전방을 경계하지만 왼쪽 가장자리로 강하게 시선이 모이는 구성은 아니다. 소총 총구는 아래쪽 지면으로 내려가 있고 특정 인물을 겨누지 않는다. 트럭은 앞면과 오른쪽 측면이 보이는 사선 방향으로 정차해 있다.",
        "built_space": "왼쪽에 연결된 피로연 천막, 원형 식탁과 접이식 의자, 전구 줄과 드럼통들이 있고, 왼쪽 전경에 주크박스 한 대가 있다. 오른쪽 전경에는 천막 일부와 식탁·의자가 남아 있으며 후방의 컨테이너와 켜진 가로등도 참조 장소와 잘 이어진다. 트럭 한 대 앞에 박철진과 후방 대원 일곱 명이 모두 지면에 서 있다. 트럭은 인물 뒤에 배치됐지만 중앙까지 뻗어 후방 오른쪽 3분의 1 제한보다 넓게 보인다.",
        "entities": "박철진의 중년 남성 얼굴, 짧은 검은 머리, 체격, 짙은 군복과 붉은 완장, 장비 벨트 및 군화는 참조와 대체로 일치한다. 뒤에는 여성으로 보이는 대원 한 명을 포함한 일곱 명이 있으며 여러 인종적 외양이 섞여 보여 한국인 현지인 기본 설정과 어긋난다. 군용 트럭, 피로연 천막, 주크박스, 켜진 조명과 일몰이 보인다. 이전 장면의 로봇이나 방사광은 없고 읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "박철진과 대원들은 모두 발을 땅에 딛고 서 있으며, 뒤 인물의 일부 가려진 다리도 서 있는 자세와 모순되지 않는다. 보이는 소총은 손으로 잡거나 멜빵으로 몸에 걸쳐 지지한다. 트럭의 타이어는 지면에 닿고 천막은 기둥과 줄로 지탱된다. 전원이 하차한 상태로 물리적으로 가능한 배치다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.857
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.607
   },
   "violations": {
    "B": [
     "[gemini-pro] 트럭 측면에 '민병대' 등 읽을 수 있는 텍스트가 노출됨 (No readable writing 지시 위반)",
     "[gemini-pro] 참조 이미지에 명확히 존재하는 주크박스의 크기와 위치를 임의로 변경함",
     "[gpt-high] 트럭 표면에 읽을 수 있는 한글 표기가 노출되어 글자 노출 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 607
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "참조 이미지의 공간적 요소(주크박스 위치 등)와 인물 외형을 정확히 유지하며, 텍스트 금지 지시를 포함한 모든 조건을 훌륭하게 충족함."
   },
   {
    "label": "B",
    "score": 607,
    "verdict_ko": "트럭 측면에 금지된 텍스트가 노출되었으며, 주크박스의 위치와 크기가 임의로 왜곡되어 치명적인 오류가 발생함.  ★위반: [gemini-pro] 트럭 측면에 '민병대' 등 읽을 수 있는 텍스트가 노출됨 (No readable writing 지시 위반) / [gemini-pro] 참조 이미지에 명확히 존재하는 주크박스의 크기와 위치를 임의로 변경함 / [gpt-high] 트럭 표면에 읽을 수 있는 한글 표기가 노출되어 글자 노출 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12_sel.png",
    "asset_id": "af0a60a2-76e1-405b-96d9-eb975306bf64",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-1001-7509-9334-3d94ac114552",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S25sh12"
  }
 },
 "S25sh19::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:02:53.600429+00:00",
  "fingerprint": "bbbd6510be567137a5cadf435d178458015654e6c2f3de17c9b0aaf9a2194f74",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S25sh19_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S25sh19_sel.png",
  "source_sha256": "dbc0f649fd386153b6f6149659a1e01ce75fb25886e8a1ae6c1c653112e448fa",
  "file": "S25sh19_cine.png",
  "staged_sha256": "59b2f8a6fe446bcbedaf14f493ed160e35cedc94f9f254102413a82ffd2d2433",
  "latency_ms": 10445
 },
 "S25sh22::signage": {
  "fp": "5063e84ef22f5e3f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::6fefa20797430b9e": {
  "subjects": [],
  "subject_text": "인천 난민촌 피로연장\n컨테이너 주거지의 빈 공터에 천막을 친 간이 행사 공간. 주변에 전등과 가로등이 있으며 한쪽에 주크박스가 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L174",
  "scope_role": "location_exterior",
  "scope_sha": "d226e83a7d00860b"
 },
 "S25sh22::bgfirst_bg": {
  "input_fingerprint": "c7f97fb556fb8097",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh22__bgfirst_bg.png",
  "asset_id": "4ea5365d-bc0a-4f55-9bd5-c4357af778b0",
  "input_asset_ids": [
   "ed35ed23-b843-4470-bb40-db03b4117fa0",
   "eac29610-01b6-43a7-9e6e-9e2387ddd674"
  ]
 },
 "S25sh22": {
  "input_fingerprint": "92c77035619899a4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The settlement's restored lighting remains established, although this alley provides concealment behind a wall. At the reception, the tents, parked militia truck, jukebox and burst lamps remain; Charlie still wears his old coat and hat. 현우: He has taken cover behind the alley wall, still wearing his outer shirt. His facial bruises and leg injury persist. 앰버: She is hiding behind the alley wall with her mask and waist tool pouch retained. 라울: He is hiding behind the alley wall after running. 페드로: He is hiding behind the alley wall after running.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The settlement's restored lighting remains established, although this alley provides concealment behind a wall. At the reception, the tents, parked militia truck, jukebox and burst lamps remain; Charlie still wears his old coat and hat. 현우: He has taken cover behind the alley wall, still wearing his outer shirt. His facial bruises and leg injury persist. 앰버: She is hiding behind the alley wall with her mask and waist tool pouch retained. 라울: He is hiding behind the alley wall after running. 페드로: He is hiding behind the alley wall after running.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The settlement's restored lighting remains established, although this alley provides concealment behind a wall. At the reception, the tents, parked militia truck, jukebox and burst lamps remain; Charlie still wears his old coat and hat. 현우: He has taken cover behind the alley wall, still wearing his outer shirt. His facial bruises and leg injury persist. 앰버: She is hiding behind the alley wall with her mask and waist tool pouch retained. 라울: He is hiding behind the alley wall after running. 페드로: He is hiding behind the alley wall after running.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh22__bgfirst_bg.png",
     "asset_id": "4ea5365d-bc0a-4f55-9bd5-c4357af778b0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S25sh22.png",
     "asset_id": "ed35ed23-b843-4470-bb40-db03b4117fa0",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_outdoor_reception_sel.png",
     "asset_id": "eac29610-01b6-43a7-9e6e-9e2387ddd674",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "다섯 명의 인물 모두 프레임 우측의 모퉁이 너머 바깥쪽을 향해 시선을 두고 있음.",
    "built_space": "중앙에 거친 질감의 콘크리트 벽이 위치하며 인물들을 가려주고 있고, 우측 상단으로 바깥 거리가 드러나는 모퉁이가 있음. 바닥은 거친 노면 형태를 띰.",
    "entities": "현우(얼굴 상처, 회색 셔츠), 앰버(방독면, 작업복, 파우치), 찰리(로봇 외형, 코트, 모자), 라울(꽁지머리, 반바지), 페드로(비니, 검은 후드) 5명 모두 참조 이미지의 외형 및 복장과 일치함.",
    "hard_violations": [],
    "physics": "모든 인물이 바닥에 두 발을 딛고 서 있거나 무릎을 굽힌 채 안정적인 자세를 유지하고 있음."
   },
   {
    "label": "B",
    "direction": "인물들이 화면 좌측 전방을 향해 시선을 던지고 있음.",
    "built_space": "화면 좌측에 붉은 벽돌 벽이 길게 늘어서 있고, 우측 안쪽으로 좁은 골목길이 이어짐.",
    "entities": "현우, 앰버, 페드로, 라울, 찰리 등 5명이 등장하며 외형적 특징은 유지됨.",
    "hard_violations": [
     "[gemini-pro] 화면 좌측 하단에 프롬프트 문장(좁고 어두운...)이 텍스트 자막으로 유출됨 (leaked text)",
     "[gpt-high] 왼쪽 아래에 촬영 지시문을 옮긴 한국어 자막이 삽입되어, 읽을 수 있는 글자와 오버레이를 금지한 조건을 직접 위반한다."
    ],
    "physics": "인물들 모두 흙바닥 위에 서서 체중을 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "화면에 어떠한 텍스트 노출도 없으며, 우측 상단의 모퉁이와 프레임 중앙의 벽면이라는 요구 구도를 정확히 구현하여 제시된 상황을 가장 충실히 묘사했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면 좌측 하단에 프롬프트의 지시문이 자막 형태로 그대로 유출되는 치명적인 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "다섯 명의 인물 모두 프레임 우측의 모퉁이 너머 바깥쪽을 향해 시선을 두고 있음.",
        "built_space": "중앙에 거친 질감의 콘크리트 벽이 위치하며 인물들을 가려주고 있고, 우측 상단으로 바깥 거리가 드러나는 모퉁이가 있음. 바닥은 거친 노면 형태를 띰.",
        "entities": "현우(얼굴 상처, 회색 셔츠), 앰버(방독면, 작업복, 파우치), 찰리(로봇 외형, 코트, 모자), 라울(꽁지머리, 반바지), 페드로(비니, 검은 후드) 5명 모두 참조 이미지의 외형 및 복장과 일치함.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 두 발을 딛고 서 있거나 무릎을 굽힌 채 안정적인 자세를 유지하고 있음."
       },
       {
        "label": "B",
        "direction": "인물들이 화면 좌측 전방을 향해 시선을 던지고 있음.",
        "built_space": "화면 좌측에 붉은 벽돌 벽이 길게 늘어서 있고, 우측 안쪽으로 좁은 골목길이 이어짐.",
        "entities": "현우, 앰버, 페드로, 라울, 찰리 등 5명이 등장하며 외형적 특징은 유지됨.",
        "hard_violations": [
         "화면 좌측 하단에 프롬프트 문장(좁고 어두운...)이 텍스트 자막으로 유출됨 (leaked text)"
        ],
        "physics": "인물들 모두 흙바닥 위에 서서 체중을 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "화면에 어떠한 텍스트 노출도 없으며, 우측 상단의 모퉁이와 프레임 중앙의 벽면이라는 요구 구도를 정확히 구현하여 제시된 상황을 가장 충실히 묘사했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면 좌측 하단에 프롬프트의 지시문이 자막 형태로 그대로 유출되는 치명적인 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "다섯 명의 인물 모두 프레임 우측의 모퉁이 너머 바깥쪽을 향해 시선을 두고 있음.",
        "built_space": "중앙에 거친 질감의 콘크리트 벽이 위치하며 인물들을 가려주고 있고, 우측 상단으로 바깥 거리가 드러나는 모퉁이가 있음. 바닥은 거친 노면 형태를 띰.",
        "entities": "현우(얼굴 상처, 회색 셔츠), 앰버(방독면, 작업복, 파우치), 찰리(로봇 외형, 코트, 모자), 라울(꽁지머리, 반바지), 페드로(비니, 검은 후드) 5명 모두 참조 이미지의 외형 및 복장과 일치함.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 두 발을 딛고 서 있거나 무릎을 굽힌 채 안정적인 자세를 유지하고 있음."
       },
       {
        "label": "B",
        "direction": "인물들이 화면 좌측 전방을 향해 시선을 던지고 있음.",
        "built_space": "화면 좌측에 붉은 벽돌 벽이 길게 늘어서 있고, 우측 안쪽으로 좁은 골목길이 이어짐.",
        "entities": "현우, 앰버, 페드로, 라울, 찰리 등 5명이 등장하며 외형적 특징은 유지됨.",
        "hard_violations": [
         "화면 좌측 하단에 프롬프트 문장(좁고 어두운...)이 텍스트 자막으로 유출됨 (leaked text)"
        ],
        "physics": "인물들 모두 흙바닥 위에 서서 체중을 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "다섯 인물의 전신과 벽 뒤 은폐, 오른쪽 모퉁이 밖을 살피는 행동을 정확히 구현했으나 앰버의 나이감·상의와 장소의 세부 일치는 아쉽다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 한국어 자막이 들어가 실격이며, 일행의 시선도 지정된 오른쪽 위 개구부가 아닌 왼쪽 화면 밖을 향한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸을 숙여 왼쪽 화면 밖을 살피고, 앰버·페드로·라울도 대체로 왼쪽을 본다. 찰리의 얼굴 역시 왼쪽 전방으로 돌아가 있다. 실제로 보이는 바깥 통로는 일행 뒤 오른쪽 위에 있어, 시선과 지정된 위험 방향이 연결되지 않는다. 무기나 겨누는 소품은 없다.",
        "built_space": "왼쪽의 긴 벽돌벽과 오른쪽 건물 벽 사이에 좁은 통로 하나가 있고, 뒤쪽 오른쪽으로 밝은 외부 공간 하나가 열린다. 양쪽 벽 끝에 수직 배수관이 보이며 바닥은 흙과 파편으로 덮여 있다. 다섯 인물은 왼쪽 벽 가까이 모여 있지만, 오른쪽 뒤 개구부를 차단하는 공통 은폐 모퉁이가 분명하지 않다. 참고의 벽돌 주택 재료는 일부 이어지지만 정확히 같은 고정 구조라고 확인할 세부는 부족하다. 페드로는 다른 인물 뒤에 가려 전신 판독이 어렵다.",
        "entities": "인물은 정확히 다섯이다. 현우는 앳된 동아시아계 남성, 헝클어진 검은 머리, 회색 겉셔츠와 얼굴의 상처를 유지한다. 앰버는 금발 여자아이이며 호흡 마스크, 남색 상의, 갈색 작업복과 허리 도구를 갖춰 참고에 가깝다. 라울은 짙은 피부의 어린 남자아이로 회색 티셔츠와 반바지를 입었다. 페드로는 짙은 머리의 젊은 남성으로 보이나 참고의 검은 비니가 없다. 찰리는 베이지 장갑판, 흰 기계 얼굴, 낡은 코트와 모자를 유지한다. 왼쪽 아래에는 읽을 수 있는 한국어 문장 두 줄이 크게 삽입되어 있다.",
        "hard_violations": [
         "왼쪽 아래에 촬영 지시문을 옮긴 한국어 자막이 삽입되어, 읽을 수 있는 글자와 오버레이를 금지한 조건을 직접 위반한다."
        ],
        "physics": "현우는 벌린 두 발로 바닥을 딛고 상체를 앞으로 기울이며, 앰버와 라울 및 찰리도 발로 지면을 지지한다. 페드로의 하체는 앞사람에게 가려져 지지점을 명확히 확인할 수 없지만 공중에 떠 있다는 증거는 없다. 마스크는 머리끈, 도구 주머니는 허리띠에 고정되어 있다. 명백한 부유나 불가능한 관절 자세는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 모퉁이 쪽으로 목과 상체를 내밀어 바깥 통로를 살핀다. 앰버도 같은 방향을 보고, 찰리·라울·페드로의 얼굴도 오른쪽 전방을 향한다. 일행의 주의 방향이 오른쪽 위에 열린 통로와 일치한다. 직접 보이는 감시 대상 인물이나 겨누는 무기는 없다.",
        "built_space": "중앙의 긴 회벽 하나가 오른쪽 모퉁이에서 꺾이며 다섯 인물의 공통 은폐 경계를 만든다. 카메라는 그 안쪽 면을 따라 비스듬히 보고, 모퉁이 너머 오른쪽에 좁은 통로 하나가 드러난다. 전경부터 각 인물의 발 아래까지 바닥이 이어져 전신이 읽힌다. 왼쪽 건물에는 창문·수직 배관·작은 전기함이 있고, 오른쪽 통로에는 켜진 가로등 두 개와 먼 차량 한 대가 보인다. 벽돌, 낡은 미장벽, 배선과 가로등은 참고의 재료와 분위기를 잇지만, 참고 사진과 정확히 같은 벽·개구부 배치인지는 확인하기 어렵다.",
        "entities": "현우·앰버·찰리·라울·페드로 다섯만 등장한다. 현우는 검은 머리의 앳된 동아시아계 남성이며 회색 겉셔츠와 뚜렷한 얼굴 멍을 유지하지만 다리 부상은 명확하지 않다. 앰버는 금발, 호흡 마스크와 허리 주머니를 유지하나 참고의 열 살 아이보다 크게 보이고 남색 반소매 대신 회색 긴소매를 입었다. 라울은 짙은 피부, 뒤로 묶은 머리, 회색 티셔츠와 반바지가 맞는다. 페드로는 검은 비니·후드·바지와 허리 체인을 유지한다. 찰리는 육중한 장갑 몸체, 긴 팔, 흰 마스크형 얼굴, 코트와 모자를 유지한다. 배경 차량의 민병대 소속 여부는 식별할 수 없으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 벌린 두 발로 바닥을 딛고 모퉁이를 향해 체중을 이동한다. 앰버는 무릎을 굽힌 채 부츠로 지면을 지지하고, 라울과 페드로는 발을 벌리고 손을 허벅지에 짚어 달린 뒤 숨을 고르는 자세를 취한다. 찰리의 넓은 기계 발은 바닥에 놓여 있고 다른 쪽 하체는 일부 가려진다. 마스크 끈과 허리띠가 소품을 지탱하며, 배경 차량도 바퀴로 노면에 서 있다. 지지 없는 부유나 명백히 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "다섯 인물의 전신과 벽 뒤 은폐, 오른쪽 모퉁이 밖을 살피는 행동을 정확히 구현했으나 앰버의 나이감·상의와 장소의 세부 일치는 아쉽다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 한국어 자막이 들어가 실격이며, 일행의 시선도 지정된 오른쪽 위 개구부가 아닌 왼쪽 화면 밖을 향한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸을 숙여 왼쪽 화면 밖을 살피고, 앰버·페드로·라울도 대체로 왼쪽을 본다. 찰리의 얼굴 역시 왼쪽 전방으로 돌아가 있다. 실제로 보이는 바깥 통로는 일행 뒤 오른쪽 위에 있어, 시선과 지정된 위험 방향이 연결되지 않는다. 무기나 겨누는 소품은 없다.",
        "built_space": "왼쪽의 긴 벽돌벽과 오른쪽 건물 벽 사이에 좁은 통로 하나가 있고, 뒤쪽 오른쪽으로 밝은 외부 공간 하나가 열린다. 양쪽 벽 끝에 수직 배수관이 보이며 바닥은 흙과 파편으로 덮여 있다. 다섯 인물은 왼쪽 벽 가까이 모여 있지만, 오른쪽 뒤 개구부를 차단하는 공통 은폐 모퉁이가 분명하지 않다. 참고의 벽돌 주택 재료는 일부 이어지지만 정확히 같은 고정 구조라고 확인할 세부는 부족하다. 페드로는 다른 인물 뒤에 가려 전신 판독이 어렵다.",
        "entities": "인물은 정확히 다섯이다. 현우는 앳된 동아시아계 남성, 헝클어진 검은 머리, 회색 겉셔츠와 얼굴의 상처를 유지한다. 앰버는 금발 여자아이이며 호흡 마스크, 남색 상의, 갈색 작업복과 허리 도구를 갖춰 참고에 가깝다. 라울은 짙은 피부의 어린 남자아이로 회색 티셔츠와 반바지를 입었다. 페드로는 짙은 머리의 젊은 남성으로 보이나 참고의 검은 비니가 없다. 찰리는 베이지 장갑판, 흰 기계 얼굴, 낡은 코트와 모자를 유지한다. 왼쪽 아래에는 읽을 수 있는 한국어 문장 두 줄이 크게 삽입되어 있다.",
        "hard_violations": [
         "왼쪽 아래에 촬영 지시문을 옮긴 한국어 자막이 삽입되어, 읽을 수 있는 글자와 오버레이를 금지한 조건을 직접 위반한다."
        ],
        "physics": "현우는 벌린 두 발로 바닥을 딛고 상체를 앞으로 기울이며, 앰버와 라울 및 찰리도 발로 지면을 지지한다. 페드로의 하체는 앞사람에게 가려져 지지점을 명확히 확인할 수 없지만 공중에 떠 있다는 증거는 없다. 마스크는 머리끈, 도구 주머니는 허리띠에 고정되어 있다. 명백한 부유나 불가능한 관절 자세는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 모퉁이 쪽으로 목과 상체를 내밀어 바깥 통로를 살핀다. 앰버도 같은 방향을 보고, 찰리·라울·페드로의 얼굴도 오른쪽 전방을 향한다. 일행의 주의 방향이 오른쪽 위에 열린 통로와 일치한다. 직접 보이는 감시 대상 인물이나 겨누는 무기는 없다.",
        "built_space": "중앙의 긴 회벽 하나가 오른쪽 모퉁이에서 꺾이며 다섯 인물의 공통 은폐 경계를 만든다. 카메라는 그 안쪽 면을 따라 비스듬히 보고, 모퉁이 너머 오른쪽에 좁은 통로 하나가 드러난다. 전경부터 각 인물의 발 아래까지 바닥이 이어져 전신이 읽힌다. 왼쪽 건물에는 창문·수직 배관·작은 전기함이 있고, 오른쪽 통로에는 켜진 가로등 두 개와 먼 차량 한 대가 보인다. 벽돌, 낡은 미장벽, 배선과 가로등은 참고의 재료와 분위기를 잇지만, 참고 사진과 정확히 같은 벽·개구부 배치인지는 확인하기 어렵다.",
        "entities": "현우·앰버·찰리·라울·페드로 다섯만 등장한다. 현우는 검은 머리의 앳된 동아시아계 남성이며 회색 겉셔츠와 뚜렷한 얼굴 멍을 유지하지만 다리 부상은 명확하지 않다. 앰버는 금발, 호흡 마스크와 허리 주머니를 유지하나 참고의 열 살 아이보다 크게 보이고 남색 반소매 대신 회색 긴소매를 입었다. 라울은 짙은 피부, 뒤로 묶은 머리, 회색 티셔츠와 반바지가 맞는다. 페드로는 검은 비니·후드·바지와 허리 체인을 유지한다. 찰리는 육중한 장갑 몸체, 긴 팔, 흰 마스크형 얼굴, 코트와 모자를 유지한다. 배경 차량의 민병대 소속 여부는 식별할 수 없으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 벌린 두 발로 바닥을 딛고 모퉁이를 향해 체중을 이동한다. 앰버는 무릎을 굽힌 채 부츠로 지면을 지지하고, 라울과 페드로는 발을 벌리고 손을 허벅지에 짚어 달린 뒤 숨을 고르는 자세를 취한다. 찰리의 넓은 기계 발은 바닥에 놓여 있고 다른 쪽 하체는 일부 가려진다. 마스크 끈과 허리띠가 소품을 지탱하며, 배경 차량도 바퀴로 노면에 서 있다. 지지 없는 부유나 명백히 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 좌측 하단에 프롬프트 문장(좁고 어두운...)이 텍스트 자막으로 유출됨 (leaked text)",
     "[gpt-high] 왼쪽 아래에 촬영 지시문을 옮긴 한국어 자막이 삽입되어, 읽을 수 있는 글자와 오버레이를 금지한 조건을 직접 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "화면에 어떠한 텍스트 노출도 없으며, 우측 상단의 모퉁이와 프레임 중앙의 벽면이라는 요구 구도를 정확히 구현하여 제시된 상황을 가장 충실히 묘사했습니다."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "화면 좌측 하단에 프롬프트의 지시문이 자막 형태로 그대로 유출되는 치명적인 위반이 발생했습니다.  ★위반: [gemini-pro] 화면 좌측 하단에 프롬프트 문장(좁고 어두운...)이 텍스트 자막으로 유출됨 (leaked text) / [gpt-high] 왼쪽 아래에 촬영 지시문을 옮긴 한국어 자막이 삽입되어, 읽을 수 있는 글자와 오버레이를 금지한 조건을 직접 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_outdoor_reception_sel.png",
    "asset_id": "eac29610-01b6-43a7-9e6e-9e2387ddd674",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-11b4-7dad-8439-03d3167d319d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh22__bgfirst_bg.png",
   "bg_asset_id": "4ea5365d-bc0a-4f55-9bd5-c4357af778b0",
   "bg_record_key": "S25sh22::bgfirst_bg",
   "chain_winner": true,
   "authority": "seed_bg"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S25sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:04:07.178706+00:00",
  "fingerprint": "305eef0d8ffa00d1348d2f7bd1c4c9472f60ebae48a13289e6f2b0f0aa6cc35e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S25sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S25sh22_sel.png",
  "source_sha256": "03bc29d934135eb79a43a4d8821dac3cf6088e6476a092a554403bb6cdd3057a",
  "file": "S25sh22_cine.png",
  "staged_sha256": "5f0ff10ac0f0bee1dc64302c4866e0fe48e7b0df65c1bbaf9c99b876bd6cb59b",
  "latency_ms": 9101
 },
 "S26sh7::signage": {
  "fp": "9f818f7c8ea421b3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S26sh7": {
  "input_fingerprint": "07562e85e19f2fcb",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 무리 속에 섞인 미연을 향해 서늘한 눈빛을 번뜩이는 박철진의 비열한 상체.\n\nLOCATION (lock): In the open reception lot at night, beside the gathered guests being held under militia guard. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 피로연장 중앙 공터 (The guests have been gathered together under threat); used as Leaves a readable depth interval between the interrogation and 미연 at the front of the gathered guests.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the reception's established nighttime illumination with controlled contrast that keeps both 박철진's glance and 미연's alarm readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception remains brightly lit, with tents, the jukebox and burst lamps still in place. The rubbish-covered sewer entrance has had its manhole cover pushed aside, and Charlie's coat-and-hat disguise remains unchanged below ground. 박철진: He remains in the reception clearing conducting the interrogation. 미연: She stands in the front row of the gathered reception guests.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 무리 속에 섞인 미연을 향해 서늘한 눈빛을 번뜩이는 박철진의 비열한 상체.\n\nLOCATION (lock): In the open reception lot at night, beside the gathered guests being held under militia guard. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 피로연장 중앙 공터 (The guests have been gathered together under threat); used as Leaves a readable depth interval between the interrogation and 미연 at the front of the gathered guests.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the reception's established nighttime illumination with controlled contrast that keeps both 박철진's glance and 미연's alarm readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception remains brightly lit, with tents, the jukebox and burst lamps still in place. The rubbish-covered sewer entrance has had its manhole cover pushed aside, and Charlie's coat-and-hat disguise remains unchanged below ground. 박철진: He remains in the reception clearing conducting the interrogation. 미연: She stands in the front row of the gathered reception guests.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 무리 속에 섞인 미연을 향해 서늘한 눈빛을 번뜩이는 박철진의 비열한 상체.\n\nLOCATION (lock): In the open reception lot at night, beside the gathered guests being held under militia guard. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 피로연장 중앙 공터 (The guests have been gathered together under threat); used as Leaves a readable depth interval between the interrogation and 미연 at the front of the gathered guests.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the reception's established nighttime illumination with controlled contrast that keeps both 박철진's glance and 미연's alarm readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception remains brightly lit, with tents, the jukebox and burst lamps still in place. The rubbish-covered sewer entrance has had its manhole cover pushed aside, and Charlie's coat-and-hat disguise remains unchanged below ground. 박철진: He remains in the reception clearing conducting the interrogation. 미연: She stands in the front row of the gathered reception guests.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진이 미연을 보지 않고 정면의 카메라를 응시함.",
    "built_space": "야외 피로연장의 주크박스와 텐트, 트럭이 레퍼런스의 구조대로 배치됨.",
    "entities": "박철진의 눈동자가 하얗게 변형되었고, 미연은 무리가 아닌 박철진 뒤에 단독으로 서 있음.",
    "hard_violations": [
     "[gemini-pro] 눈동자를 하얗게 변형함 (해부학적 변형 금지 위반)",
     "[gemini-pro] 이전 컷 레퍼런스의 민병대원을 똑같이 복사하여 배치함 (이전 컷 인물 재사용 금지 위반)",
     "[gpt-high] 이전 사진에서 박철진 이외의 인물을 이어 등장시키지 말라는 지시를 어기고, 여성 경비와 남성 경비 등 여섯 명을 같은 얼굴·복장으로 재등장시켰다."
    ],
    "physics": "인물들이 지면에 안정적으로 서 있음."
   },
   {
    "label": "B",
    "direction": "박철진이 고개를 돌려 무리 앞줄에 있는 미연을 정확히 주시함.",
    "built_space": "피로연장 중앙 공터에 하객 무리와 민병대가 나뉘어 서 있고 알맞은 깊이감이 확보됨.",
    "entities": "박철진과 미연의 인상착의가 레퍼런스와 일치하며 무리 속에 섞인 모습이 잘 연출됨.",
    "hard_violations": [
     "[gemini-pro] 이전 컷 레퍼런스 우측의 민병대원 4명을 동일하게 복사함 (이전 컷 인물 재사용 금지 위반)",
     "[gpt-high] 박철진 외에는 이전 사진의 얼굴·몸·복장을 재사용하지 말라는 명시적 금지에도 불구하고, 이전 사진의 무장 경비 여러 명을 동일한 모습으로 재등장시켰다."
    ],
    "physics": "인물들이 지면에 정상적으로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "안구가 하얗게 빛나는 치명적 오류가 있고 시선이 카메라를 향하며, 이전 컷의 배경 인물까지 그대로 복사했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "미연을 향한 시선과 군중의 배치는 훌륭하나, 이전 컷의 민병대를 똑같이 복사하여 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 미연을 보지 않고 정면의 카메라를 응시함.",
        "built_space": "야외 피로연장의 주크박스와 텐트, 트럭이 레퍼런스의 구조대로 배치됨.",
        "entities": "박철진의 눈동자가 하얗게 변형되었고, 미연은 무리가 아닌 박철진 뒤에 단독으로 서 있음.",
        "hard_violations": [
         "눈동자를 하얗게 변형함 (해부학적 변형 금지 위반)",
         "이전 컷 레퍼런스의 민병대원을 똑같이 복사하여 배치함 (이전 컷 인물 재사용 금지 위반)"
        ],
        "physics": "인물들이 지면에 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 고개를 돌려 무리 앞줄에 있는 미연을 정확히 주시함.",
        "built_space": "피로연장 중앙 공터에 하객 무리와 민병대가 나뉘어 서 있고 알맞은 깊이감이 확보됨.",
        "entities": "박철진과 미연의 인상착의가 레퍼런스와 일치하며 무리 속에 섞인 모습이 잘 연출됨.",
        "hard_violations": [
         "이전 컷 레퍼런스 우측의 민병대원 4명을 동일하게 복사함 (이전 컷 인물 재사용 금지 위반)"
        ],
        "physics": "인물들이 지면에 정상적으로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "안구가 하얗게 빛나는 치명적 오류가 있고 시선이 카메라를 향하며, 이전 컷의 배경 인물까지 그대로 복사했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "미연을 향한 시선과 군중의 배치는 훌륭하나, 이전 컷의 민병대를 똑같이 복사하여 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 미연을 보지 않고 정면의 카메라를 응시함.",
        "built_space": "야외 피로연장의 주크박스와 텐트, 트럭이 레퍼런스의 구조대로 배치됨.",
        "entities": "박철진의 눈동자가 하얗게 변형되었고, 미연은 무리가 아닌 박철진 뒤에 단독으로 서 있음.",
        "hard_violations": [
         "눈동자를 하얗게 변형함 (해부학적 변형 금지 위반)",
         "이전 컷 레퍼런스의 민병대원을 똑같이 복사하여 배치함 (이전 컷 인물 재사용 금지 위반)"
        ],
        "physics": "인물들이 지면에 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 고개를 돌려 무리 앞줄에 있는 미연을 정확히 주시함.",
        "built_space": "피로연장 중앙 공터에 하객 무리와 민병대가 나뉘어 서 있고 알맞은 깊이감이 확보됨.",
        "entities": "박철진과 미연의 인상착의가 레퍼런스와 일치하며 무리 속에 섞인 모습이 잘 연출됨.",
        "hard_violations": [
         "이전 컷 레퍼런스 우측의 민병대원 4명을 동일하게 복사함 (이전 컷 인물 재사용 금지 위반)"
        ],
        "physics": "인물들이 지면에 정상적으로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미연을 하객 무리 앞줄에 두고 공간적 간격을 만들었지만, 박철진의 눈길은 미연이 아닌 카메라로 향하며 금지된 이전 장면의 경비 인물들도 재등장한다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "미연과 하객 무리가 빠진 채 박철진이 카메라를 노려보며, 제외하도록 명시된 이전 장면의 경비 인물들을 그대로 반복한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 몸을 화면 오른쪽으로 돌리고 얼굴과 눈을 카메라 쪽으로 향한다. 미연은 그의 왼쪽 뒤 하객 앞줄에 있어 이 눈길의 도착점이 아니다. 미연과 하객들은 대체로 박철진이 있는 전경 쪽을 본다. 경비들의 소총 총구는 주로 바닥 쪽으로 내려가 있으며 미연을 직접 겨누지는 않는다.",
        "built_space": "왼쪽의 두 봉우리로 이어진 천막, 오른쪽 전경의 천막 일부, 오른쪽 뒤 군용 트럭 한 대, 왼쪽 전경 주크박스 한 대가 보인다. 원형 식탁 여러 개와 접이식 의자, 조명 기둥 및 전구 줄도 유지된다. 중앙 흙바닥과 물웅덩이는 참고 장소와 일치한다. 박철진과 뒤쪽 하객 사이에는 빈 공간이 있고 미연은 무리 앞줄에 서 있다. 다만 박철진을 허벅지까지 담고 주변을 넓게 보여 주어 요청한 상체 중심 미디엄 숏보다 넓다.",
        "entities": "박철진은 짧은 검은 머리의 중년 동아시아계 남성으로, 참고의 얼굴과 어두운 군복·장비 벨트·붉은 완장을 대체로 따른다. 완장에는 참고에 없던 검은 도형이 붙어 있다. 미연은 검은 단발과 남색 옷을 입은 중년 동아시아계 여성으로 참고와 대체로 부합한다. 여러 하객 외에 이전 사진의 여성 경비, 목도리 남성 경비 등 동일한 얼굴과 복장의 무장 인물들이 다시 보인다. 밤과 따뜻한 피로연 조명은 구현되었고 읽을 수 있는 글자는 없다. 지하의 찰리와 하수구는 이 구도에서 확인되지 않는다.",
        "hard_violations": [
         "박철진 외에는 이전 사진의 얼굴·몸·복장을 재사용하지 말라는 명시적 금지에도 불구하고, 이전 사진의 무장 경비 여러 명을 동일한 모습으로 재등장시켰다."
        ],
        "physics": "뒤쪽 인물들은 발을 흙바닥에 딛고 서 있으며, 소총은 손과 멜빵으로 지지된다. 박철진의 발은 화면 밖이지만 상체와 골반은 자연스럽게 이어지고 서 있는 자세로 읽힌다. 천막은 기둥과 줄로, 주크박스와 식탁은 지면과 다리로 지지된다. 부유하거나 지지 없이 놓인 신체·물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "박철진은 상체를 앞으로 숙이고 카메라를 정면으로 노려본다. 화면 안에는 그 시선의 목표인 미연이 없다. 뒤쪽 경비들도 대체로 카메라 쪽을 보고 소총 총구를 비스듬히 아래로 향한다. 미연을 향한 눈길이라는 핵심 관계가 성립하지 않는다.",
        "built_space": "왼쪽 두 봉우리의 천막, 오른쪽 전경 천막 일부, 군용 트럭 한 대, 주크박스 한 대가 이전 사진과 거의 같은 위치에 있다. 식탁·접이식 의자·통·조명 기둥과 전구 줄, 젖은 흙바닥도 이어진다. 그러나 중앙 공터에는 하객 무리 대신 경비 여섯 명이 서 있어 미연의 앞줄 위치와 심문 장소 사이의 깊이 관계가 없다. 박철진은 무릎 부근까지 보이므로 상체 중심 미디엄 숏보다 넓은 구도다.",
        "entities": "박철진의 검은 머리, 중년 남성 얼굴, 어두운 군복, 장비 벨트와 붉은 완장은 참고를 대체로 따른다. 완장의 검은 도형은 참고와 다르다. 눈은 아래로 숙인 얼굴에서 치켜뜬 연기로 보이나 흰자 대비가 과하게 강조되어 있다. 참고의 검은 단발·남색 옷을 입은 미연은 없고, 뒤쪽 여성은 이전 사진의 군복 차림 경비다. 그 여성을 포함한 경비 여섯 명은 이전 사진의 인물과 복장을 반복한다. 밤 조명은 맞고 판독 가능한 글자는 없다. 지하 공간과 찰리는 보이지 않는다.",
        "hard_violations": [
         "이전 사진에서 박철진 이외의 인물을 이어 등장시키지 말라는 지시를 어기고, 여성 경비와 남성 경비 등 여섯 명을 같은 얼굴·복장으로 재등장시켰다."
        ],
        "physics": "박철진은 무릎과 골반을 굽히고 상체를 앞으로 기울인다. 발은 잘렸지만 다리로 체중을 받는 자세로 성립하며 공중에 떠 있는 모습은 아니다. 경비들의 발은 지면에 닿고 총은 손과 멜빵으로 지지된다. 트럭은 바퀴로, 가구는 다리와 바닥으로, 천막은 기둥으로 지지되어 물리적으로 불가능한 배치는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "미연을 하객 무리 앞줄에 두고 공간적 간격을 만들었지만, 박철진의 눈길은 미연이 아닌 카메라로 향하며 금지된 이전 장면의 경비 인물들도 재등장한다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "미연과 하객 무리가 빠진 채 박철진이 카메라를 노려보며, 제외하도록 명시된 이전 장면의 경비 인물들을 그대로 반복한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 몸을 화면 오른쪽으로 돌리고 얼굴과 눈을 카메라 쪽으로 향한다. 미연은 그의 왼쪽 뒤 하객 앞줄에 있어 이 눈길의 도착점이 아니다. 미연과 하객들은 대체로 박철진이 있는 전경 쪽을 본다. 경비들의 소총 총구는 주로 바닥 쪽으로 내려가 있으며 미연을 직접 겨누지는 않는다.",
        "built_space": "왼쪽의 두 봉우리로 이어진 천막, 오른쪽 전경의 천막 일부, 오른쪽 뒤 군용 트럭 한 대, 왼쪽 전경 주크박스 한 대가 보인다. 원형 식탁 여러 개와 접이식 의자, 조명 기둥 및 전구 줄도 유지된다. 중앙 흙바닥과 물웅덩이는 참고 장소와 일치한다. 박철진과 뒤쪽 하객 사이에는 빈 공간이 있고 미연은 무리 앞줄에 서 있다. 다만 박철진을 허벅지까지 담고 주변을 넓게 보여 주어 요청한 상체 중심 미디엄 숏보다 넓다.",
        "entities": "박철진은 짧은 검은 머리의 중년 동아시아계 남성으로, 참고의 얼굴과 어두운 군복·장비 벨트·붉은 완장을 대체로 따른다. 완장에는 참고에 없던 검은 도형이 붙어 있다. 미연은 검은 단발과 남색 옷을 입은 중년 동아시아계 여성으로 참고와 대체로 부합한다. 여러 하객 외에 이전 사진의 여성 경비, 목도리 남성 경비 등 동일한 얼굴과 복장의 무장 인물들이 다시 보인다. 밤과 따뜻한 피로연 조명은 구현되었고 읽을 수 있는 글자는 없다. 지하의 찰리와 하수구는 이 구도에서 확인되지 않는다.",
        "hard_violations": [
         "박철진 외에는 이전 사진의 얼굴·몸·복장을 재사용하지 말라는 명시적 금지에도 불구하고, 이전 사진의 무장 경비 여러 명을 동일한 모습으로 재등장시켰다."
        ],
        "physics": "뒤쪽 인물들은 발을 흙바닥에 딛고 서 있으며, 소총은 손과 멜빵으로 지지된다. 박철진의 발은 화면 밖이지만 상체와 골반은 자연스럽게 이어지고 서 있는 자세로 읽힌다. 천막은 기둥과 줄로, 주크박스와 식탁은 지면과 다리로 지지된다. 부유하거나 지지 없이 놓인 신체·물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "박철진은 상체를 앞으로 숙이고 카메라를 정면으로 노려본다. 화면 안에는 그 시선의 목표인 미연이 없다. 뒤쪽 경비들도 대체로 카메라 쪽을 보고 소총 총구를 비스듬히 아래로 향한다. 미연을 향한 눈길이라는 핵심 관계가 성립하지 않는다.",
        "built_space": "왼쪽 두 봉우리의 천막, 오른쪽 전경 천막 일부, 군용 트럭 한 대, 주크박스 한 대가 이전 사진과 거의 같은 위치에 있다. 식탁·접이식 의자·통·조명 기둥과 전구 줄, 젖은 흙바닥도 이어진다. 그러나 중앙 공터에는 하객 무리 대신 경비 여섯 명이 서 있어 미연의 앞줄 위치와 심문 장소 사이의 깊이 관계가 없다. 박철진은 무릎 부근까지 보이므로 상체 중심 미디엄 숏보다 넓은 구도다.",
        "entities": "박철진의 검은 머리, 중년 남성 얼굴, 어두운 군복, 장비 벨트와 붉은 완장은 참고를 대체로 따른다. 완장의 검은 도형은 참고와 다르다. 눈은 아래로 숙인 얼굴에서 치켜뜬 연기로 보이나 흰자 대비가 과하게 강조되어 있다. 참고의 검은 단발·남색 옷을 입은 미연은 없고, 뒤쪽 여성은 이전 사진의 군복 차림 경비다. 그 여성을 포함한 경비 여섯 명은 이전 사진의 인물과 복장을 반복한다. 밤 조명은 맞고 판독 가능한 글자는 없다. 지하 공간과 찰리는 보이지 않는다.",
        "hard_violations": [
         "이전 사진에서 박철진 이외의 인물을 이어 등장시키지 말라는 지시를 어기고, 여성 경비와 남성 경비 등 여섯 명을 같은 얼굴·복장으로 재등장시켰다."
        ],
        "physics": "박철진은 무릎과 골반을 굽히고 상체를 앞으로 기울인다. 발은 잘렸지만 다리로 체중을 받는 자세로 성립하며 공중에 떠 있는 모습은 아니다. 경비들의 발은 지면에 닿고 총은 손과 멜빵으로 지지된다. 트럭은 바퀴로, 가구는 다리와 바닥으로, 천막은 기둥으로 지지되어 물리적으로 불가능한 배치는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.833,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.583,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 눈동자를 하얗게 변형함 (해부학적 변형 금지 위반)",
     "[gemini-pro] 이전 컷 레퍼런스의 민병대원을 똑같이 복사하여 배치함 (이전 컷 인물 재사용 금지 위반)",
     "[gpt-high] 이전 사진에서 박철진 이외의 인물을 이어 등장시키지 말라는 지시를 어기고, 여성 경비와 남성 경비 등 여섯 명을 같은 얼굴·복장으로 재등장시켰다."
    ],
    "B": [
     "[gemini-pro] 이전 컷 레퍼런스 우측의 민병대원 4명을 동일하게 복사함 (이전 컷 인물 재사용 금지 위반)",
     "[gpt-high] 박철진 외에는 이전 사진의 얼굴·몸·복장을 재사용하지 말라는 명시적 금지에도 불구하고, 이전 사진의 무장 경비 여러 명을 동일한 모습으로 재등장시켰다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 583,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 583,
    "verdict_ko": "안구가 하얗게 빛나는 치명적 오류가 있고 시선이 카메라를 향하며, 이전 컷의 배경 인물까지 그대로 복사했습니다.  ★위반: [gemini-pro] 눈동자를 하얗게 변형함 (해부학적 변형 금지 위반) / [gemini-pro] 이전 컷 레퍼런스의 민병대원을 똑같이 복사하여 배치함 (이전 컷 인물 재사용 금지 위반) / [gpt-high] 이전 사진에서 박철진 이외의 인물을 이어 등장시키지 말라는 지시를 어기고, 여성 경비와 남성 경비 등 여섯 명을 같은 얼굴·복장으로 재등장시켰다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "미연을 향한 시선과 군중의 배치는 훌륭하나, 이전 컷의 민병대를 똑같이 복사하여 지시를 위반했습니다.  ★위반: [gemini-pro] 이전 컷 레퍼런스 우측의 민병대원 4명을 동일하게 복사함 (이전 컷 인물 재사용 금지 위반) / [gpt-high] 박철진 외에는 이전 사진의 얼굴·몸·복장을 재사용하지 말라는 명시적 금지에도 불구하고, 이전 사진의 무장 경비 여러 명을 동일한 모습으로 재등장시켰다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh19_sel.png",
    "asset_id": "3f2f42c3-f8c3-459b-b94c-62d98157fff9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-1524-728d-98c2-e1afb1246812",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S25sh19"
  },
  "staged_characters_added": [
   "C10"
  ]
 },
 "S26sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:05:21.081308+00:00",
  "fingerprint": "f0c43ebac1527c61508170c83e2b2651b692b441d67b50ce7bc9ff75d310164a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S26sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S26sh7_sel.png",
  "source_sha256": "7e1aca4dca04fef52e511e9d23e433498ee37c82a1c62d8315bc44c59def3f8b",
  "file": "S26sh7_cine.png",
  "staged_sha256": "f990c7ab783839d5d863ab01901bc051d8a8b4ffe1809ede59f6718b984e51fb",
  "latency_ms": 9840
 },
 "S26sh9::signage": {
  "fp": "9a54d860a83e39c2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::bbb84cc871e390a4": {
  "subjects": [],
  "subject_text": "인천 난민촌 지하 하수도\n맨홀 아래로 사다리가 이어지는 어두운 지하 배수 통로. 길게 뻗은 통로 중간에 두 갈래로 나뉘는 분기점이 있다.",
  "identity": "canonical",
  "scope_id": "L175",
  "scope_role": "location_interior",
  "scope_sha": "86e38b0060fbf724"
 },
 "S26sh9::bgfirst_bg": {
  "input_fingerprint": "54c6631247e2cdb8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9__bgfirst_bg.png",
  "asset_id": "ff1b1858-e263-49dc-901d-bd8f3ca0ac13",
  "input_asset_ids": [
   "bad5b160-fb00-401b-90e3-64652fe59154",
   "aacb72f4-e4d3-4331-8b05-48ae8f9cc866"
  ]
 },
 "S26sh9": {
  "input_fingerprint": "89a66d80546bd881",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground sewer divides into two passages. Charlie remains a worn gorilla-shaped robot in an old coat and hat, with blue-lit eyes and a faded UBIK chest logo. 현우: He is at the sewer fork without his outer garment; his face remains bruised and his dog-bitten leg remains injured. The contact card is concealed in his shoe. 페드로: He has reached the sewer fork and is about to take a separate route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground sewer divides into two passages. Charlie remains a worn gorilla-shaped robot in an old coat and hat, with blue-lit eyes and a faded UBIK chest logo. 현우: He is at the sewer fork without his outer garment; his face remains bruised and his dog-bitten leg remains injured. The contact card is concealed in his shoe. 페드로: He has reached the sewer fork and is about to take a separate route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground sewer divides into two passages. Charlie remains a worn gorilla-shaped robot in an old coat and hat, with blue-lit eyes and a faded UBIK chest logo. 현우: He is at the sewer fork without his outer garment; his face remains bruised and his dog-bitten leg remains injured. The contact card is concealed in his shoe. 페드로: He has reached the sewer fork and is about to take a separate route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9__bgfirst_bg.png",
     "asset_id": "ff1b1858-e263-49dc-901d-bd8f3ca0ac13",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S26sh9.png",
     "asset_id": "bad5b160-fb00-401b-90e3-64652fe59154",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L175B01.png",
     "asset_id": "aacb72f4-e4d3-4331-8b05-48ae8f9cc866",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우는 페드로를 응시하고, 페드로는 시선을 아래로 피함. 현우의 손이 페드로의 어깨 옷자락을 꽉 쥐고 있음.",
    "built_space": "하수도 갈림길과 좌측 사다리 등 주어진 레퍼런스 공간을 적절히 구현함.",
    "entities": "현우는 상처 입은 얼굴이나 레퍼런스와 다른 반팔 티셔츠를 입음. 페드로는 18세 설정과 전혀 다른 덥수룩한 수염의 중년 남성으로 나타남.",
    "hard_violations": [
     "[gemini-pro] 창조된 인물: 페드로가 레퍼런스(18세, 수염 없는 얼굴)와 완전히 다른 중년 남성으로 묘사됨"
    ],
    "physics": "어깨를 쥐고 있는 손의 형태와 서 있는 자세가 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "두 인물의 시선이 화면 밖 정면을 향함. 현우의 왼손이 페드로의 오른쪽 어깨에 놓여 있음.",
    "built_space": "하수도 두 갈래 길, 좌측 사다리와 바닥 구조물이 배경 레퍼런스와 정확히 일치함.",
    "entities": "현우(멍든 얼굴, 헝클어진 머리, 회색 셔츠)와 페드로(앳된 얼굴, 비니, 검은 후드) 모두 레퍼런스와 완벽히 일치함.",
    "hard_violations": [],
    "physics": "현우의 팔과 손이 페드로의 어깨 위에 자연스럽게 얹혀 중력을 따름."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물의 외모와 의상 레퍼런스를 정확히 반영하였으며, 지정된 배경과 인물 간의 행동 배치를 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "페드로가 18세 소년이라는 설정과 레퍼런스를 무시하고 수염이 난 중년 남성으로 잘못 묘사되어 치명적인 감점 요인이 되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "두 인물의 시선이 화면 밖 정면을 향함. 현우의 왼손이 페드로의 오른쪽 어깨에 놓여 있음.",
        "built_space": "하수도 두 갈래 길, 좌측 사다리와 바닥 구조물이 배경 레퍼런스와 정확히 일치함.",
        "entities": "현우(멍든 얼굴, 헝클어진 머리, 회색 셔츠)와 페드로(앳된 얼굴, 비니, 검은 후드) 모두 레퍼런스와 완벽히 일치함.",
        "hard_violations": [],
        "physics": "현우의 팔과 손이 페드로의 어깨 위에 자연스럽게 얹혀 중력을 따름."
       },
       {
        "label": "A",
        "direction": "현우는 페드로를 응시하고, 페드로는 시선을 아래로 피함. 현우의 손이 페드로의 어깨 옷자락을 꽉 쥐고 있음.",
        "built_space": "하수도 갈림길과 좌측 사다리 등 주어진 레퍼런스 공간을 적절히 구현함.",
        "entities": "현우는 상처 입은 얼굴이나 레퍼런스와 다른 반팔 티셔츠를 입음. 페드로는 18세 설정과 전혀 다른 덥수룩한 수염의 중년 남성으로 나타남.",
        "hard_violations": [
         "창조된 인물: 페드로가 레퍼런스(18세, 수염 없는 얼굴)와 완전히 다른 중년 남성으로 묘사됨"
        ],
        "physics": "어깨를 쥐고 있는 손의 형태와 서 있는 자세가 물리적으로 타당함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물의 외모와 의상 레퍼런스를 정확히 반영하였으며, 지정된 배경과 인물 간의 행동 배치를 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "페드로가 18세 소년이라는 설정과 레퍼런스를 무시하고 수염이 난 중년 남성으로 잘못 묘사되어 치명적인 감점 요인이 되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 인물의 시선이 화면 밖 정면을 향함. 현우의 왼손이 페드로의 오른쪽 어깨에 놓여 있음.",
        "built_space": "하수도 두 갈래 길, 좌측 사다리와 바닥 구조물이 배경 레퍼런스와 정확히 일치함.",
        "entities": "현우(멍든 얼굴, 헝클어진 머리, 회색 셔츠)와 페드로(앳된 얼굴, 비니, 검은 후드) 모두 레퍼런스와 완벽히 일치함.",
        "hard_violations": [],
        "physics": "현우의 팔과 손이 페드로의 어깨 위에 자연스럽게 얹혀 중력을 따름."
       },
       {
        "label": "A",
        "direction": "현우는 페드로를 응시하고, 페드로는 시선을 아래로 피함. 현우의 손이 페드로의 어깨 옷자락을 꽉 쥐고 있음.",
        "built_space": "하수도 갈림길과 좌측 사다리 등 주어진 레퍼런스 공간을 적절히 구현함.",
        "entities": "현우는 상처 입은 얼굴이나 레퍼런스와 다른 반팔 티셔츠를 입음. 페드로는 18세 설정과 전혀 다른 덥수룩한 수염의 중년 남성으로 나타남.",
        "hard_violations": [
         "창조된 인물: 페드로가 레퍼런스(18세, 수염 없는 얼굴)와 완전히 다른 중년 남성으로 묘사됨"
        ],
        "physics": "어깨를 쥐고 있는 손의 형태와 서 있는 자세가 물리적으로 타당함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "현우의 상체와 두 갈래 하수도, 인물 외형은 충실하지만 현우가 카메라 쪽을 보고 페드로는 화면 밖을 보아, 어깨를 꽉 쥐는 결연한 상호작용보다 기념사진처럼 읽힌다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "서로를 향한 시선과 어깨 옷감을 움켜쥔 손이 핵심 행동을 더 정확히 구현하지만, 페드로가 수염 난 성인처럼 보이고 현우의 셔츠가 반팔로 바뀐 점은 뚜렷한 불일치다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 페드로가 아니라 카메라 쪽을 향한다. 페드로는 화면 오른쪽 밖을 바라본다. 현우의 뻗은 손은 페드로의 가까운 어깨에 정확히 닿지만, 두 사람의 시선은 서로 연결되지 않는다.",
        "built_space": "왼쪽 벽에 고정 사다리 하나, 그 아래 받침 위에 원형 덮개 하나가 보인다. 인물 뒤로 중앙 분리벽을 사이에 둔 통로 입구 두 개와 젖은 수로가 보이며, 콘크리트 벽과 좁은 가장자리는 장소 참조와 대체로 일치한다. 두 사람은 분기점 앞에 상체 중심으로 배치되어 있다.",
        "entities": "인물은 두 명뿐이다. 현우는 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 얼굴의 상처, 겉옷 없는 회색 긴팔 셔츠가 참조와 잘 맞는다. 페드로도 앳되고 면도한 얼굴, 검은 비니와 어두운 후드가 참조에 가깝다. 다리 부상과 신발 속 카드는 화면 밖이므로 판정할 수 없으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 팔은 어깨에서 손목까지 자연스럽게 이어지고 손바닥과 손가락이 페드로의 어깨에 접촉한다. 다만 손가락이 비교적 펴져 있고 옷감이 뚜렷하게 잡혀 올라오지 않아 강하게 움켜쥐기보다는 손을 얹은 모습에 가깝다. 두 사람의 하체와 발은 잘려 있지만 상체는 정상적인 기립 자세이며 공중에 뜬 징후는 없다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 페드로의 얼굴을 향하고, 페드로도 현우의 얼굴을 내려다본다. 현우의 굽힌 손가락은 페드로의 가까운 어깨 옷감을 붙잡아 시선과 손의 목표가 같은 상대에게 모인다.",
        "built_space": "왼쪽 벽의 고정 사다리 하나와 낮은 콘크리트 천장, 젖은 수로, 양옆의 좁은 보행 가장자리가 보인다. 뒤쪽 왼쪽 통로는 명확하고 오른쪽 통로 입구는 페드로의 머리와 어깨에 상당 부분 가려져 일부만 보인다. 분리벽과 두 방향의 통로 구조는 유지되며, 두 사람은 분기점 앞에 서 있다. 원형 덮개가 있을 전경은 인물과 프레임에 가려져 있다.",
        "entities": "현우와 페드로에 해당하는 두 사람만 보인다. 현우의 젊은 동아시아계 외형, 검은 머리와 얼굴 상처는 맞지만 회색 반팔 티셔츠는 참조의 칼라 달린 긴팔 셔츠와 다르다. 페드로는 검은 비니와 어두운 후드는 맞으나 짙은 수염과 성숙한 얼굴 때문에 참조의 앳된 18세 남성과 크게 다르게 보인다. 다리와 신발 속 카드는 프레임 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 팔꿈치가 굽혀져 있고 손가락이 페드로의 어깨 옷감을 움켜쥐며, 손 아래 천에 당겨진 주름이 생긴다. 손과 어깨의 접촉 및 팔의 연결은 물리적으로 자연스럽다. 발은 보이지 않지만 두 몸통은 기립한 상태로 이어지고, 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우의 상체와 두 갈래 하수도, 인물 외형은 충실하지만 현우가 카메라 쪽을 보고 페드로는 화면 밖을 보아, 어깨를 꽉 쥐는 결연한 상호작용보다 기념사진처럼 읽힌다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "서로를 향한 시선과 어깨 옷감을 움켜쥔 손이 핵심 행동을 더 정확히 구현하지만, 페드로가 수염 난 성인처럼 보이고 현우의 셔츠가 반팔로 바뀐 점은 뚜렷한 불일치다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 페드로가 아니라 카메라 쪽을 향한다. 페드로는 화면 오른쪽 밖을 바라본다. 현우의 뻗은 손은 페드로의 가까운 어깨에 정확히 닿지만, 두 사람의 시선은 서로 연결되지 않는다.",
        "built_space": "왼쪽 벽에 고정 사다리 하나, 그 아래 받침 위에 원형 덮개 하나가 보인다. 인물 뒤로 중앙 분리벽을 사이에 둔 통로 입구 두 개와 젖은 수로가 보이며, 콘크리트 벽과 좁은 가장자리는 장소 참조와 대체로 일치한다. 두 사람은 분기점 앞에 상체 중심으로 배치되어 있다.",
        "entities": "인물은 두 명뿐이다. 현우는 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 얼굴의 상처, 겉옷 없는 회색 긴팔 셔츠가 참조와 잘 맞는다. 페드로도 앳되고 면도한 얼굴, 검은 비니와 어두운 후드가 참조에 가깝다. 다리 부상과 신발 속 카드는 화면 밖이므로 판정할 수 없으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 팔은 어깨에서 손목까지 자연스럽게 이어지고 손바닥과 손가락이 페드로의 어깨에 접촉한다. 다만 손가락이 비교적 펴져 있고 옷감이 뚜렷하게 잡혀 올라오지 않아 강하게 움켜쥐기보다는 손을 얹은 모습에 가깝다. 두 사람의 하체와 발은 잘려 있지만 상체는 정상적인 기립 자세이며 공중에 뜬 징후는 없다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 페드로의 얼굴을 향하고, 페드로도 현우의 얼굴을 내려다본다. 현우의 굽힌 손가락은 페드로의 가까운 어깨 옷감을 붙잡아 시선과 손의 목표가 같은 상대에게 모인다.",
        "built_space": "왼쪽 벽의 고정 사다리 하나와 낮은 콘크리트 천장, 젖은 수로, 양옆의 좁은 보행 가장자리가 보인다. 뒤쪽 왼쪽 통로는 명확하고 오른쪽 통로 입구는 페드로의 머리와 어깨에 상당 부분 가려져 일부만 보인다. 분리벽과 두 방향의 통로 구조는 유지되며, 두 사람은 분기점 앞에 서 있다. 원형 덮개가 있을 전경은 인물과 프레임에 가려져 있다.",
        "entities": "현우와 페드로에 해당하는 두 사람만 보인다. 현우의 젊은 동아시아계 외형, 검은 머리와 얼굴 상처는 맞지만 회색 반팔 티셔츠는 참조의 칼라 달린 긴팔 셔츠와 다르다. 페드로는 검은 비니와 어두운 후드는 맞으나 짙은 수염과 성숙한 얼굴 때문에 참조의 앳된 18세 남성과 크게 다르게 보인다. 다리와 신발 속 카드는 프레임 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 팔꿈치가 굽혀져 있고 손가락이 페드로의 어깨 옷감을 움켜쥐며, 손 아래 천에 당겨진 주름이 생긴다. 손과 어깨의 접촉 및 팔의 연결은 물리적으로 자연스럽다. 발은 보이지 않지만 두 몸통은 기립한 상태로 이어지고, 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.857
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.857
   },
   "violations": {
    "A": [
     "[gemini-pro] 창조된 인물: 페드로가 레퍼런스(18세, 수염 없는 얼굴)와 완전히 다른 중년 남성으로 묘사됨"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1857,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1857,
    "verdict_ko": "두 인물의 외모와 의상 레퍼런스를 정확히 반영하였으며, 지정된 배경과 인물 간의 행동 배치를 충실하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "페드로가 18세 소년이라는 설정과 레퍼런스를 무시하고 수염이 난 중년 남성으로 잘못 묘사되어 치명적인 감점 요인이 되었습니다.  ★위반: [gemini-pro] 창조된 인물: 페드로가 레퍼런스(18세, 수염 없는 얼굴)와 완전히 다른 중년 남성으로 묘사됨"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L175B01.png",
    "asset_id": "aacb72f4-e4d3-4331-8b05-48ae8f9cc866",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-16d3-70be-8f4d-20f03ebf21cb",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9__bgfirst_bg.png",
   "bg_asset_id": "ff1b1858-e263-49dc-901d-bd8f3ca0ac13",
   "bg_record_key": "S26sh9::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S26sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:06:34.205643+00:00",
  "fingerprint": "62d8b379f921d816a47fd29172f9a003890e4d3470ac5d8e7aaaecfe1d664321",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S26sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S26sh9_sel.png",
  "source_sha256": "bfa0f9a9d9ee1f086b4e6ce73d30bf66c19b4efc850099069df455d444d487be",
  "file": "S26sh9_cine.png",
  "staged_sha256": "50b7851b3a17bcdb97506d629930083fa849b059b67ba06044a57cded8481cb7",
  "latency_ms": 11054
 },
 "S26sh11::signage": {
  "fp": "bdc6108b03f52d68",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S26sh11": {
  "input_fingerprint": "316c61b4342fd76d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 얼굴을 등에 기댄 앰버를 꽉 업은 채 어두운 하수도 터널을 전속력으로 내달리며, 앞으로 몸을 깊게 기울인 현우의 땀 맺힌 mid-action 얼굴.\n\nLOCATION (lock): Inside a dark underground sewer passage beyond the junction, beneath the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chosen sewer passage (현우 is running through it carrying 앰버) — The passage recedes behind their shoulders; used as A narrow, softly resolved background preserves the direction of escape without competing with the faces.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the sewer's darkness and controlled tonal separation, allowing the stated perspiration and both faces to remain legible without adding a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer route continues beyond the fork. Charlie remains in the old coat and hat, with his worn metal body and blue-lit eyes unchanged. 현우: He is running in a load-bearing posture, his outer garment removed; his facial bruises and injured leg persist. The contact card remains concealed in his shoe. 앰버: She is being carried piggyback, with worsening coughing and the outer garment covering her mouth.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 얼굴을 등에 기댄 앰버를 꽉 업은 채 어두운 하수도 터널을 전속력으로 내달리며, 앞으로 몸을 깊게 기울인 현우의 땀 맺힌 mid-action 얼굴.\n\nLOCATION (lock): Inside a dark underground sewer passage beyond the junction, beneath the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chosen sewer passage (현우 is running through it carrying 앰버) — The passage recedes behind their shoulders; used as A narrow, softly resolved background preserves the direction of escape without competing with the faces.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the sewer's darkness and controlled tonal separation, allowing the stated perspiration and both faces to remain legible without adding a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer route continues beyond the fork. Charlie remains in the old coat and hat, with his worn metal body and blue-lit eyes unchanged. 현우: He is running in a load-bearing posture, his outer garment removed; his facial bruises and injured leg persist. The contact card remains concealed in his shoe. 앰버: She is being carried piggyback, with worsening coughing and the outer garment covering her mouth.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 얼굴을 등에 기댄 앰버를 꽉 업은 채 어두운 하수도 터널을 전속력으로 내달리며, 앞으로 몸을 깊게 기울인 현우의 땀 맺힌 mid-action 얼굴.\n\nLOCATION (lock): Inside a dark underground sewer passage beyond the junction, beneath the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chosen sewer passage (현우 is running through it carrying 앰버) — The passage recedes behind their shoulders; used as A narrow, softly resolved background preserves the direction of escape without competing with the faces.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the sewer's darkness and controlled tonal separation, allowing the stated perspiration and both faces to remain legible without adding a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer route continues beyond the fork. Charlie remains in the old coat and hat, with his worn metal body and blue-lit eyes unchanged. 현우: He is running in a load-bearing posture, his outer garment removed; his facial bruises and injured leg persist. The contact card remains concealed in his shoe. 앰버: She is being carried piggyback, with worsening coughing and the outer garment covering her mouth.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 앰버 모두 전방을 향해 시선을 둠.",
    "built_space": "어두운 하수도 터널 내부이며, 좌측 벽의 사다리와 깊어지는 배경 구조가 보임.",
    "entities": "땀과 멍이 있는 현우, 옷으로 입을 가린 금발의 앰버가 모두 일치함.",
    "hard_violations": [],
    "physics": "현우가 몸을 깊게 기울여 앰버를 안정적으로 등에 업고 달리는 자세가 자연스러움."
   },
   {
    "label": "B",
    "direction": "현우가 터널 전방을 주시함.",
    "built_space": "하수도 터널 내부, 좌측의 사다리와 우측 벽의 파이프가 확인됨.",
    "entities": "상처 입은 현우와 등에 매달린 금발의 앰버.",
    "hard_violations": [
     "[gemini-pro] 현우의 왼쪽 팔에 오른손 형태가 연결된 물리적으로 불가능한 해부학적 오류",
     "[gemini-pro] 목을 감싼 앰버의 손 구조 기형"
    ],
    "physics": "업고 달리는 자세 자체는 형성되었으나 해부학적 오류로 인해 물리적 접촉과 지지가 어긋남."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 샷 크기를 완벽히 준수했으며, 인물의 땀과 상처, 입을 가린 디테일 등을 충실히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 클로즈업보다 프레이밍이 넓게 잡혔으며, 손 부위에 치명적인 해부학적 오류가 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버 모두 전방을 향해 시선을 둠.",
        "built_space": "어두운 하수도 터널 내부이며, 좌측 벽의 사다리와 깊어지는 배경 구조가 보임.",
        "entities": "땀과 멍이 있는 현우, 옷으로 입을 가린 금발의 앰버가 모두 일치함.",
        "hard_violations": [],
        "physics": "현우가 몸을 깊게 기울여 앰버를 안정적으로 등에 업고 달리는 자세가 자연스러움."
       },
       {
        "label": "B",
        "direction": "현우가 터널 전방을 주시함.",
        "built_space": "하수도 터널 내부, 좌측의 사다리와 우측 벽의 파이프가 확인됨.",
        "entities": "상처 입은 현우와 등에 매달린 금발의 앰버.",
        "hard_violations": [
         "현우의 왼쪽 팔에 오른손 형태가 연결된 물리적으로 불가능한 해부학적 오류",
         "목을 감싼 앰버의 손 구조 기형"
        ],
        "physics": "업고 달리는 자세 자체는 형성되었으나 해부학적 오류로 인해 물리적 접촉과 지지가 어긋남."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 샷 크기를 완벽히 준수했으며, 인물의 땀과 상처, 입을 가린 디테일 등을 충실히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 클로즈업보다 프레이밍이 넓게 잡혔으며, 손 부위에 치명적인 해부학적 오류가 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버 모두 전방을 향해 시선을 둠.",
        "built_space": "어두운 하수도 터널 내부이며, 좌측 벽의 사다리와 깊어지는 배경 구조가 보임.",
        "entities": "땀과 멍이 있는 현우, 옷으로 입을 가린 금발의 앰버가 모두 일치함.",
        "hard_violations": [],
        "physics": "현우가 몸을 깊게 기울여 앰버를 안정적으로 등에 업고 달리는 자세가 자연스러움."
       },
       {
        "label": "B",
        "direction": "현우가 터널 전방을 주시함.",
        "built_space": "하수도 터널 내부, 좌측의 사다리와 우측 벽의 파이프가 확인됨.",
        "entities": "상처 입은 현우와 등에 매달린 금발의 앰버.",
        "hard_violations": [
         "현우의 왼쪽 팔에 오른손 형태가 연결된 물리적으로 불가능한 해부학적 오류",
         "목을 감싼 앰버의 손 구조 기형"
        ],
        "physics": "업고 달리는 자세 자체는 형성되었으나 해부학적 오류로 인해 물리적 접촉과 지지가 어긋남."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "땀 맺힌 현우의 얼굴을 크게 잡은 클로즈업과 깊은 전방 기울기, 등에 얼굴을 기댄 앰버의 입 가림이 핵심 지시를 가장 충실하게 구현한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "업고 달리는 동작과 앰버의 입 가림은 보이지만, 허리까지 넓힌 구도와 덜 깊은 상체 기울기가 얼굴 중심 클로즈업 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽 앞을 보며 그 방향으로 몸을 기울이고 있다. 앰버는 눈을 감고 얼굴을 현우의 목 뒤와 어깨에 붙였다. 통로는 두 사람 뒤쪽으로 이어져 이동 방향과 배경의 관계는 자연스럽다.",
        "built_space": "왼쪽 벽에 고정 사다리 한 개와 그 아래 돌출 발판 하나, 발판 위 원형 덮개 하나가 보인다. 뒤쪽에는 젖은 바닥과 작은 조명 한 개가 있으며, 오른쪽에는 수직관과 연결된 굵은 수평관이 보인다. 콘크리트와 습기는 참조와 유사하지만 오른쪽 배관은 참조에서 확인되지 않는다. 인물의 허리와 양쪽 벽을 넓게 보여 주어 배경이 좁고 부드럽게 남아야 한다는 지시보다 공간의 비중이 크다.",
        "entities": "인물은 현우와 앰버 두 명뿐이다. 현우는 앳된 동아시아계 남성으로 검은 머리, 얼굴 상처, 회색 셔츠가 참조와 대체로 맞으며 피부에 땀이 보인다. 앰버는 금발의 어린 여자아이로, 겉옷이 입을 가리고 있다. 다만 머리에 참조의 보호구는 보이지 않고 원래 상의는 겉옷에 가려 확인할 수 없다. 현우의 다리 부상과 신발 속 카드는 프레임 밖이다. 다른 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 몸은 현우의 등에 붙어 있고 두 팔이 그의 어깨와 목 앞을 감싼다. 현우의 한 손은 뒤쪽에서 아이의 다리 부근을 받치고 다른 손은 앞으로 굽혀져 있어 한 팔로 지지하며 달리는 순간으로 가능하다. 발은 프레임 밖이므로 착지는 확인할 수 없지만 공중에 떠 있는 몸으로 보이지 않는다. 상체 기울기와 팔 동작은 달리기를 나타내나 깊게 숙인 전속력 자세는 약하다."
       },
       {
        "label": "B",
        "direction": "현우의 시선과 얼굴, 숙인 상체가 모두 화면 오른쪽 전방의 탈출 진행 방향을 향한다. 앰버는 눈을 아래로 내리고 현우의 등 윗부분에 얼굴을 기댄다. 어깨 뒤로 어두운 통로가 물러나 있어 전방으로 도망치는 관계가 읽힌다.",
        "built_space": "왼쪽 벽에 고정 사다리 한 개의 일부가 보이고, 뒤에는 콘크리트 통로와 젖은 바닥, 작은 조명 한 개가 흐리게 남는다. 사다리 아래 발판은 인물과 크롭에 가려져 있다. 참조의 차갑고 어두운 하수도 재질과 조명 분위기를 유지하며, 얼굴 중심 구도 뒤에 배경을 제한한다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "현우와 앰버 두 명만 등장한다. 현우의 젊은 동아시아계 얼굴, 헝클어진 검은 머리, 회색 셔츠와 뺨의 상처가 참조에 부합하며 얼굴의 땀이 선명하다. 앰버는 금발의 어린 여자아이이고 얼굴 윤곽은 참조와 대체로 맞지만, 입과 하관이 겉옷에 가려 세부 동일성은 제한적으로만 확인된다. 참조의 머리 보호구는 보이지 않는다. 다리 부상, 신발 속 카드와 하체 의상은 적절한 크롭 밖이며 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "앰버의 가슴과 얼굴이 현우의 등과 어깨에 밀착되어 등으로 체중을 받는 업기 자세다. 현우는 상체를 깊이 숙이고 보이는 팔을 뒤로 보내 하중을 받는 자세를 취한다. 손의 실제 받침과 두 사람의 하체는 프레임 밖이므로 세부 지지는 확인할 수 없지만, 보이는 접촉에는 부유나 불가능한 자세가 없다. 겉옷은 앰버의 몸과 현우의 등 위에 걸쳐져 입을 덮는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "땀 맺힌 현우의 얼굴을 크게 잡은 클로즈업과 깊은 전방 기울기, 등에 얼굴을 기댄 앰버의 입 가림이 핵심 지시를 가장 충실하게 구현한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "업고 달리는 동작과 앰버의 입 가림은 보이지만, 허리까지 넓힌 구도와 덜 깊은 상체 기울기가 얼굴 중심 클로즈업 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽 앞을 보며 그 방향으로 몸을 기울이고 있다. 앰버는 눈을 감고 얼굴을 현우의 목 뒤와 어깨에 붙였다. 통로는 두 사람 뒤쪽으로 이어져 이동 방향과 배경의 관계는 자연스럽다.",
        "built_space": "왼쪽 벽에 고정 사다리 한 개와 그 아래 돌출 발판 하나, 발판 위 원형 덮개 하나가 보인다. 뒤쪽에는 젖은 바닥과 작은 조명 한 개가 있으며, 오른쪽에는 수직관과 연결된 굵은 수평관이 보인다. 콘크리트와 습기는 참조와 유사하지만 오른쪽 배관은 참조에서 확인되지 않는다. 인물의 허리와 양쪽 벽을 넓게 보여 주어 배경이 좁고 부드럽게 남아야 한다는 지시보다 공간의 비중이 크다.",
        "entities": "인물은 현우와 앰버 두 명뿐이다. 현우는 앳된 동아시아계 남성으로 검은 머리, 얼굴 상처, 회색 셔츠가 참조와 대체로 맞으며 피부에 땀이 보인다. 앰버는 금발의 어린 여자아이로, 겉옷이 입을 가리고 있다. 다만 머리에 참조의 보호구는 보이지 않고 원래 상의는 겉옷에 가려 확인할 수 없다. 현우의 다리 부상과 신발 속 카드는 프레임 밖이다. 다른 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 몸은 현우의 등에 붙어 있고 두 팔이 그의 어깨와 목 앞을 감싼다. 현우의 한 손은 뒤쪽에서 아이의 다리 부근을 받치고 다른 손은 앞으로 굽혀져 있어 한 팔로 지지하며 달리는 순간으로 가능하다. 발은 프레임 밖이므로 착지는 확인할 수 없지만 공중에 떠 있는 몸으로 보이지 않는다. 상체 기울기와 팔 동작은 달리기를 나타내나 깊게 숙인 전속력 자세는 약하다."
       },
       {
        "label": "A",
        "direction": "현우의 시선과 얼굴, 숙인 상체가 모두 화면 오른쪽 전방의 탈출 진행 방향을 향한다. 앰버는 눈을 아래로 내리고 현우의 등 윗부분에 얼굴을 기댄다. 어깨 뒤로 어두운 통로가 물러나 있어 전방으로 도망치는 관계가 읽힌다.",
        "built_space": "왼쪽 벽에 고정 사다리 한 개의 일부가 보이고, 뒤에는 콘크리트 통로와 젖은 바닥, 작은 조명 한 개가 흐리게 남는다. 사다리 아래 발판은 인물과 크롭에 가려져 있다. 참조의 차갑고 어두운 하수도 재질과 조명 분위기를 유지하며, 얼굴 중심 구도 뒤에 배경을 제한한다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "현우와 앰버 두 명만 등장한다. 현우의 젊은 동아시아계 얼굴, 헝클어진 검은 머리, 회색 셔츠와 뺨의 상처가 참조에 부합하며 얼굴의 땀이 선명하다. 앰버는 금발의 어린 여자아이이고 얼굴 윤곽은 참조와 대체로 맞지만, 입과 하관이 겉옷에 가려 세부 동일성은 제한적으로만 확인된다. 참조의 머리 보호구는 보이지 않는다. 다리 부상, 신발 속 카드와 하체 의상은 적절한 크롭 밖이며 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "앰버의 가슴과 얼굴이 현우의 등과 어깨에 밀착되어 등으로 체중을 받는 업기 자세다. 현우는 상체를 깊이 숙이고 보이는 팔을 뒤로 보내 하중을 받는 자세를 취한다. 손의 실제 받침과 두 사람의 하체는 프레임 밖이므로 세부 지지는 확인할 수 없지만, 보이는 접촉에는 부유나 불가능한 자세가 없다. 겉옷은 앰버의 몸과 현우의 등 위에 걸쳐져 입을 덮는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.095
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.845
   },
   "violations": {
    "B": [
     "[gemini-pro] 현우의 왼쪽 팔에 오른손 형태가 연결된 물리적으로 불가능한 해부학적 오류",
     "[gemini-pro] 목을 감싼 앰버의 손 구조 기형"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 845
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 클로즈업 샷 크기를 완벽히 준수했으며, 인물의 땀과 상처, 입을 가린 디테일 등을 충실히 구현함."
   },
   {
    "label": "B",
    "score": 845,
    "verdict_ko": "지정된 클로즈업보다 프레이밍이 넓게 잡혔으며, 손 부위에 치명적인 해부학적 오류가 발생함.  ★위반: [gemini-pro] 현우의 왼쪽 팔에 오른손 형태가 연결된 물리적으로 불가능한 해부학적 오류 / [gemini-pro] 목을 감싼 앰버의 손 구조 기형"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9_sel.png",
    "asset_id": "8a3e74d9-f257-4f86-bb97-97ca70dda6c9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-1a17-75f2-bcb3-a66accc8c340",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S26sh9"
  }
 },
 "S26sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:07:43.501594+00:00",
  "fingerprint": "2363ca9fb63b5462adc4b8a1a99db65e4c55f1616ddfba46ab8aef18264519e4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S26sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S26sh11_sel.png",
  "source_sha256": "ec22baa40cc2f21d5fb781308c0c03328497ff12c4f8ee5beafc0ea445a7dc87",
  "file": "S26sh11_cine.png",
  "staged_sha256": "2e2a938bab898c09ad24cef258771f577b2da683cdb6e56968791e77d79dbad1",
  "latency_ms": 9196
 },
 "S27sh12::signage": {
  "fp": "dbf48cb7b94e26f8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S27sh12": {
  "input_fingerprint": "c9ac4c90c1da27dd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights continue switching on and off along Charlie's route, and the approaching vehicle has its headlights on. Charlie retains his old coat and hat; the small bird has already flown away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights continue switching on and off along Charlie's route, and the approaching vehicle has its headlights on. Charlie retains his old coat and hat; the small bird has already flown away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights continue switching on and off along Charlie's route, and the approaching vehicle has its headlights on. Charlie retains his old coat and hat; the small bird has already flown away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12__bgfirst_bg.png",
     "asset_id": "164bd776-d0dc-43e4-bf32-6ef9f6356b78",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S27sh12.png",
     "asset_id": "29e8ef08-4d4f-4022-ac9f-847c4599b99d",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L177B01.png",
     "asset_id": "785d4e92-6642-43a7-a3f4-70e3c4f7a9b3",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_sewer_and_manhole_sel.png",
     "asset_id": "2b4b407b-73bc-405e-8265-11c6b8d30ef7",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "트럭은 화면 전방 대각선을 향해 전조등을 비추며 주행 중이고, 찰리 역시 도로를 따라 전방으로 달리고 있어 두 경로가 수렴하는 형태임.",
    "built_space": "로케이션 레퍼런스의 건물, 가로등, 젖은 도로가 구도에 맞게 구현되었으나, 좌측 하단의 맨홀 뚜껑이 구조물 레퍼런스와 달리 닫혀 있음.",
    "entities": "좌측에 거대한 트럭이 위치하며, 찰리는 고릴라형 기계 몸체, 트렌치코트, 페도라 등 명시된 외형 요소를 정확히 갖춤.",
    "hard_violations": [],
    "physics": "찰리는 오른발로 지면을 단단히 딛고 왼발을 공중에 든 상태로 안정적인 달리기 자세를 취하고 있음."
   },
   {
    "label": "B",
    "direction": "트럭은 우측을 향해 헤드라이트를 비추며 직진하고, 찰리는 트럭의 앞을 가로질러 좌측으로 도약하고 있음.",
    "built_space": "배경의 도로와 건물이 측면 앵글로 배치되었으며, 우측 하단에 열린 맨홀 구조물이 두 개로 복제되어 나타남.",
    "entities": "트럭이 존재하며, 찰리의 복장과 기계 외형은 레퍼런스와 일치함.",
    "hard_violations": [
     "[gemini-pro] 로케이션 및 구조물 레퍼런스에 단일하게 존재하는 맨홀이 두 개로 복제되어 배치됨 (Duplicated fitting)."
    ],
    "physics": "찰리의 두 발이 모두 공중에 떠 있어 '한 발이 공중에 뜬' 지시와 다름."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "대각선으로 뻗은 도로, 차량의 배치, 한 발을 공중에 띄우고 달리는 찰리의 묘사 등 프레이밍 지시를 충실히 따랐으나 맨홀이 닫혀 있는 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시된 대각선 구도 대신 측면 앵글을 렌더링했으며, 한 발이 아닌 두 발이 모두 떠 있고 맨홀이 두 개로 복제되는 치명적 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭은 화면 전방 대각선을 향해 전조등을 비추며 주행 중이고, 찰리 역시 도로를 따라 전방으로 달리고 있어 두 경로가 수렴하는 형태임.",
        "built_space": "로케이션 레퍼런스의 건물, 가로등, 젖은 도로가 구도에 맞게 구현되었으나, 좌측 하단의 맨홀 뚜껑이 구조물 레퍼런스와 달리 닫혀 있음.",
        "entities": "좌측에 거대한 트럭이 위치하며, 찰리는 고릴라형 기계 몸체, 트렌치코트, 페도라 등 명시된 외형 요소를 정확히 갖춤.",
        "hard_violations": [],
        "physics": "찰리는 오른발로 지면을 단단히 딛고 왼발을 공중에 든 상태로 안정적인 달리기 자세를 취하고 있음."
       },
       {
        "label": "B",
        "direction": "트럭은 우측을 향해 헤드라이트를 비추며 직진하고, 찰리는 트럭의 앞을 가로질러 좌측으로 도약하고 있음.",
        "built_space": "배경의 도로와 건물이 측면 앵글로 배치되었으며, 우측 하단에 열린 맨홀 구조물이 두 개로 복제되어 나타남.",
        "entities": "트럭이 존재하며, 찰리의 복장과 기계 외형은 레퍼런스와 일치함.",
        "hard_violations": [
         "로케이션 및 구조물 레퍼런스에 단일하게 존재하는 맨홀이 두 개로 복제되어 배치됨 (Duplicated fitting)."
        ],
        "physics": "찰리의 두 발이 모두 공중에 떠 있어 '한 발이 공중에 뜬' 지시와 다름."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "대각선으로 뻗은 도로, 차량의 배치, 한 발을 공중에 띄우고 달리는 찰리의 묘사 등 프레이밍 지시를 충실히 따랐으나 맨홀이 닫혀 있는 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시된 대각선 구도 대신 측면 앵글을 렌더링했으며, 한 발이 아닌 두 발이 모두 떠 있고 맨홀이 두 개로 복제되는 치명적 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "트럭은 화면 전방 대각선을 향해 전조등을 비추며 주행 중이고, 찰리 역시 도로를 따라 전방으로 달리고 있어 두 경로가 수렴하는 형태임.",
        "built_space": "로케이션 레퍼런스의 건물, 가로등, 젖은 도로가 구도에 맞게 구현되었으나, 좌측 하단의 맨홀 뚜껑이 구조물 레퍼런스와 달리 닫혀 있음.",
        "entities": "좌측에 거대한 트럭이 위치하며, 찰리는 고릴라형 기계 몸체, 트렌치코트, 페도라 등 명시된 외형 요소를 정확히 갖춤.",
        "hard_violations": [],
        "physics": "찰리는 오른발로 지면을 단단히 딛고 왼발을 공중에 든 상태로 안정적인 달리기 자세를 취하고 있음."
       },
       {
        "label": "B",
        "direction": "트럭은 우측을 향해 헤드라이트를 비추며 직진하고, 찰리는 트럭의 앞을 가로질러 좌측으로 도약하고 있음.",
        "built_space": "배경의 도로와 건물이 측면 앵글로 배치되었으며, 우측 하단에 열린 맨홀 구조물이 두 개로 복제되어 나타남.",
        "entities": "트럭이 존재하며, 찰리의 복장과 기계 외형은 레퍼런스와 일치함.",
        "hard_violations": [
         "로케이션 및 구조물 레퍼런스에 단일하게 존재하는 맨홀이 두 개로 복제되어 배치됨 (Duplicated fitting)."
        ],
        "physics": "찰리의 두 발이 모두 공중에 떠 있어 '한 발이 공중에 뜬' 지시와 다름."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "차량이 찰리를 향하는 충돌 방향과 장소 재현은 더 정확하지만, 차량 뒤쪽이 잘리고 찰리가 중경보다 크게 배치되어 지정 와이드 구도에는 미달한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "차량 전체를 왼쪽에 담았지만 차량과 찰리의 진행 경로가 좌우로 분리되어, 찰리를 향해 돌진하는 핵심 관계가 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭 전면과 전조등은 화면 오른쪽의 찰리 쪽을 향한다. 찰리는 몸과 흰 얼굴을 왼쪽 트럭 쪽으로 돌리고 그쪽으로 달리는 자세여서, 두 대상 사이의 간격과 정면 충돌 가능성이 읽힌다. 무기나 별도로 겨냥하는 소품은 없다.",
        "built_space": "중앙의 좁은 골목, 양쪽 낡은 회벽 건물, 오른쪽 창문과 실외기 및 셔터가 장소 사진과 가깝다. 오른쪽 흰 가로등 하나와 골목 안 작은 점등 하나가 뚜렷하며, 왼쪽 조명은 차량에 대부분 가려진다. 오른쪽 노면에는 방사형 내부 철물이 있는 열린 맨홀 하나와 옆에 놓인 뚜껑 하나가 있어 구조 참고를 반영한다. 왼쪽 아래에는 별도의 닫힌 맨홀이 보인다. 차량 앞면과 옆면은 보이지만 화물칸 뒤쪽이 왼쪽 경계 밖으로 잘렸고, 도로의 대각선 원근보다 횡방향 충돌 장면이 강조된다.",
        "entities": "찰리 한 명과 대형 화물차 한 대가 보이며 추가 인물이나 새는 없다. 찰리는 흰 마스크형 얼굴, 모자, 낡은 코트, 샌드 베이지 장갑, 육중한 어깨와 긴 팔을 갖춰 캐릭터 참고에 가깝다. 얼굴이 가려져 인간의 나이·성별·민족성은 판별할 수 없다. 차량 전조등은 켜져 있고 밤 장면이다. 판독 가능한 글자는 보이지 않는다. 전조등 주변의 옅은 연무는 추가 대기 효과를 피하라는 요구와 다소 어긋난다.",
        "hard_violations": [],
        "physics": "트럭은 노면에 닿은 타이어로 지지되고, 맨홀 뚜껑은 도로 위에 놓여 있다. 찰리의 두 발은 모두 노면과 떨어져 보인다. 다만 앞으로 뻗은 다리와 뒤로 접힌 다리, 기울어진 몸통과 팔 동작이 달리기의 짧은 체공 구간을 이루며, 앞발 아래에 착지할 도로가 있다. 정적인 무지지 부유로 단정할 자세는 아니지만, 한 발을 든 달리기 순간보다 도약이 강하게 읽힌다."
       },
       {
        "label": "B",
        "direction": "트럭은 화면 왼쪽 차로에서 카메라 쪽으로 향하고, 찰리는 오른쪽에서 카메라 쪽과 화면 오른쪽을 향해 달린다. 찰리의 얼굴도 트럭이 아닌 오른쪽 진행 방향을 향한다. 트럭 전방과 찰리의 경로 사이에 넓은 횡방향 간격이 있어, 두 경로가 수렴하기보다는 나란히 분리되어 보인다.",
        "built_space": "대형 트럭 전체가 왼쪽에 들어오고 앞면과 오른쪽 측면이 보인다. 도로는 중앙 소실점으로 길게 이어지지만 찰리 역시 비교적 가까운 전경에 있다. 양쪽의 낡은 저층 건물, 셔터, 차양과 실외기는 장소의 재질감을 따른다. 가까운 가로등은 왼쪽 주황색 하나와 오른쪽 흰색 하나이며, 원경에는 여러 점등이 반복된다. 다만 장소 사진의 두 건물 사이 좁은 골목 배치는 확인되지 않는다. 왼쪽 전경의 맨홀 하나는 닫혀 있어, 구조 참고의 열린 구멍과 옆으로 옮겨진 뚜껑 상태를 재현하지 못한다.",
        "entities": "찰리 한 명과 대형 화물차 한 대가 보이며 다른 인물이나 새는 식별되지 않는다. 찰리의 모자, 낡은 코트, 베이지 장갑판, 흰 마스크형 얼굴은 일치하지만 다리가 상대적으로 길어 참고의 짧은 다리와 고릴라형 비례는 약해졌다. 가려진 얼굴로 나이·성별·민족성을 확인할 수는 없다. 밤이고 전조등은 켜져 있다. 판독 가능한 글자는 없지만 그릴 중앙에 원형 브랜드 표장처럼 보이는 장식이 있다.",
        "hard_violations": [],
        "physics": "트럭 타이어는 노면에 닿아 차체를 지지한다. 찰리는 한 다리를 뒤로 접고 다른 다리를 아래로 내민 달리기 자세이며, 아래쪽 발도 그림자와 조금 떨어져 보여 짧은 체공 순간으로 읽힌다. 무릎 굽힘과 팔의 반대 동작이 달리기 추진을 설명하고 내민 발 아래에 착지면이 있으므로, 근거 없는 정적 부유로 보기는 어렵다. 코트와 모자는 몸에 착용되어 있고 별도로 떠 있는 소품은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "차량이 찰리를 향하는 충돌 방향과 장소 재현은 더 정확하지만, 차량 뒤쪽이 잘리고 찰리가 중경보다 크게 배치되어 지정 와이드 구도에는 미달한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "차량 전체를 왼쪽에 담았지만 차량과 찰리의 진행 경로가 좌우로 분리되어, 찰리를 향해 돌진하는 핵심 관계가 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "트럭 전면과 전조등은 화면 오른쪽의 찰리 쪽을 향한다. 찰리는 몸과 흰 얼굴을 왼쪽 트럭 쪽으로 돌리고 그쪽으로 달리는 자세여서, 두 대상 사이의 간격과 정면 충돌 가능성이 읽힌다. 무기나 별도로 겨냥하는 소품은 없다.",
        "built_space": "중앙의 좁은 골목, 양쪽 낡은 회벽 건물, 오른쪽 창문과 실외기 및 셔터가 장소 사진과 가깝다. 오른쪽 흰 가로등 하나와 골목 안 작은 점등 하나가 뚜렷하며, 왼쪽 조명은 차량에 대부분 가려진다. 오른쪽 노면에는 방사형 내부 철물이 있는 열린 맨홀 하나와 옆에 놓인 뚜껑 하나가 있어 구조 참고를 반영한다. 왼쪽 아래에는 별도의 닫힌 맨홀이 보인다. 차량 앞면과 옆면은 보이지만 화물칸 뒤쪽이 왼쪽 경계 밖으로 잘렸고, 도로의 대각선 원근보다 횡방향 충돌 장면이 강조된다.",
        "entities": "찰리 한 명과 대형 화물차 한 대가 보이며 추가 인물이나 새는 없다. 찰리는 흰 마스크형 얼굴, 모자, 낡은 코트, 샌드 베이지 장갑, 육중한 어깨와 긴 팔을 갖춰 캐릭터 참고에 가깝다. 얼굴이 가려져 인간의 나이·성별·민족성은 판별할 수 없다. 차량 전조등은 켜져 있고 밤 장면이다. 판독 가능한 글자는 보이지 않는다. 전조등 주변의 옅은 연무는 추가 대기 효과를 피하라는 요구와 다소 어긋난다.",
        "hard_violations": [],
        "physics": "트럭은 노면에 닿은 타이어로 지지되고, 맨홀 뚜껑은 도로 위에 놓여 있다. 찰리의 두 발은 모두 노면과 떨어져 보인다. 다만 앞으로 뻗은 다리와 뒤로 접힌 다리, 기울어진 몸통과 팔 동작이 달리기의 짧은 체공 구간을 이루며, 앞발 아래에 착지할 도로가 있다. 정적인 무지지 부유로 단정할 자세는 아니지만, 한 발을 든 달리기 순간보다 도약이 강하게 읽힌다."
       },
       {
        "label": "A",
        "direction": "트럭은 화면 왼쪽 차로에서 카메라 쪽으로 향하고, 찰리는 오른쪽에서 카메라 쪽과 화면 오른쪽을 향해 달린다. 찰리의 얼굴도 트럭이 아닌 오른쪽 진행 방향을 향한다. 트럭 전방과 찰리의 경로 사이에 넓은 횡방향 간격이 있어, 두 경로가 수렴하기보다는 나란히 분리되어 보인다.",
        "built_space": "대형 트럭 전체가 왼쪽에 들어오고 앞면과 오른쪽 측면이 보인다. 도로는 중앙 소실점으로 길게 이어지지만 찰리 역시 비교적 가까운 전경에 있다. 양쪽의 낡은 저층 건물, 셔터, 차양과 실외기는 장소의 재질감을 따른다. 가까운 가로등은 왼쪽 주황색 하나와 오른쪽 흰색 하나이며, 원경에는 여러 점등이 반복된다. 다만 장소 사진의 두 건물 사이 좁은 골목 배치는 확인되지 않는다. 왼쪽 전경의 맨홀 하나는 닫혀 있어, 구조 참고의 열린 구멍과 옆으로 옮겨진 뚜껑 상태를 재현하지 못한다.",
        "entities": "찰리 한 명과 대형 화물차 한 대가 보이며 다른 인물이나 새는 식별되지 않는다. 찰리의 모자, 낡은 코트, 베이지 장갑판, 흰 마스크형 얼굴은 일치하지만 다리가 상대적으로 길어 참고의 짧은 다리와 고릴라형 비례는 약해졌다. 가려진 얼굴로 나이·성별·민족성을 확인할 수는 없다. 밤이고 전조등은 켜져 있다. 판독 가능한 글자는 없지만 그릴 중앙에 원형 브랜드 표장처럼 보이는 장식이 있다.",
        "hard_violations": [],
        "physics": "트럭 타이어는 노면에 닿아 차체를 지지한다. 찰리는 한 다리를 뒤로 접고 다른 다리를 아래로 내민 달리기 자세이며, 아래쪽 발도 그림자와 조금 떨어져 보여 짧은 체공 순간으로 읽힌다. 무릎 굽힘과 팔의 반대 동작이 달리기 추진을 설명하고 내민 발 아래에 착지면이 있으므로, 근거 없는 정적 부유로 보기는 어렵다. 코트와 모자는 몸에 착용되어 있고 별도로 떠 있는 소품은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.667,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 로케이션 및 구조물 레퍼런스에 단일하게 존재하는 맨홀이 두 개로 복제되어 배치됨 (Duplicated fitting)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1667,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1667,
    "verdict_ko": "대각선으로 뻗은 도로, 차량의 배치, 한 발을 공중에 띄우고 달리는 찰리의 묘사 등 프레이밍 지시를 충실히 따랐으나 맨홀이 닫혀 있는 점이 아쉽습니다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "지시된 대각선 구도 대신 측면 앵글을 렌더링했으며, 한 발이 아닌 두 발이 모두 떠 있고 맨홀이 두 개로 복제되는 치명적 위반이 발생했습니다.  ★위반: [gemini-pro] 로케이션 및 구조물 레퍼런스에 단일하게 존재하는 맨홀이 두 개로 복제되어 배치됨 (Duplicated fitting)."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L177B01.png",
    "asset_id": "785d4e92-6642-43a7-a3f4-70e3c4f7a9b3",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_sewer_and_manhole_sel.png",
    "asset_id": "2b4b407b-73bc-405e-8265-11c6b8d30ef7",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-1bd1-7b4f-886e-bda939d73876",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12__bgfirst_bg.png",
   "bg_asset_id": "164bd776-d0dc-43e4-bf32-6ef9f6356b78",
   "bg_record_key": "S27sh12::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S27sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:09:00.942338+00:00",
  "fingerprint": "96e1793f4538e0260a5d089e8fe560034655e63792e4867b95b7e326724553c7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S27sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S27sh12_sel.png",
  "source_sha256": "661dcb0862011a0b91688fc385294972559b3407334beb76c2b137e75a2899c2",
  "file": "S27sh12_cine.png",
  "staged_sha256": "638053c738cf65adb6d801273c7c5c0a183d3bf3d797cefa11e40bad36dbe8af",
  "latency_ms": 11273
 },
 "S27sh18::signage": {
  "fp": "df1c5668dd64f889",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S27sh18": {
  "input_fingerprint": "f9498243f8427013",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 헤드라이트 불빛을 등진 채 찰리 쪽을 향해 한 발을 내디딘 신부의 짙은 실루엣 전신.\n\nLOCATION (lock): On the road immediately in front of a stopped van in the refugee settlement, backlit by its headlights. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stopped vehicle (Stopped after braking in front of 찰리) — Its front faces obliquely toward the camera behind 신부; used as Anchors the approaching figure to the vehicle from which he emerged; Road between 찰리 and 신부 (신부 is crossing the remaining separation on foot) — The visible ground connects the lower-left foreground to the center-right midground; used as Makes the approach and the natural upward eyeline physically legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The vehicle's headlights backlight 신부 into a dense silhouette, withholding his facial detail at this first step while preserving his full-body outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same roadway, immediate roadside surroundings, vehicle exterior, and nighttime headlight illumination. Exclude any transient motion effects from the vehicle's approach.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van has stopped abruptly and its headlights remain on. Charlie, still wearing the old coat and hat, stands in front of it; the bird is no longer perched on him. 신부: He has stepped out of the van and is approaching on foot, still wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 헤드라이트 불빛을 등진 채 찰리 쪽을 향해 한 발을 내디딘 신부의 짙은 실루엣 전신.\n\nLOCATION (lock): On the road immediately in front of a stopped van in the refugee settlement, backlit by its headlights. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stopped vehicle (Stopped after braking in front of 찰리) — Its front faces obliquely toward the camera behind 신부; used as Anchors the approaching figure to the vehicle from which he emerged; Road between 찰리 and 신부 (신부 is crossing the remaining separation on foot) — The visible ground connects the lower-left foreground to the center-right midground; used as Makes the approach and the natural upward eyeline physically legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The vehicle's headlights backlight 신부 into a dense silhouette, withholding his facial detail at this first step while preserving his full-body outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same roadway, immediate roadside surroundings, vehicle exterior, and nighttime headlight illumination. Exclude any transient motion effects from the vehicle's approach.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van has stopped abruptly and its headlights remain on. Charlie, still wearing the old coat and hat, stands in front of it; the bird is no longer perched on him. 신부: He has stepped out of the van and is approaching on foot, still wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 헤드라이트 불빛을 등진 채 찰리 쪽을 향해 한 발을 내디딘 신부의 짙은 실루엣 전신.\n\nLOCATION (lock): On the road immediately in front of a stopped van in the refugee settlement, backlit by its headlights. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stopped vehicle (Stopped after braking in front of 찰리) — Its front faces obliquely toward the camera behind 신부; used as Anchors the approaching figure to the vehicle from which he emerged; Road between 찰리 and 신부 (신부 is crossing the remaining separation on foot) — The visible ground connects the lower-left foreground to the center-right midground; used as Makes the approach and the natural upward eyeline physically legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The vehicle's headlights backlight 신부 into a dense silhouette, withholding his facial detail at this first step while preserving his full-body outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same roadway, immediate roadside surroundings, vehicle exterior, and nighttime headlight illumination. Exclude any transient motion effects from the vehicle's approach.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van has stopped abruptly and its headlights remain on. Charlie, still wearing the old coat and hat, stands in front of it; the bird is no longer perched on him. 신부: He has stepped out of the van and is approaching on foot, still wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부가 찰리 쪽이 아닌 카메라 정면을 향해 걸어오고 있으며, 찰리 역시 다른 곳을 응시하고 있어 둘 사이의 상호작용 동선이 맞지 않습니다.",
    "built_space": "도로 배경은 유사하나, 이전 샷에서 지정된 대형 트럭 대신 형태가 완전히 다른 승합차(밴)가 배치되었습니다.",
    "entities": "찰리의 외형 및 복장, 그리고 신부의 짙은 실루엣 묘사는 프롬프트에 부합하게 구현되었습니다.",
    "hard_violations": [
     "[gemini-pro] 이전 샷 레퍼런스의 차량 외관(트럭)을 그대로 유지해야 하는 잠금(Lock) 지시를 무시하고 다른 형태의 차량(승합차)을 생성함."
    ],
    "physics": "두 캐릭터 모두 지면에 안정적으로 발을 딛고 있으며, 중력과 자세에 어긋나는 요소는 없습니다."
   },
   {
    "label": "B",
    "direction": "신부가 차량의 헤드라이트를 등지고 화면 우측에 서 있는 찰리를 향해 명확하게 발을 내디디며 이동하고 있습니다.",
    "built_space": "이전 샷 레퍼런스에 등장했던 박스형 트럭과 도로, 주변 건물들이 정확한 위치와 형태로 유지되어 있습니다.",
    "entities": "찰리는 지정된 고릴라형 사이보그 외형과 복장을 유지하고 있으며, 신부는 디테일이 가려진 짙은 실루엣으로 올바르게 묘사되었습니다.",
    "hard_violations": [],
    "physics": "신부의 걷는 자세와 찰리의 서 있는 자세 모두 지면과 자연스럽게 접촉하여 무게감을 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "이전 샷의 트럭 외관과 배경을 완벽히 유지했으며, 헤드라이트를 등지고 찰리를 향해 걷는 신부의 실루엣 동선을 정확하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스의 차량 외관을 유지하라는 지침을 위반하여 차량이 승합차로 바뀌었으며, 찰리를 향해 걷는 지시사항을 따르지 않았습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "신부가 차량의 헤드라이트를 등지고 화면 우측에 서 있는 찰리를 향해 명확하게 발을 내디디며 이동하고 있습니다.",
        "built_space": "이전 샷 레퍼런스에 등장했던 박스형 트럭과 도로, 주변 건물들이 정확한 위치와 형태로 유지되어 있습니다.",
        "entities": "찰리는 지정된 고릴라형 사이보그 외형과 복장을 유지하고 있으며, 신부는 디테일이 가려진 짙은 실루엣으로 올바르게 묘사되었습니다.",
        "hard_violations": [],
        "physics": "신부의 걷는 자세와 찰리의 서 있는 자세 모두 지면과 자연스럽게 접촉하여 무게감을 지탱하고 있습니다."
       },
       {
        "label": "A",
        "direction": "신부가 찰리 쪽이 아닌 카메라 정면을 향해 걸어오고 있으며, 찰리 역시 다른 곳을 응시하고 있어 둘 사이의 상호작용 동선이 맞지 않습니다.",
        "built_space": "도로 배경은 유사하나, 이전 샷에서 지정된 대형 트럭 대신 형태가 완전히 다른 승합차(밴)가 배치되었습니다.",
        "entities": "찰리의 외형 및 복장, 그리고 신부의 짙은 실루엣 묘사는 프롬프트에 부합하게 구현되었습니다.",
        "hard_violations": [
         "이전 샷 레퍼런스의 차량 외관(트럭)을 그대로 유지해야 하는 잠금(Lock) 지시를 무시하고 다른 형태의 차량(승합차)을 생성함."
        ],
        "physics": "두 캐릭터 모두 지면에 안정적으로 발을 딛고 있으며, 중력과 자세에 어긋나는 요소는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "이전 샷의 트럭 외관과 배경을 완벽히 유지했으며, 헤드라이트를 등지고 찰리를 향해 걷는 신부의 실루엣 동선을 정확하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스의 차량 외관을 유지하라는 지침을 위반하여 차량이 승합차로 바뀌었으며, 찰리를 향해 걷는 지시사항을 따르지 않았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부가 차량의 헤드라이트를 등지고 화면 우측에 서 있는 찰리를 향해 명확하게 발을 내디디며 이동하고 있습니다.",
        "built_space": "이전 샷 레퍼런스에 등장했던 박스형 트럭과 도로, 주변 건물들이 정확한 위치와 형태로 유지되어 있습니다.",
        "entities": "찰리는 지정된 고릴라형 사이보그 외형과 복장을 유지하고 있으며, 신부는 디테일이 가려진 짙은 실루엣으로 올바르게 묘사되었습니다.",
        "hard_violations": [],
        "physics": "신부의 걷는 자세와 찰리의 서 있는 자세 모두 지면과 자연스럽게 접촉하여 무게감을 지탱하고 있습니다."
       },
       {
        "label": "A",
        "direction": "신부가 찰리 쪽이 아닌 카메라 정면을 향해 걸어오고 있으며, 찰리 역시 다른 곳을 응시하고 있어 둘 사이의 상호작용 동선이 맞지 않습니다.",
        "built_space": "도로 배경은 유사하나, 이전 샷에서 지정된 대형 트럭 대신 형태가 완전히 다른 승합차(밴)가 배치되었습니다.",
        "entities": "찰리의 외형 및 복장, 그리고 신부의 짙은 실루엣 묘사는 프롬프트에 부합하게 구현되었습니다.",
        "hard_violations": [
         "이전 샷 레퍼런스의 차량 외관(트럭)을 그대로 유지해야 하는 잠금(Lock) 지시를 무시하고 다른 형태의 차량(승합차)을 생성함."
        ],
        "physics": "두 캐릭터 모두 지면에 안정적으로 발을 딛고 있으며, 중력과 자세에 어긋나는 요소는 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "신부가 찰리 쪽으로 내딛는 역광 전신과 이전 차량 외관을 유지하지만, 신부가 왼쪽 전경에 있어 지정된 접근 공간의 배치는 어긋난다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "왼쪽 전경과 오른쪽 중경의 배치는 가깝지만, 신부가 찰리보다 카메라 정면으로 걸으며 이전 장면의 차량도 다른 차종으로 바뀌었다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 얼굴과 앞발은 화면 오른쪽의 찰리 쪽을 향한다. 정확한 눈동자는 역광으로 보이지 않지만 접근 대상은 찰리로 읽힌다. 찰리는 몸통이 거의 카메라 정면을 향해 신부와의 시선 교환은 약하다. 차량 전면과 켜진 전조등은 신부 뒤에서 카메라 쪽을 향한다.",
        "built_space": "왼쪽에 이전 장면과 같은 대형 상자형 화물차 한 대가 있고, 전면의 위아래 좌우 등 네 개 중 일부가 신부에게 가려진다. 양쪽 낡은 건물, 셔터, 왼쪽 차양, 가까운 좌우 가로등 각 한 개, 전선과 왼쪽 전경 맨홀이 유지된다. 신부는 왼쪽 전경, 찰리는 오른쪽 중경에 있어, 요구된 왼쪽 전경에서 오른쪽 중경의 신부로 이어지는 접근 배치는 반대다. 젖은 노면의 광원 반사는 가능한 위치에 있다.",
        "entities": "신부 한 명과 찰리 한 명만 보인다. 신부는 짧은 머리의 남성으로 검은 성직자 옷과 흰 칼라를 착용하며, 얼굴이 어두워 한국인 60대의 세부 인상과 주름은 확인할 수 없다. 찰리의 낡은 모자와 코트, 모래색 장갑판, 흰 마스크형 얼굴, 긴 팔이 참조와 대체로 맞는다. 새는 없다. 차량은 문구상의 밴보다는 이전 장면에 실제로 나온 화물차 외관을 따른다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부는 앞발 뒤꿈치가 노면에 닿는 보행 순간이며 반대쪽 발도 노면 가까이 있어 정상적인 체중 이동으로 설명된다. 찰리는 두 발로 도로에 서 있고 차량은 바퀴로 지지된다. 공중에 지지 없이 떠 있는 신체나 소품은 없다."
       },
       {
        "label": "B",
        "direction": "신부의 얼굴, 몸통과 내딛는 발은 거의 카메라 정면을 향한다. 찰리는 신부의 왼쪽 전경에 있으므로 이 보행 방향은 찰리를 향한 대각선 접근으로 명확하게 이어지지 않는다. 찰리의 얼굴은 약간 오른쪽으로 돌아 신부 방향에 더 가깝다. 차량 전조등은 신부 뒤에서 카메라 쪽을 비춘다.",
        "built_space": "찰리가 왼쪽 전경, 신부가 오른쪽 중경에 있고 그 뒤에 비스듬히 정차한 밴 한 대가 있어 요구된 도로 연결 배치에 가깝다. 밴에는 밝은 전조등 두 개가 보인다. 양쪽 건물과 셔터, 왼쪽 차양, 가까운 좌우 가로등 각 한 개, 전선과 맨홀은 이전 장소를 대체로 유지한다. 그러나 이전 장면의 높은 화물차가 낮은 승합형 밴으로 교체되어 차량 외관 연속성이 깨진다. 노면 반사는 광원과 지면 관계에 맞는다.",
        "entities": "신부 한 명과 찰리 한 명이 보인다. 신부의 검은 옷과 작은 흰 성직자 칼라는 보이지만 얼굴은 실루엣에 가려 나이와 참조 얼굴의 세부를 확인할 수 없다. 찰리는 모자, 낡은 코트, 모래색 장갑판, 흰 얼굴과 육중하고 긴 팔을 유지한다. 새는 없다. 차량은 밴이라는 명칭에는 맞지만 이전 장면의 실제 차량과는 다르다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부는 앞쪽 발을 노면에 디디고 뒤쪽 발을 옮기는 보행 자세로, 지지와 체중 이동이 성립한다. 찰리의 두 발은 도로에 닿아 있고 밴도 바퀴로 지지된다. 떠 있는 물체나 불가능한 관절 배치는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "신부가 찰리 쪽으로 내딛는 역광 전신과 이전 차량 외관을 유지하지만, 신부가 왼쪽 전경에 있어 지정된 접근 공간의 배치는 어긋난다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "왼쪽 전경과 오른쪽 중경의 배치는 가깝지만, 신부가 찰리보다 카메라 정면으로 걸으며 이전 장면의 차량도 다른 차종으로 바뀌었다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 얼굴과 앞발은 화면 오른쪽의 찰리 쪽을 향한다. 정확한 눈동자는 역광으로 보이지 않지만 접근 대상은 찰리로 읽힌다. 찰리는 몸통이 거의 카메라 정면을 향해 신부와의 시선 교환은 약하다. 차량 전면과 켜진 전조등은 신부 뒤에서 카메라 쪽을 향한다.",
        "built_space": "왼쪽에 이전 장면과 같은 대형 상자형 화물차 한 대가 있고, 전면의 위아래 좌우 등 네 개 중 일부가 신부에게 가려진다. 양쪽 낡은 건물, 셔터, 왼쪽 차양, 가까운 좌우 가로등 각 한 개, 전선과 왼쪽 전경 맨홀이 유지된다. 신부는 왼쪽 전경, 찰리는 오른쪽 중경에 있어, 요구된 왼쪽 전경에서 오른쪽 중경의 신부로 이어지는 접근 배치는 반대다. 젖은 노면의 광원 반사는 가능한 위치에 있다.",
        "entities": "신부 한 명과 찰리 한 명만 보인다. 신부는 짧은 머리의 남성으로 검은 성직자 옷과 흰 칼라를 착용하며, 얼굴이 어두워 한국인 60대의 세부 인상과 주름은 확인할 수 없다. 찰리의 낡은 모자와 코트, 모래색 장갑판, 흰 마스크형 얼굴, 긴 팔이 참조와 대체로 맞는다. 새는 없다. 차량은 문구상의 밴보다는 이전 장면에 실제로 나온 화물차 외관을 따른다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부는 앞발 뒤꿈치가 노면에 닿는 보행 순간이며 반대쪽 발도 노면 가까이 있어 정상적인 체중 이동으로 설명된다. 찰리는 두 발로 도로에 서 있고 차량은 바퀴로 지지된다. 공중에 지지 없이 떠 있는 신체나 소품은 없다."
       },
       {
        "label": "A",
        "direction": "신부의 얼굴, 몸통과 내딛는 발은 거의 카메라 정면을 향한다. 찰리는 신부의 왼쪽 전경에 있으므로 이 보행 방향은 찰리를 향한 대각선 접근으로 명확하게 이어지지 않는다. 찰리의 얼굴은 약간 오른쪽으로 돌아 신부 방향에 더 가깝다. 차량 전조등은 신부 뒤에서 카메라 쪽을 비춘다.",
        "built_space": "찰리가 왼쪽 전경, 신부가 오른쪽 중경에 있고 그 뒤에 비스듬히 정차한 밴 한 대가 있어 요구된 도로 연결 배치에 가깝다. 밴에는 밝은 전조등 두 개가 보인다. 양쪽 건물과 셔터, 왼쪽 차양, 가까운 좌우 가로등 각 한 개, 전선과 맨홀은 이전 장소를 대체로 유지한다. 그러나 이전 장면의 높은 화물차가 낮은 승합형 밴으로 교체되어 차량 외관 연속성이 깨진다. 노면 반사는 광원과 지면 관계에 맞는다.",
        "entities": "신부 한 명과 찰리 한 명이 보인다. 신부의 검은 옷과 작은 흰 성직자 칼라는 보이지만 얼굴은 실루엣에 가려 나이와 참조 얼굴의 세부를 확인할 수 없다. 찰리는 모자, 낡은 코트, 모래색 장갑판, 흰 얼굴과 육중하고 긴 팔을 유지한다. 새는 없다. 차량은 밴이라는 명칭에는 맞지만 이전 장면의 실제 차량과는 다르다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부는 앞쪽 발을 노면에 디디고 뒤쪽 발을 옮기는 보행 자세로, 지지와 체중 이동이 성립한다. 찰리의 두 발은 도로에 닿아 있고 밴도 바퀴로 지지된다. 떠 있는 물체나 불가능한 관절 배치는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.946,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.696,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷 레퍼런스의 차량 외관(트럭)을 그대로 유지해야 하는 잠금(Lock) 지시를 무시하고 다른 형태의 차량(승합차)을 생성함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 696
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "이전 샷의 트럭 외관과 배경을 완벽히 유지했으며, 헤드라이트를 등지고 찰리를 향해 걷는 신부의 실루엣 동선을 정확하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 696,
    "verdict_ko": "레퍼런스의 차량 외관을 유지하라는 지침을 위반하여 차량이 승합차로 바뀌었으며, 찰리를 향해 걷는 지시사항을 따르지 않았습니다.  ★위반: [gemini-pro] 이전 샷 레퍼런스의 차량 외관(트럭)을 그대로 유지해야 하는 잠금(Lock) 지시를 무시하고 다른 형태의 차량(승합차)을 생성함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12_sel.png",
    "asset_id": "238e6601-3959-4022-a815-97a07f670c77",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-1f3c-7947-8da2-fc59ec17c5c3",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S27sh12"
  },
  "lane_policy": "ab_select_bypass:prev",
  "staged_characters_added": [
   "C06"
  ]
 },
 "S27sh18::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:10:24.696238+00:00",
  "fingerprint": "180355e152bd8f1f0a50f4b756b24d5d5b55548f243fd0809cc704b1b28e9890",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S27sh18_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S27sh18_sel.png",
  "source_sha256": "3a953baa74c487780bc64bba54d97cd856188ab588f0956e2b0cc80e2cb9c027",
  "file": "S27sh18_cine.png",
  "staged_sha256": "be8c8d5ff5c9c590c27c5e3a6dd9e50060c6f7f515ab73729239b8ca01b51bb2",
  "latency_ms": 10297
 },
 "S27sh21::signage": {
  "fp": "3e5eeffaf2a5c370",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S27sh21": {
  "input_fingerprint": "fbfc4f057284d965",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 골목 모퉁이에 숨어 찰리가 타는 자동차를 예의주시하는 구도환의 날카로운 눈매 클로즈업.\n\nLOCATION (lock): At a dark alley corner overlooking the stopped van on the refugee settlement road. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Alley corner (Conceals 구도환 as he watches the vehicle) — The near edge is visible along the left side of the image; used as Provides a restrained foreground occlusion that explains his hidden position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the alley's stated darkness with enough neutral tonal separation to read 구도환's eyes, without transferring the road's headlight effect onto his face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van remains stopped with its headlights on, and Charlie is boarding in the old coat and hat. The vehicle has a crude halogen cross ornament on its roof. 구도환: He is observing the boarding from outside the van.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 골목 모퉁이에 숨어 찰리가 타는 자동차를 예의주시하는 구도환의 날카로운 눈매 클로즈업.\n\nLOCATION (lock): At a dark alley corner overlooking the stopped van on the refugee settlement road. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Alley corner (Conceals 구도환 as he watches the vehicle) — The near edge is visible along the left side of the image; used as Provides a restrained foreground occlusion that explains his hidden position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the alley's stated darkness with enough neutral tonal separation to read 구도환's eyes, without transferring the road's headlight effect onto his face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van remains stopped with its headlights on, and Charlie is boarding in the old coat and hat. The vehicle has a crude halogen cross ornament on its roof. 구도환: He is observing the boarding from outside the van.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 골목 모퉁이에 숨어 찰리가 타는 자동차를 예의주시하는 구도환의 날카로운 눈매 클로즈업.\n\nLOCATION (lock): At a dark alley corner overlooking the stopped van on the refugee settlement road. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Alley corner (Conceals 구도환 as he watches the vehicle) — The near edge is visible along the left side of the image; used as Provides a restrained foreground occlusion that explains his hidden position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the alley's stated darkness with enough neutral tonal separation to read 구도환's eyes, without transferring the road's headlight effect onto his face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van remains stopped with its headlights on, and Charlie is boarding in the old coat and hat. The vehicle has a crude halogen cross ornament on its roof. 구도환: He is observing the boarding from outside the van.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "우측 배경의 트럭 쪽을 향해 시선이 고정됨.",
    "built_space": "좌측 전경에 벽돌 벽이 위치하며 배경의 거리에는 대형 트럭이 주차됨.",
    "entities": "구도환의 얼굴은 일치하나 지시에 없는 비니를 씀. 십자가 승합차와 찰리가 없음.",
    "hard_violations": [
     "[gemini-pro] 지시에 없는 물건(비니) 추가",
     "[gemini-pro] 프롬프트에 명시된 십자가 승합차 누락 및 임의의 트럭 배치"
    ],
    "physics": "땅에 서서 벽에 기댄 자세가 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "우측 배경의 승합차와 탑승하는 인물들을 날카롭게 주시함.",
    "built_space": "좌측에 은폐용 벽이 있고 배경에는 헤드라이트를 켠 승합차가 위치함.",
    "entities": "구도환의 인상과 코트가 레퍼런스와 정확히 일치함. 십자가가 달린 승합차와 탑승하는 인물들이 있음.",
    "hard_violations": [
     "[gpt-high] 구도환 이외의 인물이 차량 옆에 최소 두 명 등장하여 명시적인 단독 등장 제한을 위반한다."
    ],
    "physics": "벽 뒤에 숨어 내다보는 자세와 배경 인물들의 움직임에 어색함이 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시에 없는 비니를 착용했고 필수 요소인 십자가 승합차와 찰리의 탑승 장면을 누락함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "캐릭터의 외형과 의상을 정확히 재현했으며 십자가 승합차와 탑승 장면 등 프롬프트의 지시를 충실히 이행함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "우측 배경의 트럭 쪽을 향해 시선이 고정됨.",
        "built_space": "좌측 전경에 벽돌 벽이 위치하며 배경의 거리에는 대형 트럭이 주차됨.",
        "entities": "구도환의 얼굴은 일치하나 지시에 없는 비니를 씀. 십자가 승합차와 찰리가 없음.",
        "hard_violations": [
         "지시에 없는 물건(비니) 추가",
         "프롬프트에 명시된 십자가 승합차 누락 및 임의의 트럭 배치"
        ],
        "physics": "땅에 서서 벽에 기댄 자세가 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "우측 배경의 승합차와 탑승하는 인물들을 날카롭게 주시함.",
        "built_space": "좌측에 은폐용 벽이 있고 배경에는 헤드라이트를 켠 승합차가 위치함.",
        "entities": "구도환의 인상과 코트가 레퍼런스와 정확히 일치함. 십자가가 달린 승합차와 탑승하는 인물들이 있음.",
        "hard_violations": [],
        "physics": "벽 뒤에 숨어 내다보는 자세와 배경 인물들의 움직임에 어색함이 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시에 없는 비니를 착용했고 필수 요소인 십자가 승합차와 찰리의 탑승 장면을 누락함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "캐릭터의 외형과 의상을 정확히 재현했으며 십자가 승합차와 탑승 장면 등 프롬프트의 지시를 충실히 이행함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "우측 배경의 트럭 쪽을 향해 시선이 고정됨.",
        "built_space": "좌측 전경에 벽돌 벽이 위치하며 배경의 거리에는 대형 트럭이 주차됨.",
        "entities": "구도환의 얼굴은 일치하나 지시에 없는 비니를 씀. 십자가 승합차와 찰리가 없음.",
        "hard_violations": [
         "지시에 없는 물건(비니) 추가",
         "프롬프트에 명시된 십자가 승합차 누락 및 임의의 트럭 배치"
        ],
        "physics": "땅에 서서 벽에 기댄 자세가 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "우측 배경의 승합차와 탑승하는 인물들을 날카롭게 주시함.",
        "built_space": "좌측에 은폐용 벽이 있고 배경에는 헤드라이트를 켠 승합차가 위치함.",
        "entities": "구도환의 인상과 코트가 레퍼런스와 정확히 일치함. 십자가가 달린 승합차와 탑승하는 인물들이 있음.",
        "hard_violations": [],
        "physics": "벽 뒤에 숨어 내다보는 자세와 배경 인물들의 움직임에 어색함이 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴·머리·외투와 차량의 십자가는 비교적 충실하지만, 금지된 추가 인물들이 등장하고 구도환의 시선이 뒤쪽 차량을 향하지 않는다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "단독 인물의 어두운 모퉁이 클로즈업은 더 적합하지만, 시선이 차량과 반대쪽이며 참조에 없는 비니와 달라진 외투, 차량 지붕의 십자가 누락이 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "구도환의 두 눈은 화면 왼쪽 위, 카메라 가까운 화면 밖을 향한다. 감시해야 할 밴은 그의 오른쪽 뒤에 있어 시선이 차량이나 승차 지점에 닿지 않는다. 밴 전면과 켜진 전조등은 카메라 쪽을 향한다. 승차문 근처 인물 하나는 차량 쪽으로 몸을 돌리고 있다.",
        "built_space": "왼쪽 전경의 두꺼운 벽 모서리 하나가 화면 약 3분의 1을 차지하며 얼굴 왼쪽과 몸 일부를 가린다. 오른쪽 뒤에는 밴 한 대, 지붕 십자가 하나, 전면의 밝은 전조등 두 곳과 열린 측면 승차구가 보인다. 낡은 저층 건물, 가공 전선, 젖은 도로와 여러 가로등은 이전 장소의 분위기에 부합하지만 같은 고정 설비 배치인지는 확인하기 어렵다. 차량과 승차 인물들을 상당히 보여주는 구도로, 눈매 중심의 긴밀한 클로즈업보다는 배경 설명 비중이 크다.",
        "entities": "전경에는 검은 머리의 중년 한국인 남성으로 보이는 구도환이 있으며, 얼굴과 털 안감의 낡은 녹색 외투, 검은 목 부분 의상이 인물 참조와 비교적 잘 맞는다. 뒤에는 최소 두 명의 별도 인물이 명확히 보이고 승차구에도 사람 얼굴처럼 보이는 형상이 있어, 구도환만 등장해야 한다는 제한을 어긴다. 밴과 발광 십자가는 존재한다. 배경 인물의 정확한 나이·민족성 및 찰리의 모자 착용 여부는 흐림 때문에 확인할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "구도환 이외의 인물이 차량 옆에 최소 두 명 등장하여 명시적인 단독 등장 제한을 위반한다."
        ],
        "physics": "구도환은 벽 뒤에서 상체를 조금 내민 자세이며 하체는 프레임 밖이다. 공중에 떠 있다고 볼 징후는 없다. 배경 인물들의 다리는 도로까지 이어지고, 승차 중인 인물의 손은 문 쪽에 닿아 있다. 밴은 바퀴로 도로에 놓이고 십자가는 지붕 위에 설치되어 있다. 젖은 노면의 광원 반사도 가능한 배치다."
       },
       {
        "label": "B",
        "direction": "구도환은 눈을 화면 왼쪽 위로 돌려 화면 밖을 응시한다. 밴은 오른쪽 뒤에 있으므로 그 차량을 예의주시하는 시선은 아니다. 차량 전면과 두 전조등은 카메라 쪽을 향한다. 무기나 손에 든 지시 물체는 없다.",
        "built_space": "왼쪽에는 벽돌 모퉁이 하나가 있고 구도환은 그 오른쪽 가장자리에 몸을 붙이고 있다. 얼굴과 어깨를 크게 담고 오른쪽 먼 배경에 차량 한 대를 작게 두어 클로즈업의 우선순위는 더 잘 지킨다. 배경에는 좁은 도로, 낡은 건물과 여러 조명이 보인다. 젖은 도로와 야간 분위기는 이어지지만, 참조의 회벽 중심 거리와 동일한 고정 구조인지 입증할 세부는 부족하다. 차량 지붕에는 요구된 십자가가 보이지 않는다.",
        "entities": "보이는 사람은 중년 한국인 남성으로 읽히는 구도환 한 명뿐이다. 얼굴은 참조와 대체로 유사하지만, 참조에 없던 검은 비니가 검은 머리를 덮고 외투에는 참조의 뚜렷한 털 안감 깃이 보이지 않는다. 먼 배경에는 전조등 두 개가 켜진 밴 형태의 차량이 있다. 찰리나 다른 인물은 보이지 않으며, 이 클로즈업에서 그들을 추가하지 않은 것은 적절하다. 눈은 정상적인 홍채와 동공을 지니고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "구도환은 벽 모서리에 어깨와 머리를 가까이 둔 자연스러운 기립 상체 자세다. 발은 촬영 범위 밖이므로 지면 접촉은 확인할 수 없지만 부유나 불가능한 자세는 보이지 않는다. 비니는 머리에 씌워져 있고 외투는 어깨에 걸쳐져 있다. 배경 차량은 도로 위에 놓여 있으며 지지 없는 물체나 불가능한 반사는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴·머리·외투와 차량의 십자가는 비교적 충실하지만, 금지된 추가 인물들이 등장하고 구도환의 시선이 뒤쪽 차량을 향하지 않는다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "단독 인물의 어두운 모퉁이 클로즈업은 더 적합하지만, 시선이 차량과 반대쪽이며 참조에 없는 비니와 달라진 외투, 차량 지붕의 십자가 누락이 남는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "구도환의 두 눈은 화면 왼쪽 위, 카메라 가까운 화면 밖을 향한다. 감시해야 할 밴은 그의 오른쪽 뒤에 있어 시선이 차량이나 승차 지점에 닿지 않는다. 밴 전면과 켜진 전조등은 카메라 쪽을 향한다. 승차문 근처 인물 하나는 차량 쪽으로 몸을 돌리고 있다.",
        "built_space": "왼쪽 전경의 두꺼운 벽 모서리 하나가 화면 약 3분의 1을 차지하며 얼굴 왼쪽과 몸 일부를 가린다. 오른쪽 뒤에는 밴 한 대, 지붕 십자가 하나, 전면의 밝은 전조등 두 곳과 열린 측면 승차구가 보인다. 낡은 저층 건물, 가공 전선, 젖은 도로와 여러 가로등은 이전 장소의 분위기에 부합하지만 같은 고정 설비 배치인지는 확인하기 어렵다. 차량과 승차 인물들을 상당히 보여주는 구도로, 눈매 중심의 긴밀한 클로즈업보다는 배경 설명 비중이 크다.",
        "entities": "전경에는 검은 머리의 중년 한국인 남성으로 보이는 구도환이 있으며, 얼굴과 털 안감의 낡은 녹색 외투, 검은 목 부분 의상이 인물 참조와 비교적 잘 맞는다. 뒤에는 최소 두 명의 별도 인물이 명확히 보이고 승차구에도 사람 얼굴처럼 보이는 형상이 있어, 구도환만 등장해야 한다는 제한을 어긴다. 밴과 발광 십자가는 존재한다. 배경 인물의 정확한 나이·민족성 및 찰리의 모자 착용 여부는 흐림 때문에 확인할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "구도환 이외의 인물이 차량 옆에 최소 두 명 등장하여 명시적인 단독 등장 제한을 위반한다."
        ],
        "physics": "구도환은 벽 뒤에서 상체를 조금 내민 자세이며 하체는 프레임 밖이다. 공중에 떠 있다고 볼 징후는 없다. 배경 인물들의 다리는 도로까지 이어지고, 승차 중인 인물의 손은 문 쪽에 닿아 있다. 밴은 바퀴로 도로에 놓이고 십자가는 지붕 위에 설치되어 있다. 젖은 노면의 광원 반사도 가능한 배치다."
       },
       {
        "label": "A",
        "direction": "구도환은 눈을 화면 왼쪽 위로 돌려 화면 밖을 응시한다. 밴은 오른쪽 뒤에 있으므로 그 차량을 예의주시하는 시선은 아니다. 차량 전면과 두 전조등은 카메라 쪽을 향한다. 무기나 손에 든 지시 물체는 없다.",
        "built_space": "왼쪽에는 벽돌 모퉁이 하나가 있고 구도환은 그 오른쪽 가장자리에 몸을 붙이고 있다. 얼굴과 어깨를 크게 담고 오른쪽 먼 배경에 차량 한 대를 작게 두어 클로즈업의 우선순위는 더 잘 지킨다. 배경에는 좁은 도로, 낡은 건물과 여러 조명이 보인다. 젖은 도로와 야간 분위기는 이어지지만, 참조의 회벽 중심 거리와 동일한 고정 구조인지 입증할 세부는 부족하다. 차량 지붕에는 요구된 십자가가 보이지 않는다.",
        "entities": "보이는 사람은 중년 한국인 남성으로 읽히는 구도환 한 명뿐이다. 얼굴은 참조와 대체로 유사하지만, 참조에 없던 검은 비니가 검은 머리를 덮고 외투에는 참조의 뚜렷한 털 안감 깃이 보이지 않는다. 먼 배경에는 전조등 두 개가 켜진 밴 형태의 차량이 있다. 찰리나 다른 인물은 보이지 않으며, 이 클로즈업에서 그들을 추가하지 않은 것은 적절하다. 눈은 정상적인 홍채와 동공을 지니고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "구도환은 벽 모서리에 어깨와 머리를 가까이 둔 자연스러운 기립 상체 자세다. 발은 촬영 범위 밖이므로 지면 접촉은 확인할 수 없지만 부유나 불가능한 자세는 보이지 않는다. 비니는 머리에 씌워져 있고 외투는 어깨에 걸쳐져 있다. 배경 차량은 도로 위에 놓여 있으며 지지 없는 물체나 불가능한 반사는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.6
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.35
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시에 없는 물건(비니) 추가",
     "[gemini-pro] 프롬프트에 명시된 십자가 승합차 누락 및 임의의 트럭 배치"
    ],
    "B": [
     "[gpt-high] 구도환 이외의 인물이 차량 옆에 최소 두 명 등장하여 명시적인 단독 등장 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1321,
   "B": 1350
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "지시에 없는 비니를 착용했고 필수 요소인 십자가 승합차와 찰리의 탑승 장면을 누락함.  ★위반: [gemini-pro] 지시에 없는 물건(비니) 추가 / [gemini-pro] 프롬프트에 명시된 십자가 승합차 누락 및 임의의 트럭 배치"
   },
   {
    "label": "B",
    "score": 1350,
    "verdict_ko": "캐릭터의 외형과 의상을 정확히 재현했으며 십자가 승합차와 탑승 장면 등 프롬프트의 지시를 충실히 이행함.  ★위반: [gpt-high] 구도환 이외의 인물이 차량 옆에 최소 두 명 등장하여 명시적인 단독 등장 제한을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh18_sel.png",
    "asset_id": "d5de13f4-c83f-4719-adf9-3bc295063278",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1023014>",
    "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-20f9-7346-8072-f6e2c654d682",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S27sh18"
  }
 },
 "S27sh21::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:11:45.032316+00:00",
  "fingerprint": "35cc722fcd07e55a343a9baf85a3e1f078096d2912f199bcb697be79a5efefc7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S27sh21_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S27sh21_sel.png",
  "source_sha256": "a2f75129a261b0aaea8aec91a70e146f3c5bdfd0f689e635d5464e2794aefb4f",
  "file": "S27sh21_cine.png",
  "staged_sha256": "c7404cfb85cc581358b7c076412d4f567dd4aef14ca9c5380d53b7d3e0661e09",
  "latency_ms": 10130
 },
 "S28sh2::signage": {
  "fp": "9f143c5992eeaeae",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::22ae2a4d0af72fb5": {
  "subjects": [],
  "subject_text": "신부의 밴 내부\n운전석 뒤로 뒷좌석과 적재 공간이 이어지는 밴 실내. 물품 상자들이 듬성듬성 놓여 있고 상자 사이에 빈 공간이 남아 있다.",
  "identity": "canonical",
  "scope_id": "L179",
  "scope_role": "location_interior",
  "scope_sha": "b707d6881f7a896b"
 },
 "S28sh2::bgfirst_bg": {
  "input_fingerprint": "03f810a54ad37b98",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh2__bgfirst_bg.png",
  "asset_id": "bc133181-b478-4d81-83c2-7aa1d682c287",
  "input_asset_ids": [
   "999f65d7-3a65-487a-8bd7-c6b9ebc6b8a4",
   "f6306063-3d19-4618-946d-060af53e1589"
  ]
 },
 "S28sh2": {
  "input_fingerprint": "253aaaa4acf7c795",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Scattered cargo boxes occupy the van's rear seating area, and a crude halogen cross is mounted on the roof. Charlie remains concealed among the boxes in his old coat and hat. 현우: He is concealed among the rear cargo boxes, with facial bruises and an injured leg; his outer garment remains removed. The contact card remains concealed in his shoe. 앰버: She is concealed among the rear boxes, still exhausted, with the outer garment used as a mouth covering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Scattered cargo boxes occupy the van's rear seating area, and a crude halogen cross is mounted on the roof. Charlie remains concealed among the boxes in his old coat and hat. 현우: He is concealed among the rear cargo boxes, with facial bruises and an injured leg; his outer garment remains removed. The contact card remains concealed in his shoe. 앰버: She is concealed among the rear boxes, still exhausted, with the outer garment used as a mouth covering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Scattered cargo boxes occupy the van's rear seating area, and a crude halogen cross is mounted on the roof. Charlie remains concealed among the boxes in his old coat and hat. 현우: He is concealed among the rear cargo boxes, with facial bruises and an injured leg; his outer garment remains removed. The contact card remains concealed in his shoe. 앰버: She is concealed among the rear boxes, still exhausted, with the outer garment used as a mouth covering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh2__bgfirst_bg.png",
     "asset_id": "bc133181-b478-4d81-83c2-7aa1d682c287",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S28sh2.png",
     "asset_id": "999f65d7-3a65-487a-8bd7-c6b9ebc6b8a4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
     "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "세 인물 모두 시선이 바닥을 향하고 있음.",
    "built_space": "밴 내부에서 앞유리를 바라보는 구도. 상자들이 주변에 배치되어 있으나 인물들이 위치한 중앙 공간은 넓게 비워져 있음.",
    "entities": "현우(참조 이미지의 겉옷을 그대로 착용함), 앰버(무릎에 얼굴을 묻고 있음), 찰리(코트와 모자를 착용하고 로봇 외형이 일치함).",
    "hard_violations": [],
    "physics": "세 인물 모두 밴의 바닥에 안정적으로 체중을 싣고 앉아 있음."
   },
   {
    "label": "B",
    "direction": "현우는 앞의 상자를, 앰버는 정면을, 찰리는 현우 쪽을 바라봄.",
    "built_space": "밴 내부. 종이상자들이 빽빽하게 쌓여 인물들이 몸을 숨긴 비좁은 틈새 공간을 형성함.",
    "entities": "현우(겉옷을 탈의하고 다리의 멍과 부상이 보임), 앰버(옷으로 입을 가림), 찰리(코트와 모자 착용, 외형 일치).",
    "hard_violations": [
     "[gpt-high] 왼쪽 종이상자에 '취급주의' 등 판독 가능한 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
    ],
    "physics": "바닥에 웅크린 채 상자에 기대어 자연스럽게 자세를 유지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 겉옷 탈의와 다리 부상, 앰버의 입을 가린 상태를 정확히 반영했으며, 상자 틈새에 욱여넣어진 비좁은 구도를 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 여전히 겉옷을 입고 있어 지시사항을 위반했으며, 인물들이 상자 사이가 아닌 텅 빈 중앙 공간에 나란히 앉아 있어 긴장감이 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "세 인물 모두 시선이 바닥을 향하고 있음.",
        "built_space": "밴 내부에서 앞유리를 바라보는 구도. 상자들이 주변에 배치되어 있으나 인물들이 위치한 중앙 공간은 넓게 비워져 있음.",
        "entities": "현우(참조 이미지의 겉옷을 그대로 착용함), 앰버(무릎에 얼굴을 묻고 있음), 찰리(코트와 모자를 착용하고 로봇 외형이 일치함).",
        "hard_violations": [],
        "physics": "세 인물 모두 밴의 바닥에 안정적으로 체중을 싣고 앉아 있음."
       },
       {
        "label": "B",
        "direction": "현우는 앞의 상자를, 앰버는 정면을, 찰리는 현우 쪽을 바라봄.",
        "built_space": "밴 내부. 종이상자들이 빽빽하게 쌓여 인물들이 몸을 숨긴 비좁은 틈새 공간을 형성함.",
        "entities": "현우(겉옷을 탈의하고 다리의 멍과 부상이 보임), 앰버(옷으로 입을 가림), 찰리(코트와 모자 착용, 외형 일치).",
        "hard_violations": [],
        "physics": "바닥에 웅크린 채 상자에 기대어 자연스럽게 자세를 유지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 겉옷 탈의와 다리 부상, 앰버의 입을 가린 상태를 정확히 반영했으며, 상자 틈새에 욱여넣어진 비좁은 구도를 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 여전히 겉옷을 입고 있어 지시사항을 위반했으며, 인물들이 상자 사이가 아닌 텅 빈 중앙 공간에 나란히 앉아 있어 긴장감이 떨어집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "세 인물 모두 시선이 바닥을 향하고 있음.",
        "built_space": "밴 내부에서 앞유리를 바라보는 구도. 상자들이 주변에 배치되어 있으나 인물들이 위치한 중앙 공간은 넓게 비워져 있음.",
        "entities": "현우(참조 이미지의 겉옷을 그대로 착용함), 앰버(무릎에 얼굴을 묻고 있음), 찰리(코트와 모자를 착용하고 로봇 외형이 일치함).",
        "hard_violations": [],
        "physics": "세 인물 모두 밴의 바닥에 안정적으로 체중을 싣고 앉아 있음."
       },
       {
        "label": "B",
        "direction": "현우는 앞의 상자를, 앰버는 정면을, 찰리는 현우 쪽을 바라봄.",
        "built_space": "밴 내부. 종이상자들이 빽빽하게 쌓여 인물들이 몸을 숨긴 비좁은 틈새 공간을 형성함.",
        "entities": "현우(겉옷을 탈의하고 다리의 멍과 부상이 보임), 앰버(옷으로 입을 가림), 찰리(코트와 모자 착용, 외형 일치).",
        "hard_violations": [],
        "physics": "바닥에 웅크린 채 상자에 기대어 자연스럽게 자세를 유지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "읽을 수 있는 상자 문구가 명시적 금지 조건을 위반하며, 앰버와 찰리의 전신이 가려지고 카메라도 지정된 앞 조수석 쪽 시점과 반대다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "세 인물의 굳은 웅크림과 앰버의 입 가림을 더 충실히 구현하고 판독 가능한 글자도 없지만, 카메라 방향과 찰리의 전신 가시성은 요구에 미달한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽의 찰리 쪽을 보고, 앰버는 카메라 쪽을 정면으로 바라본다. 찰리는 왼쪽 아래의 일행과 상자 쪽으로 고개를 숙인다. 지정된 응시 표적은 없지만 앰버의 정면 응시는 숨죽인 은신 연기를 약화한다. 무기나 이동하는 물체는 없다.",
        "built_space": "앞좌석 등받이와 머리받침이 각각 두 개, 운전대 한 개, 중앙 대시보드 한 곳, 실내 거울 한 개가 보인다. 양옆 창과 낡은 내장재는 장소 참조와 유사하다. 그러나 카메라는 뒤쪽에서 앞유리를 향하므로 앞 조수석 쪽에서 후방 공간을 대각선으로 보는 구도가 아니다. 현우는 왼쪽, 앰버는 중앙, 찰리는 오른쪽 상자 틈에 있으나 중앙 앞 상자가 앰버의 하체를, 전경 상자가 찰리의 발을 가린다. 참조의 뒤쪽 벤치 좌석은 확인되지 않는다.",
        "entities": "등장 개체는 현우, 앰버, 찰리 세 명뿐이다. 현우는 앳된 동아시아계 남성으로 검은 머리와 얼굴 멍, 다리 상처가 보이고 외투를 벗었지만, 민소매와 반바지는 참조의 셔츠·긴 바지와 다르다. 앰버는 금발의 어린 여자아이이며 어두운 겉옷으로 입을 가렸다. 얼굴과 체격은 대체로 맞지만 하체와 원래 의상은 가려졌다. 찰리는 모자와 낡은 코트, 베이지 장갑판, 흰 기계형 얼굴과 긴 팔이 참조에 부합한다. 낡은 종이상자는 충분하지만 왼쪽 상자의 '취급주의'가 읽힌다. 신발 속 카드와 지붕 위 십자가는 보이지 않아 확인할 수 없다.",
        "hard_violations": [
         "왼쪽 종이상자에 '취급주의' 등 판독 가능한 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
        ],
        "physics": "현우는 다리를 접어 바닥에 앉고 신발을 바닥에 대며 양손으로 옆 상자를 짚는다. 앰버는 상자 사이에 낮게 앉아 있으나 하체와 실제 지지 접점은 가려져 있다. 찰리는 접힌 다리 위로 몸을 낮추고 한 손을 상자 위에 얹는다. 두 인물의 가려진 접점을 부유로 판단할 근거는 없다. 상자들은 바닥이나 다른 상자에 놓여 있으며, 명백한 무지지 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 끌어안은 무릎 쪽으로 시선을 낮추고, 찰리도 접은 팔과 무릎 쪽으로 고개를 숙인다. 세 인물 모두 카메라를 응시하지 않아 긴장 속에서 몸을 움츠린 순간으로 읽힌다. 겨누는 물체나 이동 방향을 확인해야 할 동작은 없다.",
        "built_space": "앞좌석 두 개와 머리받침 두 개, 운전대 한 개, 중앙 대시보드와 실내 거울 한 개가 보인다. 창틀, 안전벨트 고정부, 회색 내장재와 밤의 외부 풍경은 참조 장소에 가깝다. 다만 이 역시 뒤에서 앞좌석을 보는 방향으로, 지정된 앞 조수석 쪽 시점은 아니다. 현우·앰버·찰리가 왼쪽부터 각기 다른 상자 틈을 차지하며 상자 윗면과 옆면이 보인다. 현우와 앰버의 웅크린 몸은 발까지 대부분 읽히지만 찰리의 하체와 발은 앞 상자에 가려진다. 참조의 뒤쪽 벤치 좌석은 확인되지 않는다.",
        "entities": "세 인물 외에 추가 인물은 없다. 현우의 젊은 동아시아계 남성 외형, 검은 머리, 회색 셔츠와 녹색 바지는 참조에 가깝고 얼굴 멍도 보인다. 별도 외투는 보이지 않으며 다리 부상은 바지 때문에 확인하기 어렵다. 앰버는 금발의 어린 여자아이로 남색 상의, 갈색 작업복과 부츠가 참조에 가깝고, 천을 입에 댄 채 무릎을 끌어안고 있다. 얼굴이 숙여져 세부 동일성은 제한적으로만 확인된다. 찰리의 육중한 체격, 긴 장갑 팔, 베이지 장갑판, 흰 얼굴, 낡은 코트와 모자는 참조와 부합한다. 상자에는 판독 가능한 글자가 보이지 않는다. 숨긴 카드와 외부 지붕 십자가는 확인 범위 밖이다.",
        "hard_violations": [],
        "physics": "현우와 앰버는 바닥에 엉덩이를 낮추고 다리를 접었으며 부츠가 바닥에 닿는다. 두 사람의 손과 팔은 정강이나 무릎을 감싸 실제 접촉을 이룬다. 찰리는 다리를 접고 팔을 무릎 위에 포갠 자세로 보이나 발과 좌면 접점은 상자에 가려져 있다. 부유를 나타내는 틈이나 불가능한 자세는 보이지 않는다. 상자는 바닥과 아래쪽 상자에 지지되어 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "읽을 수 있는 상자 문구가 명시적 금지 조건을 위반하며, 앰버와 찰리의 전신이 가려지고 카메라도 지정된 앞 조수석 쪽 시점과 반대다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "세 인물의 굳은 웅크림과 앰버의 입 가림을 더 충실히 구현하고 판독 가능한 글자도 없지만, 카메라 방향과 찰리의 전신 가시성은 요구에 미달한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽의 찰리 쪽을 보고, 앰버는 카메라 쪽을 정면으로 바라본다. 찰리는 왼쪽 아래의 일행과 상자 쪽으로 고개를 숙인다. 지정된 응시 표적은 없지만 앰버의 정면 응시는 숨죽인 은신 연기를 약화한다. 무기나 이동하는 물체는 없다.",
        "built_space": "앞좌석 등받이와 머리받침이 각각 두 개, 운전대 한 개, 중앙 대시보드 한 곳, 실내 거울 한 개가 보인다. 양옆 창과 낡은 내장재는 장소 참조와 유사하다. 그러나 카메라는 뒤쪽에서 앞유리를 향하므로 앞 조수석 쪽에서 후방 공간을 대각선으로 보는 구도가 아니다. 현우는 왼쪽, 앰버는 중앙, 찰리는 오른쪽 상자 틈에 있으나 중앙 앞 상자가 앰버의 하체를, 전경 상자가 찰리의 발을 가린다. 참조의 뒤쪽 벤치 좌석은 확인되지 않는다.",
        "entities": "등장 개체는 현우, 앰버, 찰리 세 명뿐이다. 현우는 앳된 동아시아계 남성으로 검은 머리와 얼굴 멍, 다리 상처가 보이고 외투를 벗었지만, 민소매와 반바지는 참조의 셔츠·긴 바지와 다르다. 앰버는 금발의 어린 여자아이이며 어두운 겉옷으로 입을 가렸다. 얼굴과 체격은 대체로 맞지만 하체와 원래 의상은 가려졌다. 찰리는 모자와 낡은 코트, 베이지 장갑판, 흰 기계형 얼굴과 긴 팔이 참조에 부합한다. 낡은 종이상자는 충분하지만 왼쪽 상자의 '취급주의'가 읽힌다. 신발 속 카드와 지붕 위 십자가는 보이지 않아 확인할 수 없다.",
        "hard_violations": [
         "왼쪽 종이상자에 '취급주의' 등 판독 가능한 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
        ],
        "physics": "현우는 다리를 접어 바닥에 앉고 신발을 바닥에 대며 양손으로 옆 상자를 짚는다. 앰버는 상자 사이에 낮게 앉아 있으나 하체와 실제 지지 접점은 가려져 있다. 찰리는 접힌 다리 위로 몸을 낮추고 한 손을 상자 위에 얹는다. 두 인물의 가려진 접점을 부유로 판단할 근거는 없다. 상자들은 바닥이나 다른 상자에 놓여 있으며, 명백한 무지지 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우와 앰버는 끌어안은 무릎 쪽으로 시선을 낮추고, 찰리도 접은 팔과 무릎 쪽으로 고개를 숙인다. 세 인물 모두 카메라를 응시하지 않아 긴장 속에서 몸을 움츠린 순간으로 읽힌다. 겨누는 물체나 이동 방향을 확인해야 할 동작은 없다.",
        "built_space": "앞좌석 두 개와 머리받침 두 개, 운전대 한 개, 중앙 대시보드와 실내 거울 한 개가 보인다. 창틀, 안전벨트 고정부, 회색 내장재와 밤의 외부 풍경은 참조 장소에 가깝다. 다만 이 역시 뒤에서 앞좌석을 보는 방향으로, 지정된 앞 조수석 쪽 시점은 아니다. 현우·앰버·찰리가 왼쪽부터 각기 다른 상자 틈을 차지하며 상자 윗면과 옆면이 보인다. 현우와 앰버의 웅크린 몸은 발까지 대부분 읽히지만 찰리의 하체와 발은 앞 상자에 가려진다. 참조의 뒤쪽 벤치 좌석은 확인되지 않는다.",
        "entities": "세 인물 외에 추가 인물은 없다. 현우의 젊은 동아시아계 남성 외형, 검은 머리, 회색 셔츠와 녹색 바지는 참조에 가깝고 얼굴 멍도 보인다. 별도 외투는 보이지 않으며 다리 부상은 바지 때문에 확인하기 어렵다. 앰버는 금발의 어린 여자아이로 남색 상의, 갈색 작업복과 부츠가 참조에 가깝고, 천을 입에 댄 채 무릎을 끌어안고 있다. 얼굴이 숙여져 세부 동일성은 제한적으로만 확인된다. 찰리의 육중한 체격, 긴 장갑 팔, 베이지 장갑판, 흰 얼굴, 낡은 코트와 모자는 참조와 부합한다. 상자에는 판독 가능한 글자가 보이지 않는다. 숨긴 카드와 외부 지붕 십자가는 확인 범위 밖이다.",
        "hard_violations": [],
        "physics": "현우와 앰버는 바닥에 엉덩이를 낮추고 다리를 접었으며 부츠가 바닥에 닿는다. 두 사람의 손과 팔은 정강이나 무릎을 감싸 실제 접촉을 이룬다. 찰리는 다리를 접고 팔을 무릎 위에 포갠 자세로 보이나 발과 좌면 접점은 상자에 가려져 있다. 부유를 나타내는 틈이나 불가능한 자세는 보이지 않는다. 상자는 바닥과 아래쪽 상자에 지지되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.083
   },
   "violations": {
    "B": [
     "[gpt-high] 왼쪽 종이상자에 '취급주의' 등 판독 가능한 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1083,
   "A": 1571
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "현우의 겉옷 탈의와 다리 부상, 앰버의 입을 가린 상태를 정확히 반영했으며, 상자 틈새에 욱여넣어진 비좁은 구도를 훌륭하게 구현했습니다.  ★위반: [gpt-high] 왼쪽 종이상자에 '취급주의' 등 판독 가능한 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "현우가 여전히 겉옷을 입고 있어 지시사항을 위반했으며, 인물들이 상자 사이가 아닌 텅 빈 중앙 공간에 나란히 앉아 있어 긴장감이 떨어집니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
    "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-2299-7695-8d82-e905f66c18ab",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh2__bgfirst_bg.png",
   "bg_asset_id": "bc133181-b478-4d81-83c2-7aa1d682c287",
   "bg_record_key": "S28sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S28sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:13:01.753282+00:00",
  "fingerprint": "efd5dbbfc71b15e816544f78d1ebf8785bbe45498ee522edb8009c3d2538ccc6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S28sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S28sh2_sel.png",
  "source_sha256": "1bb235639bac7d20f7770fae5d2bbfa8e1be08a4fc18aef10555eabd7b197740",
  "file": "S28sh2_cine.png",
  "staged_sha256": "9acd80196959392bee75261e451ced2737198e05ae2ddb13f0ad863e3efc9968",
  "latency_ms": 12310
 },
 "S28sh5::signage": {
  "fp": "0227ca4fd6663347",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S28sh5::bgfirst_bg": {
  "input_fingerprint": "ee9bf078ea28e12d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh5__bgfirst_bg.png",
  "asset_id": "2fb1b93e-74b7-4cf0-85b0-9fb67e335dcd",
  "input_asset_ids": [
   "029b4a24-f56c-4f1f-aeee-61bd3bdf819c",
   "f6306063-3d19-4618-946d-060af53e1589"
  ]
 },
 "S28sh5": {
  "input_fingerprint": "198ac8749577c253",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rear cargo boxes and roof-mounted halogen cross remain in place. Charlie is concealed among the boxes, still wearing his old coat and hat.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원들 right now, so 민병대원들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rear cargo boxes and roof-mounted halogen cross remain in place. Charlie is concealed among the boxes, still wearing his old coat and hat.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원들 right now, so 민병대원들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rear cargo boxes and roof-mounted halogen cross remain in place. Charlie is concealed among the boxes, still wearing his old coat and hat.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원들 right now, so 민병대원들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh5__bgfirst_bg.png",
     "asset_id": "2fb1b93e-74b7-4cf0-85b0-9fb67e335dcd",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S28sh5.png",
     "asset_id": "029b4a24-f56c-4f1f-aeee-61bd3bdf819c",
     "role": "conti_light"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
     "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
     "role": "location_plate"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "중앙의 민병대원이 손을 들어 밴을 멈춰 세우고 있으며, 운전자는 정면을 주시함.",
    "built_space": "승합차 내부. 대시보드와 앞좌석, 뒤편의 화물 상자 구역이 보임.",
    "entities": "운전석에 앉은 지시문에 없는 인물, 상자 사이에 숨은 찰리, 외부의 민병대원 3명. 대원들 주변 허공에 흰색 화살표 마커가 떠 있음. 지정된 형태의 총기는 구현되지 않음.",
    "hard_violations": [
     "[gemini-pro] 지시문에 없는 인물(운전자) 추가",
     "[gemini-pro] 화면에 흰색 화살표 마커 유출",
     "[gpt-high] 실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
     "[gpt-high] 샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
    ],
    "physics": "운전자와 찰리는 차량 내부에 자리 잡고 지탱되며, 외부 대원들은 지면에 두 발로 서 있음."
   },
   {
    "label": "B",
    "direction": "중앙의 민병대원이 밴 정면(카메라)을 향해 권총을 정조준하고 있음.",
    "built_space": "승합차 앞좌석 사이에서 밖을 내다보는 시점. 빈 앞좌석과 대시보드, 전방 도로의 바리케이드가 보임.",
    "entities": "도로를 막아선 4명의 민병대원. 찰리와 화물 상자는 프레임 밖에 있어 보이지 않음. 참조된 특수 총기는 정확히 묘사되지 않고 일반 권총과 소총들로 대체됨.",
    "hard_violations": [
     "[gpt-high] 운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
    ],
    "physics": "대원들은 지면에 안정적으로 서서 손으로 무기를 파지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구도의 한계로 찰리가 프레임에서 제외되었으나, 앞유리창 너머로 위협을 가하는 민병대의 전경을 규칙 위반 없이 성공적으로 묘사했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시문에 명시되지 않은 운전자가 프레임에 추가되었고, 화면 양측에 화살표 형태의 마커가 유출되어 치명적인 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 민병대원이 손을 들어 밴을 멈춰 세우고 있으며, 운전자는 정면을 주시함.",
        "built_space": "승합차 내부. 대시보드와 앞좌석, 뒤편의 화물 상자 구역이 보임.",
        "entities": "운전석에 앉은 지시문에 없는 인물, 상자 사이에 숨은 찰리, 외부의 민병대원 3명. 대원들 주변 허공에 흰색 화살표 마커가 떠 있음. 지정된 형태의 총기는 구현되지 않음.",
        "hard_violations": [
         "지시문에 없는 인물(운전자) 추가",
         "화면에 흰색 화살표 마커 유출"
        ],
        "physics": "운전자와 찰리는 차량 내부에 자리 잡고 지탱되며, 외부 대원들은 지면에 두 발로 서 있음."
       },
       {
        "label": "B",
        "direction": "중앙의 민병대원이 밴 정면(카메라)을 향해 권총을 정조준하고 있음.",
        "built_space": "승합차 앞좌석 사이에서 밖을 내다보는 시점. 빈 앞좌석과 대시보드, 전방 도로의 바리케이드가 보임.",
        "entities": "도로를 막아선 4명의 민병대원. 찰리와 화물 상자는 프레임 밖에 있어 보이지 않음. 참조된 특수 총기는 정확히 묘사되지 않고 일반 권총과 소총들로 대체됨.",
        "hard_violations": [],
        "physics": "대원들은 지면에 안정적으로 서서 손으로 무기를 파지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구도의 한계로 찰리가 프레임에서 제외되었으나, 앞유리창 너머로 위협을 가하는 민병대의 전경을 규칙 위반 없이 성공적으로 묘사했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시문에 명시되지 않은 운전자가 프레임에 추가되었고, 화면 양측에 화살표 형태의 마커가 유출되어 치명적인 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 민병대원이 손을 들어 밴을 멈춰 세우고 있으며, 운전자는 정면을 주시함.",
        "built_space": "승합차 내부. 대시보드와 앞좌석, 뒤편의 화물 상자 구역이 보임.",
        "entities": "운전석에 앉은 지시문에 없는 인물, 상자 사이에 숨은 찰리, 외부의 민병대원 3명. 대원들 주변 허공에 흰색 화살표 마커가 떠 있음. 지정된 형태의 총기는 구현되지 않음.",
        "hard_violations": [
         "지시문에 없는 인물(운전자) 추가",
         "화면에 흰색 화살표 마커 유출"
        ],
        "physics": "운전자와 찰리는 차량 내부에 자리 잡고 지탱되며, 외부 대원들은 지면에 두 발로 서 있음."
       },
       {
        "label": "B",
        "direction": "중앙의 민병대원이 밴 정면(카메라)을 향해 권총을 정조준하고 있음.",
        "built_space": "승합차 앞좌석 사이에서 밖을 내다보는 시점. 빈 앞좌석과 대시보드, 전방 도로의 바리케이드가 보임.",
        "entities": "도로를 막아선 4명의 민병대원. 찰리와 화물 상자는 프레임 밖에 있어 보이지 않음. 참조된 특수 총기는 정확히 묘사되지 않고 일반 권총과 소총들로 대체됨.",
        "hard_violations": [],
        "physics": "대원들은 지면에 안정적으로 서서 손으로 무기를 파지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "앞유리 너머 차량을 가로막는 민병대의 위협이라는 핵심 구도는 더 충실하지만, 노출된 로고와 한국인 설정에 맞지 않는 인물 구성, 참조 총기와의 차이가 남는다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "흰 화살표가 화면에 노출되고 지시되지 않은 운전자가 추가되었으며, 화물칸까지 넓힌 구도가 앞유리 밖 민병대의 위협을 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙 남성은 차 안을 똑바로 보며 권총 총구를 앞유리와 카메라 쪽으로 향한다. 나머지 네 명은 대체로 차량을 향해 서 있고, 긴 총들은 몸 앞에서 아래쪽 또는 비스듬한 옆쪽을 향한다. 왼쪽에서 두 번째 인물의 총은 거의 수직으로 내려가 있다. 모든 총이 차량을 겨눠야 한다는 지시는 없으므로 이 방향 배치는 위협적인 길막 장면과 양립한다.",
        "built_space": "밴 내부에서 앞유리 하나를 통해 도로를 직접 본다. 앞좌석 두 개의 일부, 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 앞유리 아래 와이퍼 두 개가 보인다. 참조와 유사한 낡은 대시보드와 중앙 조작부가 유지된다. 밖에는 다섯 명이 도로 폭을 가로질러 서 있고 그 뒤로 가로 차단봉 하나가 놓여 있다. 낮은 건물과 가로등이 밤길을 따라 이어지며, 위협을 반사상으로 대체하지 않았다.",
        "entities": "민병대는 성인 남성 다섯 명이며 낡은 야전 재킷과 모자, 총기를 착용하거나 들고 있다. 중앙 인물은 백인으로, 오른쪽 끝 인물 등은 흑인으로 보이므로 별도 설명 없는 현지인은 한국인이라는 설정과 맞지 않는다. 중앙의 짧은 권총과 주변의 긴 소총들은 참조의 개머리판이 결합된 권총형 총기와 다르다. 운전대에는 식별 가능한 현대 로고가 있다. 후방 상자, 찰리, 지붕 십자가는 이 구도에서 보이지 않으므로 부재로 판정하지 않는다.",
        "hard_violations": [
         "운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
        ],
        "physics": "중앙 인물은 두 손으로 권총을 받쳐 들고, 나머지 인물들도 손과 팔로 총을 지지한다. 인물들의 하체 일부는 대시보드와 창틀에 가리지만 지면에 서 있는 자세로 자연스럽게 이어지며 공중에 떠 있는 징후는 없다. 차단봉은 수직 기둥으로 지지되고 실내 부품들도 정상적으로 장착되어 있다."
       },
       {
        "label": "B",
        "direction": "중앙 민병대원은 차량 내부를 바라보면서 손바닥을 차량 쪽으로 들어 정지를 지시한다. 그의 총구는 화면 오른쪽 아래를 향한다. 왼쪽 대원의 얼굴과 소총은 도로 중앙 및 오른쪽 아래를 향하고, 오른쪽 대원은 오른쪽으로 이동하며 총을 낮추고 있다. 운전자는 앞길을 보고 있으며, 후방의 모자 쓴 인물은 상자 사이로 몸을 숙인다.",
        "built_space": "카메라는 화물 공간 쪽에서 앞좌석 두 개와 앞유리 하나를 본다. 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 중앙 화면 하나가 있으며 참조의 라디오 중심 조작부와는 다르다. 왼쪽 큰 상자 하나와 중앙·오른쪽의 여러 상자 및 운반함이 화면 아래를 크게 차지한다. 밖에는 민병대원 세 명, 도로 양옆의 콘크리트 방벽, 교차형 장애물과 오른쪽 초소가 보인다. 참조에서 확인되는 장소보다 검문소 시설을 훨씬 구체적으로 추가했고, 실내 비중이 커져 외부 위협의 크기가 줄었다.",
        "entities": "외부의 세 대원은 동아시아계 성인 남성으로 보이며 군복과 방탄 장비를 착용한다. 총들은 참조의 개머리판 부착 권총보다는 긴 총몸과 전방 구조를 가진 소총형으로 보인다. 내부에는 샷 텍스트에 없는 운전자 한 명이 추가되어 있다. 후방에는 낡은 외투와 모자를 쓴 찰리로 해석되는 인물이 있지만 상체가 상당히 드러나 은폐 상태가 약하다. 화면 밖 지붕 십자가는 확인할 수 없다. 외부 인물과 건물 주변에 흰 화살표 표식 여러 개가 붙어 있다.",
        "hard_violations": [
         "실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
         "샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
        ],
        "physics": "중앙 대원은 한 손으로 총을 잡고 다른 손을 들어 정지 신호를 하며, 총은 몸에 붙어 지지된다. 양옆 대원들도 손으로 총을 잡고 있고 오른쪽 대원은 다리를 벌린 보행 자세로 지면을 딛는다. 운전자는 좌석에 앉아 운전대를 잡는다. 후방 인물은 상자 뒤로 숙여 하체 지지가 가려져 있으나 떠 있다고 볼 근거는 없다. 상자와 운반함은 바닥 또는 다른 상자 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "앞유리 너머 차량을 가로막는 민병대의 위협이라는 핵심 구도는 더 충실하지만, 노출된 로고와 한국인 설정에 맞지 않는 인물 구성, 참조 총기와의 차이가 남는다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "흰 화살표가 화면에 노출되고 지시되지 않은 운전자가 추가되었으며, 화물칸까지 넓힌 구도가 앞유리 밖 민병대의 위협을 약화한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "중앙 남성은 차 안을 똑바로 보며 권총 총구를 앞유리와 카메라 쪽으로 향한다. 나머지 네 명은 대체로 차량을 향해 서 있고, 긴 총들은 몸 앞에서 아래쪽 또는 비스듬한 옆쪽을 향한다. 왼쪽에서 두 번째 인물의 총은 거의 수직으로 내려가 있다. 모든 총이 차량을 겨눠야 한다는 지시는 없으므로 이 방향 배치는 위협적인 길막 장면과 양립한다.",
        "built_space": "밴 내부에서 앞유리 하나를 통해 도로를 직접 본다. 앞좌석 두 개의 일부, 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 앞유리 아래 와이퍼 두 개가 보인다. 참조와 유사한 낡은 대시보드와 중앙 조작부가 유지된다. 밖에는 다섯 명이 도로 폭을 가로질러 서 있고 그 뒤로 가로 차단봉 하나가 놓여 있다. 낮은 건물과 가로등이 밤길을 따라 이어지며, 위협을 반사상으로 대체하지 않았다.",
        "entities": "민병대는 성인 남성 다섯 명이며 낡은 야전 재킷과 모자, 총기를 착용하거나 들고 있다. 중앙 인물은 백인으로, 오른쪽 끝 인물 등은 흑인으로 보이므로 별도 설명 없는 현지인은 한국인이라는 설정과 맞지 않는다. 중앙의 짧은 권총과 주변의 긴 소총들은 참조의 개머리판이 결합된 권총형 총기와 다르다. 운전대에는 식별 가능한 현대 로고가 있다. 후방 상자, 찰리, 지붕 십자가는 이 구도에서 보이지 않으므로 부재로 판정하지 않는다.",
        "hard_violations": [
         "운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
        ],
        "physics": "중앙 인물은 두 손으로 권총을 받쳐 들고, 나머지 인물들도 손과 팔로 총을 지지한다. 인물들의 하체 일부는 대시보드와 창틀에 가리지만 지면에 서 있는 자세로 자연스럽게 이어지며 공중에 떠 있는 징후는 없다. 차단봉은 수직 기둥으로 지지되고 실내 부품들도 정상적으로 장착되어 있다."
       },
       {
        "label": "A",
        "direction": "중앙 민병대원은 차량 내부를 바라보면서 손바닥을 차량 쪽으로 들어 정지를 지시한다. 그의 총구는 화면 오른쪽 아래를 향한다. 왼쪽 대원의 얼굴과 소총은 도로 중앙 및 오른쪽 아래를 향하고, 오른쪽 대원은 오른쪽으로 이동하며 총을 낮추고 있다. 운전자는 앞길을 보고 있으며, 후방의 모자 쓴 인물은 상자 사이로 몸을 숙인다.",
        "built_space": "카메라는 화물 공간 쪽에서 앞좌석 두 개와 앞유리 하나를 본다. 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 중앙 화면 하나가 있으며 참조의 라디오 중심 조작부와는 다르다. 왼쪽 큰 상자 하나와 중앙·오른쪽의 여러 상자 및 운반함이 화면 아래를 크게 차지한다. 밖에는 민병대원 세 명, 도로 양옆의 콘크리트 방벽, 교차형 장애물과 오른쪽 초소가 보인다. 참조에서 확인되는 장소보다 검문소 시설을 훨씬 구체적으로 추가했고, 실내 비중이 커져 외부 위협의 크기가 줄었다.",
        "entities": "외부의 세 대원은 동아시아계 성인 남성으로 보이며 군복과 방탄 장비를 착용한다. 총들은 참조의 개머리판 부착 권총보다는 긴 총몸과 전방 구조를 가진 소총형으로 보인다. 내부에는 샷 텍스트에 없는 운전자 한 명이 추가되어 있다. 후방에는 낡은 외투와 모자를 쓴 찰리로 해석되는 인물이 있지만 상체가 상당히 드러나 은폐 상태가 약하다. 화면 밖 지붕 십자가는 확인할 수 없다. 외부 인물과 건물 주변에 흰 화살표 표식 여러 개가 붙어 있다.",
        "hard_violations": [
         "실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
         "샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
        ],
        "physics": "중앙 대원은 한 손으로 총을 잡고 다른 손을 들어 정지 신호를 하며, 총은 몸에 붙어 지지된다. 양옆 대원들도 손으로 총을 잡고 있고 오른쪽 대원은 다리를 벌린 보행 자세로 지면을 딛는다. 운전자는 좌석에 앉아 운전대를 잡는다. 후방 인물은 상자 뒤로 숙여 하체 지지가 가려져 있으나 떠 있다고 볼 근거는 없다. 상자와 운반함은 바닥 또는 다른 상자 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.762,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.512,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시문에 없는 인물(운전자) 추가",
     "[gemini-pro] 화면에 흰색 화살표 마커 유출",
     "[gpt-high] 실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
     "[gpt-high] 샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
    ],
    "B": [
     "[gpt-high] 운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 512
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "구도의 한계로 찰리가 프레임에서 제외되었으나, 앞유리창 너머로 위협을 가하는 민병대의 전경을 규칙 위반 없이 성공적으로 묘사했습니다.  ★위반: [gpt-high] 운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 512,
    "verdict_ko": "지시문에 명시되지 않은 운전자가 프레임에 추가되었고, 화면 양측에 화살표 형태의 마커가 유출되어 치명적인 위반이 발생했습니다.  ★위반: [gemini-pro] 지시문에 없는 인물(운전자) 추가 / [gemini-pro] 화면에 흰색 화살표 마커 유출 / [gpt-high] 실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다. / [gpt-high] 샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
    "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
    "role": "location_plate"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-25e5-78e6-a8cb-7c77c6f87238",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh5__bgfirst_bg.png",
   "bg_asset_id": "2fb1b93e-74b7-4cf0-85b0-9fb67e335dcd",
   "bg_record_key": "S28sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S28sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:29:41.624724+00:00",
  "fingerprint": "1b838b6bbcd292dc6e45f3ba0c6f7848667a51cd292343588a51fa89870ed32d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S28sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S28sh5_sel.png",
  "source_sha256": "6115c6a9dc5bb761a567220ea203835422021d49f560fa5eb86bfe98756071de",
  "file": "S28sh5_cine.png",
  "staged_sha256": "38740b888248d9b956da3b795a54c63c33ad44b3a63011d92ca3c83d72712eb5",
  "latency_ms": 10970
 },
 "S28sh7::confined_fp_apt": {
  "applies": true,
  "reason_ko": "차량(밴)의 운전석 내부를 배경으로 하는 숏으로, 운전대 및 열린 운전석 창문과 외부 인물(민병대원들) 간의 공간적 위치 관계를 정확히 묘사해야 하므로 평면도 레이아웃 가이드가 필요합니다.",
  "input_fingerprint": "3be4f6485ed8fa97"
 },
 "S28sh7::signage": {
  "fp": "231d94694e2e23f2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::9fd979762448": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_9fd979762448.png",
  "place_text": "At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop.",
  "input_fingerprint": "5612a230c9bd725a"
 },
 "S28sh7::confined_fp": {
  "reads": {
   "controls": "The steering wheel is located at the front-left station (driver's seat).",
   "mirrors": "No mirrors are depicted in the diagram.",
   "camera": "The camera is positioned in the middle of the cabin, forward of the seats, pointing directly to the left towards the driver's seat and the adjacent side window.",
   "occupants": "신부 is seated in the left (driver's) seat. The right (passenger) seat is marked but unoccupied."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned inside the center of the cabin, pointing directly left at the driver's seat in a close-up framing. In the center foreground, 신부 is seated, seen in profile facing left. On the left edge of the frame, slightly in front of him, the right side of the steering wheel is visible. Beyond his profile on the far left, the driver's side window is entirely lowered, opening to the exterior night. The view focuses entirely on the driver and the window, with the rest of the cabin behind the camera or off-screen to the right.",
  "fixed": true,
  "input_fingerprint": "25f45d1e1eff4916"
 },
 "S28sh7": {
  "input_fingerprint": "08ee15010f84ca79",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 운전석 창문 너머의 민병대원들을 향해 활짝 웃으며 찡긋 눈웃음을 짓는 신부의 여유로운 얼굴.\n\nLOCATION (lock): At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Open driver's window (Lowered for the exchange with the militia) — Its open aperture is seen obliquely beyond 신부's profile; no glass lies across the opening; used as Establishes the direction of his smile and offscreen attention; Driver's seat (Occupied by 신부) — A limited portion is visible behind his shoulder; used as Keeps the close portrait physically grounded inside the van.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient light across the cabin and open window, keeping the smile and wink legible without attributing facial illumination to the roof ornament.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The driver's window is lowered for the inspection, with the halogen cross fixed on the roof. The rear boxes continue to conceal Charlie in his old coat and hat. 신부: He remains seated at the wheel in his clerical collar, smiling through the lowered window.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned inside the center of the cabin, pointing directly left at the driver's seat in a close-up framing. In the center foreground, 신부 is seated, seen in profile facing left. On the left edge of the frame, slightly in front of him, the right side of the steering wheel is visible. Beyond his profile on the far left, the driver's side window is entirely lowered, opening to the exterior night. The view focuses entirely on the driver and the window, with the rest of the cabin behind the camera or off-screen to the right.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 운전석 창문 너머의 민병대원들을 향해 활짝 웃으며 찡긋 눈웃음을 짓는 신부의 여유로운 얼굴.\n\nLOCATION (lock): At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient light across the cabin and open window, keeping the smile and wink legible without attributing facial illumination to the roof ornament.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The driver's window is lowered for the inspection, with the halogen cross fixed on the roof. The rear boxes continue to conceal Charlie in his old coat and hat. 신부: He remains seated at the wheel in his clerical collar, smiling through the lowered window.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned inside the center of the cabin, pointing directly left at the driver's seat in a close-up framing. In the center foreground, 신부 is seated, seen in profile facing left. On the left edge of the frame, slightly in front of him, the right side of the steering wheel is visible. Beyond his profile on the far left, the driver's side window is entirely lowered, opening to the exterior night. The view focuses entirely on the driver and the window, with the rest of the cabin behind the camera or off-screen to the right.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 운전석 창문 너머의 민병대원들을 향해 활짝 웃으며 찡긋 눈웃음을 짓는 신부의 여유로운 얼굴.\n\nLOCATION (lock): At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient light across the cabin and open window, keeping the smile and wink legible without attributing facial illumination to the roof ornament.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The driver's window is lowered for the inspection, with the halogen cross fixed on the roof. The rear boxes continue to conceal Charlie in his old coat and hat. 신부: He remains seated at the wheel in his clerical collar, smiling through the lowered window.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh7_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "신부",
     "path": "<bytes:863986>",
     "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh7_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "신부",
     "path": "<bytes:863986>",
     "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부의 시선과 미소가 지시된 측면 운전석 창문(화면 우측 뒤편)이 아닌 차량 정면(화면 좌측)을 향하고 있음.",
    "built_space": "밴 운전석 내부 구도이나 카메라가 조수석 창문 밖으로 다소 물러나 전경 하단에 조수석 문틀과 잠금 핀이 보이며, 스티어링 휠이 화면 좌측 전방에 위치함.",
    "entities": "60대 한국인 남성 신부의 외양, 검은 셔츠와 로만 칼라가 기준 이미지와 완벽히 일치하며 지시된 찡긋하는 눈웃음이 묘사됨.",
    "hard_violations": [],
    "physics": "신부가 운전석 시트에 체중을 싣고 자연스럽게 앉아 있음."
   },
   {
    "label": "B",
    "direction": "신부의 시선이 측면 창문(화면 중앙 배경)이 아닌 정면 유리창 쪽(화면 좌측)을 향해 빗나감.",
    "built_space": "조수석 창문 밖에서 밴 내부를 촬영한 구도로 하단에 문틀이 보임. 배경에 운전석 창문과 바깥 풍경이 묘사됨.",
    "entities": "신부의 외양과 복장은 일치하나, 프롬프트에서 금지한 추가 인물인 민병대원 2명이 창문 밖에 렌더링됨.",
    "hard_violations": [
     "[gemini-pro] 프롬프트의 엄격한 인물 제한(신부 외 등장 불가)을 어기고 민병대원을 추가함",
     "[gemini-pro] 화면 밖(offscreen)으로 설정된 시선 대상을 화면 안에 직접 렌더링함",
     "[gpt-high] 신부만 화면에 등장하도록 제한한 인물 조건과 민병대원을 화면 밖 대상으로 둔 구성을 어기고 민병대원 두 명을 배경에 추가했다."
    ],
    "physics": "신부가 차량 시트에 안정적으로 앉아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "신부의 외양과 찡긋하는 눈웃음을 훌륭하게 재현하고 제약 사항을 준수했으나, 시선이 창문이 아닌 정면을 향한 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "신부 외의 인물(민병대원)을 배경에 등장시켜 인물 추가 금지 및 오프스크린 제약을 명백히 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선과 미소가 지시된 측면 운전석 창문(화면 우측 뒤편)이 아닌 차량 정면(화면 좌측)을 향하고 있음.",
        "built_space": "밴 운전석 내부 구도이나 카메라가 조수석 창문 밖으로 다소 물러나 전경 하단에 조수석 문틀과 잠금 핀이 보이며, 스티어링 휠이 화면 좌측 전방에 위치함.",
        "entities": "60대 한국인 남성 신부의 외양, 검은 셔츠와 로만 칼라가 기준 이미지와 완벽히 일치하며 지시된 찡긋하는 눈웃음이 묘사됨.",
        "hard_violations": [],
        "physics": "신부가 운전석 시트에 체중을 싣고 자연스럽게 앉아 있음."
       },
       {
        "label": "B",
        "direction": "신부의 시선이 측면 창문(화면 중앙 배경)이 아닌 정면 유리창 쪽(화면 좌측)을 향해 빗나감.",
        "built_space": "조수석 창문 밖에서 밴 내부를 촬영한 구도로 하단에 문틀이 보임. 배경에 운전석 창문과 바깥 풍경이 묘사됨.",
        "entities": "신부의 외양과 복장은 일치하나, 프롬프트에서 금지한 추가 인물인 민병대원 2명이 창문 밖에 렌더링됨.",
        "hard_violations": [
         "프롬프트의 엄격한 인물 제한(신부 외 등장 불가)을 어기고 민병대원을 추가함",
         "화면 밖(offscreen)으로 설정된 시선 대상을 화면 안에 직접 렌더링함"
        ],
        "physics": "신부가 차량 시트에 안정적으로 앉아 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "신부의 외양과 찡긋하는 눈웃음을 훌륭하게 재현하고 제약 사항을 준수했으나, 시선이 창문이 아닌 정면을 향한 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "신부 외의 인물(민병대원)을 배경에 등장시켜 인물 추가 금지 및 오프스크린 제약을 명백히 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선과 미소가 지시된 측면 운전석 창문(화면 우측 뒤편)이 아닌 차량 정면(화면 좌측)을 향하고 있음.",
        "built_space": "밴 운전석 내부 구도이나 카메라가 조수석 창문 밖으로 다소 물러나 전경 하단에 조수석 문틀과 잠금 핀이 보이며, 스티어링 휠이 화면 좌측 전방에 위치함.",
        "entities": "60대 한국인 남성 신부의 외양, 검은 셔츠와 로만 칼라가 기준 이미지와 완벽히 일치하며 지시된 찡긋하는 눈웃음이 묘사됨.",
        "hard_violations": [],
        "physics": "신부가 운전석 시트에 체중을 싣고 자연스럽게 앉아 있음."
       },
       {
        "label": "B",
        "direction": "신부의 시선이 측면 창문(화면 중앙 배경)이 아닌 정면 유리창 쪽(화면 좌측)을 향해 빗나감.",
        "built_space": "조수석 창문 밖에서 밴 내부를 촬영한 구도로 하단에 문틀이 보임. 배경에 운전석 창문과 바깥 풍경이 묘사됨.",
        "entities": "신부의 외양과 복장은 일치하나, 프롬프트에서 금지한 추가 인물인 민병대원 2명이 창문 밖에 렌더링됨.",
        "hard_violations": [
         "프롬프트의 엄격한 인물 제한(신부 외 등장 불가)을 어기고 민병대원을 추가함",
         "화면 밖(offscreen)으로 설정된 시선 대상을 화면 안에 직접 렌더링함"
        ],
        "physics": "신부가 차량 시트에 안정적으로 앉아 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "화면 밖에 있어야 할 민병대원 두 명을 추가했고, 얼굴 클로즈업 대신 운전석 전경을 넓게 담았으며 웃음의 방향도 운전석 창밖 상대를 명확히 향하지 않는다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "신부만 보이며 활짝 웃는 표정과 찡긋한 눈, 야간 분위기는 잘 맞지만, 얼굴 클로즈업보다 넓고 운전석 창밖 상대를 향한 고개와 시선이 불명확하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 얼굴과 가늘게 뜬 눈은 화면 왼쪽의 차량 앞쪽을 향한다. 가까운 운전석 창밖 상대를 향해 고개를 돌린 모습은 분명하지 않다. 보이는 민병대원 두 명은 신부 뒤편의 반대쪽 창 너머에 있어 신부의 웃음이 그들에게 향하지 않는다. 두 사람의 총기는 대체로 아래쪽을 향하며 신부를 겨누지 않는다.",
        "built_space": "가까운 쪽 열린 창틀 하나와 반대편 측면 창 하나, 운전대 하나, 운전석 머리받침 하나, 실내 후사경 하나, 양쪽 기둥의 손잡이 두 개가 보인다. 신부는 운전대 뒤에 있고 머리받침은 뒤쪽에 있어 착석 관계는 성립한다. 가까운 창의 개구부에는 유리가 보이지 않는다. 다만 카메라는 창밖에서 실내를 넓게 바라보며, 얼굴보다 창틀·대시보드·반대편 문이 차지하는 면적이 커서 요구한 클로즈업과 다르다. 불가능한 반사는 확인되지 않는다.",
        "entities": "신부는 주름진 얼굴과 짧은 회흑색 머리의 한국인 노년 남성으로 보이며, 참조의 얼굴과 검은 성직자복·흰 칼라에 대체로 부합한다. 치아가 드러나는 미소와 한쪽 눈을 더 감은 표정이 보인다. 그러나 배경에 군복과 헬멧을 착용한 민병대원 두 명 및 총기가 추가로 등장한다. 지붕 십자가와 후방 상자·찰리는 프레임 밖이므로 평가하지 않는다. 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [
         "신부만 화면에 등장하도록 제한한 인물 조건과 민병대원을 화면 밖 대상으로 둔 구성을 어기고 민병대원 두 명을 배경에 추가했다."
        ],
        "physics": "신부의 상체는 운전석 등받이와 머리받침 앞에 자연스럽게 놓여 있고, 좌석에 앉은 자세로 읽힌다. 골반과 다리는 프레임 밖이다. 배경 인물들은 직립해 있으며 총기는 몸 앞에서 팔과 손으로 받치고 있다. 떠 있는 신체나 지지 없이 공중에 놓인 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "신부는 화면 왼쪽, 운전대가 있는 차량 앞쪽으로 얼굴을 향한 채 활짝 웃고 한쪽 눈을 찡긋한다. 눈에 보이는 상대는 없으며, 가까운 운전석 창밖 민병대원에게 시선을 보낸다는 방향 관계는 뚜렷하지 않다. 무기나 움직이는 물체는 없다.",
        "built_space": "가까운 창의 하단과 측면 테두리, 반대편 측면 창 하나와 문 안쪽, 운전대 하나, 운전석 머리받침 하나, 기둥 손잡이 하나, 천장 손잡이 하나와 측면 거울 하나가 보인다. 신부는 운전대 뒤와 머리받침 앞에 앉아 있어 좌석 배치는 물리적으로 가능하다. 창 개구부를 가로막는 유리는 보이지 않는다. A보다 얼굴이 크지만 가슴과 운전대, 문을 상당 부분 포함해 요청한 얼굴 클로즈업보다 넓다. 불가능한 반사나 중복 설비는 보이지 않는다.",
        "entities": "등장인물은 신부 한 명이다. 한국인 60대 남성으로 보이는 얼굴, 이마와 눈가의 주름, 회흑색 머리, 검은 성직자복과 흰 칼라가 참조와 대체로 맞는다. 활짝 드러난 치아와 찡긋한 눈이 요구한 여유로운 표정을 전달한다. 목걸이 줄은 보이지만 펜던트 형태는 확인할 수 없다. 지붕 십자가와 후방 상자·찰리는 프레임 밖이며, 읽을 수 있는 문자나 도식은 없다.",
        "hard_violations": [],
        "physics": "신부의 등과 어깨 뒤에 좌석 및 머리받침이 있고, 상체는 자연스러운 착석 상태로 지지된다. 하체와 손은 잘려 있어 접촉 상태를 추가로 판단할 수 없다. 운전대와 손잡이, 거울은 차량 구조에 부착되어 있으며 지지 없는 물체나 불가능한 신체 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면 밖에 있어야 할 민병대원 두 명을 추가했고, 얼굴 클로즈업 대신 운전석 전경을 넓게 담았으며 웃음의 방향도 운전석 창밖 상대를 명확히 향하지 않는다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "신부만 보이며 활짝 웃는 표정과 찡긋한 눈, 야간 분위기는 잘 맞지만, 얼굴 클로즈업보다 넓고 운전석 창밖 상대를 향한 고개와 시선이 불명확하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 얼굴과 가늘게 뜬 눈은 화면 왼쪽의 차량 앞쪽을 향한다. 가까운 운전석 창밖 상대를 향해 고개를 돌린 모습은 분명하지 않다. 보이는 민병대원 두 명은 신부 뒤편의 반대쪽 창 너머에 있어 신부의 웃음이 그들에게 향하지 않는다. 두 사람의 총기는 대체로 아래쪽을 향하며 신부를 겨누지 않는다.",
        "built_space": "가까운 쪽 열린 창틀 하나와 반대편 측면 창 하나, 운전대 하나, 운전석 머리받침 하나, 실내 후사경 하나, 양쪽 기둥의 손잡이 두 개가 보인다. 신부는 운전대 뒤에 있고 머리받침은 뒤쪽에 있어 착석 관계는 성립한다. 가까운 창의 개구부에는 유리가 보이지 않는다. 다만 카메라는 창밖에서 실내를 넓게 바라보며, 얼굴보다 창틀·대시보드·반대편 문이 차지하는 면적이 커서 요구한 클로즈업과 다르다. 불가능한 반사는 확인되지 않는다.",
        "entities": "신부는 주름진 얼굴과 짧은 회흑색 머리의 한국인 노년 남성으로 보이며, 참조의 얼굴과 검은 성직자복·흰 칼라에 대체로 부합한다. 치아가 드러나는 미소와 한쪽 눈을 더 감은 표정이 보인다. 그러나 배경에 군복과 헬멧을 착용한 민병대원 두 명 및 총기가 추가로 등장한다. 지붕 십자가와 후방 상자·찰리는 프레임 밖이므로 평가하지 않는다. 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [
         "신부만 화면에 등장하도록 제한한 인물 조건과 민병대원을 화면 밖 대상으로 둔 구성을 어기고 민병대원 두 명을 배경에 추가했다."
        ],
        "physics": "신부의 상체는 운전석 등받이와 머리받침 앞에 자연스럽게 놓여 있고, 좌석에 앉은 자세로 읽힌다. 골반과 다리는 프레임 밖이다. 배경 인물들은 직립해 있으며 총기는 몸 앞에서 팔과 손으로 받치고 있다. 떠 있는 신체나 지지 없이 공중에 놓인 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "신부는 화면 왼쪽, 운전대가 있는 차량 앞쪽으로 얼굴을 향한 채 활짝 웃고 한쪽 눈을 찡긋한다. 눈에 보이는 상대는 없으며, 가까운 운전석 창밖 민병대원에게 시선을 보낸다는 방향 관계는 뚜렷하지 않다. 무기나 움직이는 물체는 없다.",
        "built_space": "가까운 창의 하단과 측면 테두리, 반대편 측면 창 하나와 문 안쪽, 운전대 하나, 운전석 머리받침 하나, 기둥 손잡이 하나, 천장 손잡이 하나와 측면 거울 하나가 보인다. 신부는 운전대 뒤와 머리받침 앞에 앉아 있어 좌석 배치는 물리적으로 가능하다. 창 개구부를 가로막는 유리는 보이지 않는다. A보다 얼굴이 크지만 가슴과 운전대, 문을 상당 부분 포함해 요청한 얼굴 클로즈업보다 넓다. 불가능한 반사나 중복 설비는 보이지 않는다.",
        "entities": "등장인물은 신부 한 명이다. 한국인 60대 남성으로 보이는 얼굴, 이마와 눈가의 주름, 회흑색 머리, 검은 성직자복과 흰 칼라가 참조와 대체로 맞는다. 활짝 드러난 치아와 찡긋한 눈이 요구한 여유로운 표정을 전달한다. 목걸이 줄은 보이지만 펜던트 형태는 확인할 수 없다. 지붕 십자가와 후방 상자·찰리는 프레임 밖이며, 읽을 수 있는 문자나 도식은 없다.",
        "hard_violations": [],
        "physics": "신부의 등과 어깨 뒤에 좌석 및 머리받침이 있고, 상체는 자연스러운 착석 상태로 지지된다. 하체와 손은 잘려 있어 접촉 상태를 추가로 판단할 수 없다. 운전대와 손잡이, 거울은 차량 구조에 부착되어 있으며 지지 없는 물체나 불가능한 신체 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트의 엄격한 인물 제한(신부 외 등장 불가)을 어기고 민병대원을 추가함",
     "[gemini-pro] 화면 밖(offscreen)으로 설정된 시선 대상을 화면 안에 직접 렌더링함",
     "[gpt-high] 신부만 화면에 등장하도록 제한한 인물 조건과 민병대원을 화면 밖 대상으로 둔 구성을 어기고 민병대원 두 명을 배경에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "신부의 외양과 찡긋하는 눈웃음을 훌륭하게 재현하고 제약 사항을 준수했으나, 시선이 창문이 아닌 정면을 향한 점이 아쉽습니다."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "신부 외의 인물(민병대원)을 배경에 등장시켜 인물 추가 금지 및 오프스크린 제약을 명백히 위반했습니다.  ★위반: [gemini-pro] 프롬프트의 엄격한 인물 제한(신부 외 등장 불가)을 어기고 민병대원을 추가함 / [gemini-pro] 화면 밖(offscreen)으로 설정된 시선 대상을 화면 안에 직접 렌더링함 / [gpt-high] 신부만 화면에 등장하도록 제한한 인물 조건과 민병대원을 화면 밖 대상으로 둔 구성을 어기고 민병대원 두 명을 배경에 추가했다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh7_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "신부",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-291d-7a27-af23-b9790e21cbd1",
  "confined_fp": {
   "base_key": "confinedfp::9fd979762448",
   "apt_reason": "차량(밴)의 운전석 내부를 배경으로 하는 숏으로, 운전대 및 열린 운전석 창문과 외부 인물(민병대원들) 간의 공간적 위치 관계를 정확히 묘사해야 하므로 평면도 레이아웃 가이드가 필요합니다.",
   "fixed": true,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S28sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:15:10.321260+00:00",
  "fingerprint": "c7d9a16c9147d2241eba30a789215879d3a082a3bce726235929c73bf9bbe877",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S28sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S28sh7_sel.png",
  "source_sha256": "7d67be84c08a12e73759e5f710bcfa5735e4bd4de592cb021e9ee0fbb9cbb957",
  "file": "S28sh7_cine.png",
  "staged_sha256": "1d9b7fafce0f16588c30d463093658662d07457dd1232b93a5c0160e228dd5c9",
  "latency_ms": 9839
 },
 "S29sh6::signage": {
  "fp": "0b03705f58a7f802",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::b7ea15bbe8cf36a0": {
  "subjects": [],
  "subject_text": "인천 성당 지하 기도실\n십자가와 작은 책상, 의자, 작은 풍금이 놓인 단출한 지하 공간. 구석에는 낡은 라디오가 있고 촛불과 랜턴이 실내를 밝힌다.",
  "identity": "canonical",
  "scope_id": "L180",
  "scope_role": "location_interior",
  "scope_sha": "e2c03ad952b754fc"
 },
 "S29sh6::bgfirst_bg": {
  "input_fingerprint": "211f365f6147facf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6__bgfirst_bg.png",
  "asset_id": "4df963a5-17a9-4cdd-a2c1-6834d045d891",
  "input_asset_ids": [
   "51f9964d-9eec-4e5f-a76a-7feb0e6f9752",
   "973d81a2-fcfc-4412-82af-c60138fde795"
  ]
 },
 "S29sh6": {
  "input_fingerprint": "e256cbd230da258c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement prayer room contains a cross, a small desk, chairs and a small organ, with candlelight brought down the stairs. Charlie retains his old coat and hat. 현우: He is in the basement prayer room with facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe. 앰버: She is in the basement, tired and hungry, retaining the outer garment used as a mouth covering. 신부: He is in the basement wearing his clerical collar and carrying a lit candle.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement prayer room contains a cross, a small desk, chairs and a small organ, with candlelight brought down the stairs. Charlie retains his old coat and hat. 현우: He is in the basement prayer room with facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe. 앰버: She is in the basement, tired and hungry, retaining the outer garment used as a mouth covering. 신부: He is in the basement wearing his clerical collar and carrying a lit candle.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement prayer room contains a cross, a small desk, chairs and a small organ, with candlelight brought down the stairs. Charlie retains his old coat and hat. 현우: He is in the basement prayer room with facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe. 앰버: She is in the basement, tired and hungry, retaining the outer garment used as a mouth covering. 신부: He is in the basement wearing his clerical collar and carrying a lit candle.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6__bgfirst_bg.png",
     "asset_id": "4df963a5-17a9-4cdd-a2c1-6834d045d891",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S29sh6.png",
     "asset_id": "51f9964d-9eec-4e5f-a76a-7feb0e6f9752",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:863986>",
     "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L180B02.png",
     "asset_id": "973d81a2-fcfc-4412-82af-c60138fde795",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:863986>",
     "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "신부, 현우, 앰버는 소리가 나는 위쪽 천장(화면 좌측 상단)을 향해 시선을 두고 있으나, 찰리는 위를 보지 않고 카메라 정면을 응시하고 있어 '일제히' 위를 보는 액션이 완전히 일치하지 않음.",
    "built_space": "레퍼런스와 동일한 지하실 기도실 배경. 왼쪽 전경에 위로 올라가는 계단이 배치되어 앵커 역할을 제대로 수행함. 뒤쪽으로 책상과 의자, 벽면의 십자가가 보이며, 화면 우측 후경에 풍금이 위치해 지정된 공간 구성 요소를 모두 정확한 스케일로 포함하고 있음.",
    "entities": "현우는 겉옷 없이 회색 셔츠를 입고 있으며 프롬프트에 명시된 얼굴의 멍과 '다친 다리(붕대 묘사)'를 정확히 표현함. 앰버는 겉옷으로 입을 가린 어린아이의 모습임. 신부는 로만 칼라를 착용하고 촛불을 든 노년 남성임. 찰리는 레퍼런스 이미지의 비인간적인 '고릴라형 몸체(육중한 어깨와 팔, 짧은 다리)'와 헬멧, 코트를 완벽하게 재현함.",
    "hard_violations": [
     "[gpt-high] 찰리의 어깨 장갑판에 판독 가능한 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "모든 인물은 바닥에 안정적으로 서 있음. 현우와 앰버는 타격음에 놀라 어깨를 움츠리고 몸을 굽힌(cringing) 자세를 자연스럽게 보여주어 상황의 물리적 반응을 잘 나타냄."
   },
   {
    "label": "B",
    "direction": "현우, 앰버, 찰리, 신부 네 명 모두 명시된 대로 천장/계단 위쪽을 향해 시선을 일제히 향하고 있음.",
    "built_space": "기도실 배경으로 왼쪽의 계단, 중앙 후경의 책상과 십자가, 우측의 풍금 등 요구된 모든 배경 요소가 적절한 위치와 크기로 배치됨.",
    "entities": "현우는 얼굴에 멍이 있으나 명시된 '다친 다리'에 대한 묘사가 없으며, 티셔츠 위에 셔츠를 걸친 형태로 레퍼런스와 다름. 앰버는 입을 옷으로 가리고 있음. 신부는 조건에 맞음. 찰리는 프롬프트와 레퍼런스에서 가장 중요한 특징인 '고릴라형 몸체, 육중하고 긴 팔, 짧은 다리'를 완전히 상실하고 일반적인 인간 남성의 체형으로 렌더링되어 심각한 캐릭터 불일치를 보임.",
    "hard_violations": [],
    "physics": "모든 인물이 바닥에 서 있음. 시선은 위를 향하고 있으나, '어깨를 움츠린 채'라는 긴장감 있는 자세 묘사가 부족하여 다들 다소 꼿꼿하게 서 있는 모습임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찰리가 정면을 응시하여 '일제히 위를 올려다보는' 행동 묘사에서 다소 아쉬움이 있으나, 찰리의 육중한 고릴라형 체형, 현우의 다친 다리(붕대) 등 캐릭터의 핵심 디테일과 공간 분위기를 완벽하게 재현하여 우수한 결과를 냈습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "모든 인물이 위를 쳐다보는 시선 방향은 맞췄으나, 찰리가 프롬프트에 명시된 '고릴라형 몸체'를 완전히 상실하고 일반적인 인간 체형으로 잘못 렌더링되었으며 현우의 다친 다리 묘사도 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부, 현우, 앰버는 소리가 나는 위쪽 천장(화면 좌측 상단)을 향해 시선을 두고 있으나, 찰리는 위를 보지 않고 카메라 정면을 응시하고 있어 '일제히' 위를 보는 액션이 완전히 일치하지 않음.",
        "built_space": "레퍼런스와 동일한 지하실 기도실 배경. 왼쪽 전경에 위로 올라가는 계단이 배치되어 앵커 역할을 제대로 수행함. 뒤쪽으로 책상과 의자, 벽면의 십자가가 보이며, 화면 우측 후경에 풍금이 위치해 지정된 공간 구성 요소를 모두 정확한 스케일로 포함하고 있음.",
        "entities": "현우는 겉옷 없이 회색 셔츠를 입고 있으며 프롬프트에 명시된 얼굴의 멍과 '다친 다리(붕대 묘사)'를 정확히 표현함. 앰버는 겉옷으로 입을 가린 어린아이의 모습임. 신부는 로만 칼라를 착용하고 촛불을 든 노년 남성임. 찰리는 레퍼런스 이미지의 비인간적인 '고릴라형 몸체(육중한 어깨와 팔, 짧은 다리)'와 헬멧, 코트를 완벽하게 재현함.",
        "hard_violations": [],
        "physics": "모든 인물은 바닥에 안정적으로 서 있음. 현우와 앰버는 타격음에 놀라 어깨를 움츠리고 몸을 굽힌(cringing) 자세를 자연스럽게 보여주어 상황의 물리적 반응을 잘 나타냄."
       },
       {
        "label": "B",
        "direction": "현우, 앰버, 찰리, 신부 네 명 모두 명시된 대로 천장/계단 위쪽을 향해 시선을 일제히 향하고 있음.",
        "built_space": "기도실 배경으로 왼쪽의 계단, 중앙 후경의 책상과 십자가, 우측의 풍금 등 요구된 모든 배경 요소가 적절한 위치와 크기로 배치됨.",
        "entities": "현우는 얼굴에 멍이 있으나 명시된 '다친 다리'에 대한 묘사가 없으며, 티셔츠 위에 셔츠를 걸친 형태로 레퍼런스와 다름. 앰버는 입을 옷으로 가리고 있음. 신부는 조건에 맞음. 찰리는 프롬프트와 레퍼런스에서 가장 중요한 특징인 '고릴라형 몸체, 육중하고 긴 팔, 짧은 다리'를 완전히 상실하고 일반적인 인간 남성의 체형으로 렌더링되어 심각한 캐릭터 불일치를 보임.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 서 있음. 시선은 위를 향하고 있으나, '어깨를 움츠린 채'라는 긴장감 있는 자세 묘사가 부족하여 다들 다소 꼿꼿하게 서 있는 모습임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찰리가 정면을 응시하여 '일제히 위를 올려다보는' 행동 묘사에서 다소 아쉬움이 있으나, 찰리의 육중한 고릴라형 체형, 현우의 다친 다리(붕대) 등 캐릭터의 핵심 디테일과 공간 분위기를 완벽하게 재현하여 우수한 결과를 냈습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "모든 인물이 위를 쳐다보는 시선 방향은 맞췄으나, 찰리가 프롬프트에 명시된 '고릴라형 몸체'를 완전히 상실하고 일반적인 인간 체형으로 잘못 렌더링되었으며 현우의 다친 다리 묘사도 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부, 현우, 앰버는 소리가 나는 위쪽 천장(화면 좌측 상단)을 향해 시선을 두고 있으나, 찰리는 위를 보지 않고 카메라 정면을 응시하고 있어 '일제히' 위를 보는 액션이 완전히 일치하지 않음.",
        "built_space": "레퍼런스와 동일한 지하실 기도실 배경. 왼쪽 전경에 위로 올라가는 계단이 배치되어 앵커 역할을 제대로 수행함. 뒤쪽으로 책상과 의자, 벽면의 십자가가 보이며, 화면 우측 후경에 풍금이 위치해 지정된 공간 구성 요소를 모두 정확한 스케일로 포함하고 있음.",
        "entities": "현우는 겉옷 없이 회색 셔츠를 입고 있으며 프롬프트에 명시된 얼굴의 멍과 '다친 다리(붕대 묘사)'를 정확히 표현함. 앰버는 겉옷으로 입을 가린 어린아이의 모습임. 신부는 로만 칼라를 착용하고 촛불을 든 노년 남성임. 찰리는 레퍼런스 이미지의 비인간적인 '고릴라형 몸체(육중한 어깨와 팔, 짧은 다리)'와 헬멧, 코트를 완벽하게 재현함.",
        "hard_violations": [],
        "physics": "모든 인물은 바닥에 안정적으로 서 있음. 현우와 앰버는 타격음에 놀라 어깨를 움츠리고 몸을 굽힌(cringing) 자세를 자연스럽게 보여주어 상황의 물리적 반응을 잘 나타냄."
       },
       {
        "label": "B",
        "direction": "현우, 앰버, 찰리, 신부 네 명 모두 명시된 대로 천장/계단 위쪽을 향해 시선을 일제히 향하고 있음.",
        "built_space": "기도실 배경으로 왼쪽의 계단, 중앙 후경의 책상과 십자가, 우측의 풍금 등 요구된 모든 배경 요소가 적절한 위치와 크기로 배치됨.",
        "entities": "현우는 얼굴에 멍이 있으나 명시된 '다친 다리'에 대한 묘사가 없으며, 티셔츠 위에 셔츠를 걸친 형태로 레퍼런스와 다름. 앰버는 입을 옷으로 가리고 있음. 신부는 조건에 맞음. 찰리는 프롬프트와 레퍼런스에서 가장 중요한 특징인 '고릴라형 몸체, 육중하고 긴 팔, 짧은 다리'를 완전히 상실하고 일반적인 인간 남성의 체형으로 렌더링되어 심각한 캐릭터 불일치를 보임.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 서 있음. 시선은 위를 향하고 있으나, '어깨를 움츠린 채'라는 긴장감 있는 자세 묘사가 부족하여 다들 다소 꼿꼿하게 서 있는 모습임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "왼쪽 상승 계단과 일행의 전경은 맞지만, 앰버가 너무 성숙하고 찰리의 체형·현우의 겉옷 상태 및 일제히 움츠리는 반응이 부정확하다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물 정체성·부상·움츠린 자세와 장소 재현은 더 충실하지만, 찰리 어깨의 읽을 수 있는 숫자 표기가 명시적인 무문자 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 신부는 왼쪽 계단 위쪽의 화면 밖을 올려다본다. 앰버도 왼쪽을 보지만 눈길의 상승 각도는 약하다. 찰리는 얼굴을 거의 수평으로 왼쪽에 향하고 있어, 네 명 모두가 위층의 타격음에 일제히 올려다보는 동작은 완성되지 않는다. 배경 십자가를 바라보는 사람은 없다.",
        "built_space": "왼쪽 전경의 계단 한 줄이 위로 이어져 화면 밖으로 나간다. 일행은 계단 오른쪽의 기도실 바닥에 서 있다. 뒤에는 책상 하나, 등받이 의자 두 개, 오른쪽 오르간 하나와 작은 좌석 하나, 벽 십자가 하나, 왼쪽 책장 하나가 보인다. 천장등 하나와 왼쪽 벽등 하나가 보이며, 낡은 벽과 콘크리트 바닥은 참조 장소와 대체로 맞는다. 넓은 전신 구도와 배경 가구의 상대적 크기는 적절하다.",
        "entities": "추가 인물 없이 네 인물이 있다. 현우는 검은 머리의 젊은 동아시아계 남성이고 얼굴에 상처 흔적이 있지만, 티셔츠 위에 열린 겉셔츠를 입고 있으며 다리 부상은 드러나지 않는다. 앰버는 금발이지만 10세 아동보다 훨씬 성숙한 체격과 얼굴이며 참조 의상도 다르다. 손으로 입가의 천을 붙잡고 있다. 찰리는 낡은 코트와 모자, 베이지색 장갑과 흰 가면을 갖췄지만, 긴 다리와 보통 길이의 팔 때문에 참조의 육중한 고릴라형 체형과 크게 다르다. 신부는 나이 든 동아시아계 남성으로 성직 칼라를 착용하고 불붙은 초를 들고 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "네 인물 모두 발이 바닥에 닿아 있고 떠 있는 신체는 없다. 신부의 손이 초를 잡고, 앰버의 손이 입가 천을 지탱한다. 가구는 바닥에 놓여 있다. 현우의 비대칭 체중 배분은 가능하지만, 찰리와 신부는 비교적 곧게 서 있어 타격음에 어깨를 움츠린 순간의 표현이 약하다."
       },
       {
        "label": "B",
        "direction": "신부와 현우는 왼쪽 계단의 화면 밖 상부를 뚜렷하게 올려다본다. 앰버도 몸을 숙인 채 왼쪽 위로 눈길을 보낸다. 찰리는 얼굴이 왼쪽을 향하지만 시선은 거의 수평이어서 위층을 올려다보는 공통 반응에서 벗어난다. 십자가는 시선의 목표가 아닌 배경 요소로 남는다.",
        "built_space": "왼쪽 전경의 계단 한 줄이 화면 밖 위쪽으로 이어지고, 네 인물은 그 오른쪽 기도실 바닥에 배치되어 있다. 뒤에는 책상 하나와 등받이 의자 두 개, 오른쪽 오르간 하나와 작은 좌석 하나, 벽 십자가 하나, 왼쪽 책장 하나가 보인다. 천장등 하나, 왼쪽 벽등 하나와 오른쪽의 일부 가려진 벽등이 있으며, 오른쪽 아래 수납함도 참조 장소에 대응한다. 책상과 오르간은 인물 뒤에 비스듬히 보이고 실제적인 크기를 유지한다.",
        "entities": "네 인물만 등장한다. 현우는 참조에 가까운 앳된 동아시아계 남성으로 헝클어진 검은 머리, 회색 셔츠와 녹색 바지를 갖추고 있으며 얼굴 상처와 다리 붕대가 보인다. 별도의 겉옷은 없다. 앰버는 금발의 어린 여자아이로 읽히며 겉옷과 입을 가린 천을 유지하지만 참조의 작업복과 머리 장비는 다르다. 찰리는 긴 팔·짧은 다리·베이지 장갑판·흰 기계형 얼굴·낡은 코트와 모자가 참조에 가깝다. 신부는 나이 든 동아시아계 남성으로 칼라와 십자가 목걸이를 착용하고 촛불을 들고 있다. 다만 찰리의 화면 오른쪽 어깨 장갑판에는 읽을 수 있는 검은 숫자 표기가 있다.",
        "hard_violations": [
         "찰리의 어깨 장갑판에 판독 가능한 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "모든 인물의 발이 바닥에 닿아 있다. 현우는 무릎을 굽히고 체중을 나누며 한 손으로 반대쪽 팔을 잡고, 앰버는 몸을 웅크린 채 손으로 입가 천을 붙든다. 신부의 손이 초를 지탱하며 찰리도 넓은 두 발로 서 있다. 지지 없이 떠 있는 신체나 물체는 없다. 현우와 앰버의 자세는 놀라 움츠린 순간으로 자연스럽지만 찰리의 반응은 덜 명확하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "왼쪽 상승 계단과 일행의 전경은 맞지만, 앰버가 너무 성숙하고 찰리의 체형·현우의 겉옷 상태 및 일제히 움츠리는 반응이 부정확하다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물 정체성·부상·움츠린 자세와 장소 재현은 더 충실하지만, 찰리 어깨의 읽을 수 있는 숫자 표기가 명시적인 무문자 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 신부는 왼쪽 계단 위쪽의 화면 밖을 올려다본다. 앰버도 왼쪽을 보지만 눈길의 상승 각도는 약하다. 찰리는 얼굴을 거의 수평으로 왼쪽에 향하고 있어, 네 명 모두가 위층의 타격음에 일제히 올려다보는 동작은 완성되지 않는다. 배경 십자가를 바라보는 사람은 없다.",
        "built_space": "왼쪽 전경의 계단 한 줄이 위로 이어져 화면 밖으로 나간다. 일행은 계단 오른쪽의 기도실 바닥에 서 있다. 뒤에는 책상 하나, 등받이 의자 두 개, 오른쪽 오르간 하나와 작은 좌석 하나, 벽 십자가 하나, 왼쪽 책장 하나가 보인다. 천장등 하나와 왼쪽 벽등 하나가 보이며, 낡은 벽과 콘크리트 바닥은 참조 장소와 대체로 맞는다. 넓은 전신 구도와 배경 가구의 상대적 크기는 적절하다.",
        "entities": "추가 인물 없이 네 인물이 있다. 현우는 검은 머리의 젊은 동아시아계 남성이고 얼굴에 상처 흔적이 있지만, 티셔츠 위에 열린 겉셔츠를 입고 있으며 다리 부상은 드러나지 않는다. 앰버는 금발이지만 10세 아동보다 훨씬 성숙한 체격과 얼굴이며 참조 의상도 다르다. 손으로 입가의 천을 붙잡고 있다. 찰리는 낡은 코트와 모자, 베이지색 장갑과 흰 가면을 갖췄지만, 긴 다리와 보통 길이의 팔 때문에 참조의 육중한 고릴라형 체형과 크게 다르다. 신부는 나이 든 동아시아계 남성으로 성직 칼라를 착용하고 불붙은 초를 들고 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "네 인물 모두 발이 바닥에 닿아 있고 떠 있는 신체는 없다. 신부의 손이 초를 잡고, 앰버의 손이 입가 천을 지탱한다. 가구는 바닥에 놓여 있다. 현우의 비대칭 체중 배분은 가능하지만, 찰리와 신부는 비교적 곧게 서 있어 타격음에 어깨를 움츠린 순간의 표현이 약하다."
       },
       {
        "label": "A",
        "direction": "신부와 현우는 왼쪽 계단의 화면 밖 상부를 뚜렷하게 올려다본다. 앰버도 몸을 숙인 채 왼쪽 위로 눈길을 보낸다. 찰리는 얼굴이 왼쪽을 향하지만 시선은 거의 수평이어서 위층을 올려다보는 공통 반응에서 벗어난다. 십자가는 시선의 목표가 아닌 배경 요소로 남는다.",
        "built_space": "왼쪽 전경의 계단 한 줄이 화면 밖 위쪽으로 이어지고, 네 인물은 그 오른쪽 기도실 바닥에 배치되어 있다. 뒤에는 책상 하나와 등받이 의자 두 개, 오른쪽 오르간 하나와 작은 좌석 하나, 벽 십자가 하나, 왼쪽 책장 하나가 보인다. 천장등 하나, 왼쪽 벽등 하나와 오른쪽의 일부 가려진 벽등이 있으며, 오른쪽 아래 수납함도 참조 장소에 대응한다. 책상과 오르간은 인물 뒤에 비스듬히 보이고 실제적인 크기를 유지한다.",
        "entities": "네 인물만 등장한다. 현우는 참조에 가까운 앳된 동아시아계 남성으로 헝클어진 검은 머리, 회색 셔츠와 녹색 바지를 갖추고 있으며 얼굴 상처와 다리 붕대가 보인다. 별도의 겉옷은 없다. 앰버는 금발의 어린 여자아이로 읽히며 겉옷과 입을 가린 천을 유지하지만 참조의 작업복과 머리 장비는 다르다. 찰리는 긴 팔·짧은 다리·베이지 장갑판·흰 기계형 얼굴·낡은 코트와 모자가 참조에 가깝다. 신부는 나이 든 동아시아계 남성으로 칼라와 십자가 목걸이를 착용하고 촛불을 들고 있다. 다만 찰리의 화면 오른쪽 어깨 장갑판에는 읽을 수 있는 검은 숫자 표기가 있다.",
        "hard_violations": [
         "찰리의 어깨 장갑판에 판독 가능한 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "모든 인물의 발이 바닥에 닿아 있다. 현우는 무릎을 굽히고 체중을 나누며 한 손으로 반대쪽 팔을 잡고, 앰버는 몸을 웅크린 채 손으로 입가 천을 붙든다. 신부의 손이 초를 지탱하며 찰리도 넓은 두 발로 서 있다. 지지 없이 떠 있는 신체나 물체는 없다. 현우와 앰버의 자세는 놀라 움츠린 순간으로 자연스럽지만 찰리의 반응은 덜 명확하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.8,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.55,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gpt-high] 찰리의 어깨 장갑판에 판독 가능한 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1550,
   "B": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1550,
    "verdict_ko": "찰리가 정면을 응시하여 '일제히 위를 올려다보는' 행동 묘사에서 다소 아쉬움이 있으나, 찰리의 육중한 고릴라형 체형, 현우의 다친 다리(붕대) 등 캐릭터의 핵심 디테일과 공간 분위기를 완벽하게 재현하여 우수한 결과를 냈습니다.  ★위반: [gpt-high] 찰리의 어깨 장갑판에 판독 가능한 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "모든 인물이 위를 쳐다보는 시선 방향은 맞췄으나, 찰리가 프롬프트에 명시된 '고릴라형 몸체'를 완전히 상실하고 일반적인 인간 체형으로 잘못 렌더링되었으며 현우의 다친 다리 묘사도 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L180B02.png",
    "asset_id": "973d81a2-fcfc-4412-82af-c60138fde795",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-2c3b-7eb9-a08b-ce3e14c9fb45",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6__bgfirst_bg.png",
   "bg_asset_id": "4df963a5-17a9-4cdd-a2c1-6834d045d891",
   "bg_record_key": "S29sh6::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S29sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:16:47.904649+00:00",
  "fingerprint": "89f0653eaff4df82e9531686a92e32d70454665bfa2530662a34448fa45f6f39",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S29sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S29sh6_sel.png",
  "source_sha256": "37c8df978587205a371328f4e61272455f4ccedb12cb0dc8739cbd969486b69d",
  "file": "S29sh6_cine.png",
  "staged_sha256": "fea40e3414d2aeb3fcb581c055d8ac9838dac00a6d5e0f937450f2f3de384bf7",
  "latency_ms": 9664
 },
 "S29sh10::signage": {
  "fp": "48ac4b585de92106",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S29sh10::bgfirst_bg": {
  "input_fingerprint": "dc99fe6c7d163142",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh10__bgfirst_bg.png",
  "asset_id": "bd7dfe1d-fc8f-4e44-ac36-c3cfb142c309",
  "input_asset_ids": [
   "552fe877-7268-4883-a1c6-0d965dfe77b6",
   "db270a75-23c6-4def-abb3-d877c7750bce"
  ]
 },
 "S29sh10": {
  "input_fingerprint": "86e5183f59ea3e31",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The church entrance has been opened, and its small viewing window is open. The basement furnishings remain unchanged, with Charlie still downstairs in the old coat and hat. 신부: He is outside the church entrance, wearing his clerical collar. 구도환: He stands outside the church entrance, engaged in conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The church entrance has been opened, and its small viewing window is open. The basement furnishings remain unchanged, with Charlie still downstairs in the old coat and hat. 신부: He is outside the church entrance, wearing his clerical collar. 구도환: He stands outside the church entrance, engaged in conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The church entrance has been opened, and its small viewing window is open. The basement furnishings remain unchanged, with Charlie still downstairs in the old coat and hat. 신부: He is outside the church entrance, wearing his clerical collar. 구도환: He stands outside the church entrance, engaged in conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh10__bgfirst_bg.png",
     "asset_id": "bd7dfe1d-fc8f-4e44-ac36-c3cfb142c309",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S29sh10.png",
     "asset_id": "552fe877-7268-4883-a1c6-0d965dfe77b6",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:863986>",
     "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1023014>",
     "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L182B01.png",
     "asset_id": "db270a75-23c6-4def-abb3-d877c7750bce",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:863986>",
     "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1023014>",
     "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시점에서 두 인물(신부와 구도환)이 서로를 마주보며 대화하는 방향성이 바르게 연출됨.",
    "built_space": "전경에 관찰창 프레임이 위치하나, 창 너머 풍경 왼쪽에 실내 벽면(흰색/파란색 투톤)과 경첩이 달린 열린 문짝이 또다시 나타나 공간적으로 불가능한 구조임.",
    "entities": "신부(검은 셔츠)와 구도환(야상 코트)의 외형 및 의상이 레퍼런스와 일치함.",
    "hard_violations": [
     "[gemini-pro] physically impossible staging (출입문의 관찰창 너머로 실내 복도 벽과 문짝이 중복해서 보이는 구조적 모순)"
    ],
    "physics": "두 사람 모두 지면에 안정적으로 서서 체중을 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "신부와 구도환이 서로 마주보는 시선 처리가 확인됨.",
    "built_space": "전경의 사각형 창틀 너머로 실내 벽과 문의 가장자리가 보이는 물리적으로 불가능한 배경 겹침 오류가 동일하게 발생함.",
    "entities": "구도환은 레퍼런스와 일치하나, 신부가 검은 셔츠 대신 회색 체크무늬 셔츠를 입고 있어 설정에 어긋남.",
    "hard_violations": [
     "[gemini-pro] physically impossible staging (관찰창 안쪽 시야에 실내 벽면과 출입문이 다시 나타나는 물리적 오류)"
    ],
    "physics": "인물들 모두 바닥에 정상적으로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "관찰창 너머로 실내 복도와 문이 다시 등장하는 치명적인 공간 구조 오류가 있으나, 두 인물의 의상과 외형은 레퍼런스를 비교적 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "A와 동일한 물리적 공간 배치 오류가 발생했으며, 신부의 의상이 레퍼런스(검은 셔츠)와 전혀 다른 셔츠로 잘못 묘사되어 완성도가 더 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시점에서 두 인물(신부와 구도환)이 서로를 마주보며 대화하는 방향성이 바르게 연출됨.",
        "built_space": "전경에 관찰창 프레임이 위치하나, 창 너머 풍경 왼쪽에 실내 벽면(흰색/파란색 투톤)과 경첩이 달린 열린 문짝이 또다시 나타나 공간적으로 불가능한 구조임.",
        "entities": "신부(검은 셔츠)와 구도환(야상 코트)의 외형 및 의상이 레퍼런스와 일치함.",
        "hard_violations": [
         "physically impossible staging (출입문의 관찰창 너머로 실내 복도 벽과 문짝이 중복해서 보이는 구조적 모순)"
        ],
        "physics": "두 사람 모두 지면에 안정적으로 서서 체중을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "신부와 구도환이 서로 마주보는 시선 처리가 확인됨.",
        "built_space": "전경의 사각형 창틀 너머로 실내 벽과 문의 가장자리가 보이는 물리적으로 불가능한 배경 겹침 오류가 동일하게 발생함.",
        "entities": "구도환은 레퍼런스와 일치하나, 신부가 검은 셔츠 대신 회색 체크무늬 셔츠를 입고 있어 설정에 어긋남.",
        "hard_violations": [
         "physically impossible staging (관찰창 안쪽 시야에 실내 벽면과 출입문이 다시 나타나는 물리적 오류)"
        ],
        "physics": "인물들 모두 바닥에 정상적으로 서 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "관찰창 너머로 실내 복도와 문이 다시 등장하는 치명적인 공간 구조 오류가 있으나, 두 인물의 의상과 외형은 레퍼런스를 비교적 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "A와 동일한 물리적 공간 배치 오류가 발생했으며, 신부의 의상이 레퍼런스(검은 셔츠)와 전혀 다른 셔츠로 잘못 묘사되어 완성도가 더 떨어집니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시점에서 두 인물(신부와 구도환)이 서로를 마주보며 대화하는 방향성이 바르게 연출됨.",
        "built_space": "전경에 관찰창 프레임이 위치하나, 창 너머 풍경 왼쪽에 실내 벽면(흰색/파란색 투톤)과 경첩이 달린 열린 문짝이 또다시 나타나 공간적으로 불가능한 구조임.",
        "entities": "신부(검은 셔츠)와 구도환(야상 코트)의 외형 및 의상이 레퍼런스와 일치함.",
        "hard_violations": [
         "physically impossible staging (출입문의 관찰창 너머로 실내 복도 벽과 문짝이 중복해서 보이는 구조적 모순)"
        ],
        "physics": "두 사람 모두 지면에 안정적으로 서서 체중을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "신부와 구도환이 서로 마주보는 시선 처리가 확인됨.",
        "built_space": "전경의 사각형 창틀 너머로 실내 벽과 문의 가장자리가 보이는 물리적으로 불가능한 배경 겹침 오류가 동일하게 발생함.",
        "entities": "구도환은 레퍼런스와 일치하나, 신부가 검은 셔츠 대신 회색 체크무늬 셔츠를 입고 있어 설정에 어긋남.",
        "hard_violations": [
         "physically impossible staging (관찰창 안쪽 시야에 실내 벽면과 출입문이 다시 나타나는 물리적 오류)"
        ],
        "physics": "인물들 모두 바닥에 정상적으로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "서로 마주 보는 야간 대화는 맞지만, 관찰창이 참조보다 정사각형에 가까워지고 허벅지까지 드러나는 넓은 구도로 은밀한 제한 시야와 미디엄 쇼트에서 멀어진다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "상체 크롭은 미디엄보다 다소 타이트하지만, 중앙의 열린 가로형 관찰창과 덮개, 문밖에서 서로 바라보는 두 인물로 지정된 은밀한 시점과 장소를 더 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 신부는 오른쪽 구도환의 얼굴을, 오른쪽 구도환은 왼쪽 신부의 얼굴을 바라본다. 몸도 서로 마주 향하며 카메라를 보는 사람은 없다. 무기나 이동 동작은 없다.",
        "built_space": "전경 중앙에 금속 관찰창 하나와 하단 걸쇠 하나가 보인다. 창 너머 왼쪽에는 출입문틀과 잠금 부품이 있고, 두 사람은 그 바깥에 서 있다. 외부의 콘크리트 담장, 철제 울타리, 가로등과 젖은 바닥은 참조와 유사하다. 다만 관찰창은 참조의 낮고 긴 가로형보다 훨씬 높은 직사각형이며, 열린 덮개는 화면에서 확인되지 않는다. 인물은 머리부터 허벅지 부근까지 보여 지정된 미디엄보다 넓게 잡혔다.",
        "entities": "보이는 사람은 신부와 구도환 두 명뿐이며, 각각 나이 든 동아시아계 남성과 검은 머리의 중년 동아시아계 남성으로 설정에 부합한다. 신부의 희끗한 머리와 주름, 구도환의 얼굴과 털 안감 외투는 참조와 대체로 맞는다. 신부의 흰 성직자 칼라는 보이지만 셔츠는 참조의 낡은 검정 셔츠보다 밝은 회갈색이고, 가슴의 작은 장식은 참조의 굵은 십자가 형태로 명확히 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통과 보이는 다리는 수직으로 이어져 평범하게 서 있는 자세다. 발의 지면 접촉은 하단 창틀에 가려 확인할 수 없지만 공중에 떠 있다는 단서는 없다. 손은 팔에 자연스럽게 이어져 몸 옆에 내려와 있고, 금속 창틀과 걸쇠는 문에 고정되어 있다."
       },
       {
        "label": "B",
        "direction": "왼쪽 신부는 오른쪽 아래 방향의 구도환 얼굴을 보고, 구도환은 왼쪽 위의 신부 얼굴을 바라본다. 구도환의 얼굴이 카메라 쪽으로 일부 열려 있지만 시선의 대상은 신부다. 두 사람의 대화 관계가 분명하며 이동하거나 겨누는 물체는 없다.",
        "built_space": "전경 중앙에 가로로 긴 관찰창 하나, 위로 열린 금속 덮개 하나, 하단 걸쇠 하나가 있다. 참조의 관찰창 구조를 A보다 잘 유지한다. 그 너머 양쪽 출입구 가장자리와 문밖에 선 두 사람이 보이고, 뒤에는 콘크리트 담장과 철제 울타리, 가로등, 나무와 건물 창문이 배치되어 있다. 내부에서 작은 개구부를 통해 외부를 보는 시점이 명확하다. 인물은 가슴 부근에서 잘려 지정된 미디엄보다 약간 타이트하다.",
        "entities": "추가 인물 없이 신부와 구도환 두 명만 보인다. 신부의 회색 섞인 머리, 나이 든 옆얼굴과 검은 셔츠, 구도환의 검은 머리와 중년 얼굴, 검은 목 부분 의복과 털 안감 외투가 참조에 가깝다. 신부는 등과 옆면을 보여 흰 칼라 전면과 십자가를 확인하기 어렵고, 외투 허리끈은 크롭 밖이다. 이를 누락으로 단정할 수 없다. 읽을 수 있는 글자나 표시도 없다.",
        "hard_violations": [],
        "physics": "두 사람은 문밖에서 상체를 세우고 자연스럽게 서로를 향한다. 하체와 발은 관찰창 아래에 가려져 직접적인 지면 접촉은 보이지 않지만, 부유나 불가능한 자세의 징후는 없다. 열린 덮개는 상단 연결부에 붙어 있고 창틀과 걸쇠도 문에 고정되어 있어 지지 관계가 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "서로 마주 보는 야간 대화는 맞지만, 관찰창이 참조보다 정사각형에 가까워지고 허벅지까지 드러나는 넓은 구도로 은밀한 제한 시야와 미디엄 쇼트에서 멀어진다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "상체 크롭은 미디엄보다 다소 타이트하지만, 중앙의 열린 가로형 관찰창과 덮개, 문밖에서 서로 바라보는 두 인물로 지정된 은밀한 시점과 장소를 더 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 신부는 오른쪽 구도환의 얼굴을, 오른쪽 구도환은 왼쪽 신부의 얼굴을 바라본다. 몸도 서로 마주 향하며 카메라를 보는 사람은 없다. 무기나 이동 동작은 없다.",
        "built_space": "전경 중앙에 금속 관찰창 하나와 하단 걸쇠 하나가 보인다. 창 너머 왼쪽에는 출입문틀과 잠금 부품이 있고, 두 사람은 그 바깥에 서 있다. 외부의 콘크리트 담장, 철제 울타리, 가로등과 젖은 바닥은 참조와 유사하다. 다만 관찰창은 참조의 낮고 긴 가로형보다 훨씬 높은 직사각형이며, 열린 덮개는 화면에서 확인되지 않는다. 인물은 머리부터 허벅지 부근까지 보여 지정된 미디엄보다 넓게 잡혔다.",
        "entities": "보이는 사람은 신부와 구도환 두 명뿐이며, 각각 나이 든 동아시아계 남성과 검은 머리의 중년 동아시아계 남성으로 설정에 부합한다. 신부의 희끗한 머리와 주름, 구도환의 얼굴과 털 안감 외투는 참조와 대체로 맞는다. 신부의 흰 성직자 칼라는 보이지만 셔츠는 참조의 낡은 검정 셔츠보다 밝은 회갈색이고, 가슴의 작은 장식은 참조의 굵은 십자가 형태로 명확히 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통과 보이는 다리는 수직으로 이어져 평범하게 서 있는 자세다. 발의 지면 접촉은 하단 창틀에 가려 확인할 수 없지만 공중에 떠 있다는 단서는 없다. 손은 팔에 자연스럽게 이어져 몸 옆에 내려와 있고, 금속 창틀과 걸쇠는 문에 고정되어 있다."
       },
       {
        "label": "A",
        "direction": "왼쪽 신부는 오른쪽 아래 방향의 구도환 얼굴을 보고, 구도환은 왼쪽 위의 신부 얼굴을 바라본다. 구도환의 얼굴이 카메라 쪽으로 일부 열려 있지만 시선의 대상은 신부다. 두 사람의 대화 관계가 분명하며 이동하거나 겨누는 물체는 없다.",
        "built_space": "전경 중앙에 가로로 긴 관찰창 하나, 위로 열린 금속 덮개 하나, 하단 걸쇠 하나가 있다. 참조의 관찰창 구조를 A보다 잘 유지한다. 그 너머 양쪽 출입구 가장자리와 문밖에 선 두 사람이 보이고, 뒤에는 콘크리트 담장과 철제 울타리, 가로등, 나무와 건물 창문이 배치되어 있다. 내부에서 작은 개구부를 통해 외부를 보는 시점이 명확하다. 인물은 가슴 부근에서 잘려 지정된 미디엄보다 약간 타이트하다.",
        "entities": "추가 인물 없이 신부와 구도환 두 명만 보인다. 신부의 회색 섞인 머리, 나이 든 옆얼굴과 검은 셔츠, 구도환의 검은 머리와 중년 얼굴, 검은 목 부분 의복과 털 안감 외투가 참조에 가깝다. 신부는 등과 옆면을 보여 흰 칼라 전면과 십자가를 확인하기 어렵고, 외투 허리끈은 크롭 밖이다. 이를 누락으로 단정할 수 없다. 읽을 수 있는 글자나 표시도 없다.",
        "hard_violations": [],
        "physics": "두 사람은 문밖에서 상체를 세우고 자연스럽게 서로를 향한다. 하체와 발은 관찰창 아래에 가려져 직접적인 지면 접촉은 보이지 않지만, 부유나 불가능한 자세의 징후는 없다. 열린 덮개는 상단 연결부에 붙어 있고 창틀과 걸쇠도 문에 고정되어 있어 지지 관계가 자연스럽다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.417
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible staging (출입문의 관찰창 너머로 실내 복도 벽과 문짝이 중복해서 보이는 구조적 모순)"
    ],
    "B": [
     "[gemini-pro] physically impossible staging (관찰창 안쪽 시야에 실내 벽면과 출입문이 다시 나타나는 물리적 오류)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "관찰창 너머로 실내 복도와 문이 다시 등장하는 치명적인 공간 구조 오류가 있으나, 두 인물의 의상과 외형은 레퍼런스를 비교적 잘 구현했습니다.  ★위반: [gemini-pro] physically impossible staging (출입문의 관찰창 너머로 실내 복도 벽과 문짝이 중복해서 보이는 구조적 모순)"
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "A와 동일한 물리적 공간 배치 오류가 발생했으며, 신부의 의상이 레퍼런스(검은 셔츠)와 전혀 다른 셔츠로 잘못 묘사되어 완성도가 더 떨어집니다.  ★위반: [gemini-pro] physically impossible staging (관찰창 안쪽 시야에 실내 벽면과 출입문이 다시 나타나는 물리적 오류)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L182B01.png",
    "asset_id": "db270a75-23c6-4def-abb3-d877c7750bce",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1023014>",
    "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-2f95-73ef-a3de-de1495e3e607",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh10__bgfirst_bg.png",
   "bg_asset_id": "bd7dfe1d-fc8f-4e44-ac36-c3cfb142c309",
   "bg_record_key": "S29sh10::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S29sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:18:17.900669+00:00",
  "fingerprint": "e992c27cc9a53adb0395402e18736a28fbee0eebe5071e09076c2804c07e6372",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S29sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S29sh10_sel.png",
  "source_sha256": "924b1f39522d1921906bb8a6bee5b8d66c2adacbcb86a99cbe9b26606cffaec7",
  "file": "S29sh10_cine.png",
  "staged_sha256": "a590442bf3e64ab7cbce8a2fcf22c18f5c2251d92698cc20df5fe6dda1a7519a",
  "latency_ms": 8663
 },
 "S29sh12::signage": {
  "fp": "4d57f2d70cc579d7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S29sh12": {
  "input_fingerprint": "92a79f19580a8fb2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서치라이트 불빛 주변의 어두운 하늘을 벌떼처럼 빽빽하게 덮은 채 앞으로 향한 비행 자세로 허공에 뜬 소형 드론들의 실루엣.\n\nLOCATION (lock): In the night sky above the refugee settlement, where small drones fly around sweeping helicopter searchlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Night sky (Dark behind the airborne drones); used as Negative space between clusters preserves the silhouettes' readability; Small drone swarm (Densely clustered and flying forward along the searchlight) — Undersides and oblique side profiles are visible, with their forward axes directed toward screen right; used as Distributed silhouettes convey overwhelming numbers without enlarging any single drone.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The helicopter's searchlight separates the drone silhouettes from the dark night sky with controlled contrast, without adding a dreamlike treatment to the flashback.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Helicopter searchlights sweep the night sky, with several drones flying along the beams. The church basement still contains the cross, desk, chairs and organ, and Charlie remains there in his old coat and hat.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서치라이트 불빛 주변의 어두운 하늘을 벌떼처럼 빽빽하게 덮은 채 앞으로 향한 비행 자세로 허공에 뜬 소형 드론들의 실루엣.\n\nLOCATION (lock): In the night sky above the refugee settlement, where small drones fly around sweeping helicopter searchlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Night sky (Dark behind the airborne drones); used as Negative space between clusters preserves the silhouettes' readability; Small drone swarm (Densely clustered and flying forward along the searchlight) — Undersides and oblique side profiles are visible, with their forward axes directed toward screen right; used as Distributed silhouettes convey overwhelming numbers without enlarging any single drone.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The helicopter's searchlight separates the drone silhouettes from the dark night sky with controlled contrast, without adding a dreamlike treatment to the flashback.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Helicopter searchlights sweep the night sky, with several drones flying along the beams. The church basement still contains the cross, desk, chairs and organ, and Charlie remains there in his old coat and hat.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서치라이트 불빛 주변의 어두운 하늘을 벌떼처럼 빽빽하게 덮은 채 앞으로 향한 비행 자세로 허공에 뜬 소형 드론들의 실루엣.\n\nLOCATION (lock): In the night sky above the refugee settlement, where small drones fly around sweeping helicopter searchlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Night sky (Dark behind the airborne drones); used as Negative space between clusters preserves the silhouettes' readability; Small drone swarm (Densely clustered and flying forward along the searchlight) — Undersides and oblique side profiles are visible, with their forward axes directed toward screen right; used as Distributed silhouettes convey overwhelming numbers without enlarging any single drone.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The helicopter's searchlight separates the drone silhouettes from the dark night sky with controlled contrast, without adding a dreamlike treatment to the flashback.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Helicopter searchlights sweep the night sky, with several drones flying along the beams. The church basement still contains the cross, desk, chairs and organ, and Charlie remains there in his old coat and hat.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "헬리콥터 탐조등이 우측 하단을 향해 빛을 비추고 있으며, 드론들의 전면 축이 화면 우측을 향해 비행하고 있음.",
    "built_space": "화면 최하단에 난민촌의 구조물 지붕들이 배치되어 있음.",
    "entities": "1대의 헬리콥터, 1개의 탐조등 불빛, 벌떼처럼 빽빽하게 밀집된 다수의 소형 쿼드코프터 드론 실루엣이 보임.",
    "hard_violations": [],
    "physics": "헬리콥터와 드론들 모두 로터의 회전을 통한 양력으로 공중에 안정적으로 떠 있음."
   },
   {
    "label": "B",
    "direction": "화면 상단 양쪽에서 두 개의 탐조등이 중앙 하단을 향해 교차하며 빛을 비추고 있으며, 드론들은 대체로 우측이나 정면을 향함.",
    "built_space": "화면 하단에 넓은 면적의 난민촌 건물과 골목들이 위치해 있음.",
    "entities": "화면 상단 양 끝에 2대의 헬리콥터(일부), 2개의 탐조등 불빛, 다수의 소형 쿼드코프터 드론들이 존재함.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
    ],
    "physics": "헬리콥터와 드론들 모두 양력을 받아 허공에 떠 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "단일 헬리콥터 조명과 우측을 향해 빽빽하게 밀집된 드론 떼를 프롬프트와 레퍼런스에 맞게 정확히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "레퍼런스와 설정에 없는 두 번째 헬리콥터와 교차하는 탐조등 불빛을 임의로 복제하여 추가한 점이 치명적인 오류입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "헬리콥터 탐조등이 우측 하단을 향해 빛을 비추고 있으며, 드론들의 전면 축이 화면 우측을 향해 비행하고 있음.",
        "built_space": "화면 최하단에 난민촌의 구조물 지붕들이 배치되어 있음.",
        "entities": "1대의 헬리콥터, 1개의 탐조등 불빛, 벌떼처럼 빽빽하게 밀집된 다수의 소형 쿼드코프터 드론 실루엣이 보임.",
        "hard_violations": [],
        "physics": "헬리콥터와 드론들 모두 로터의 회전을 통한 양력으로 공중에 안정적으로 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 상단 양쪽에서 두 개의 탐조등이 중앙 하단을 향해 교차하며 빛을 비추고 있으며, 드론들은 대체로 우측이나 정면을 향함.",
        "built_space": "화면 하단에 넓은 면적의 난민촌 건물과 골목들이 위치해 있음.",
        "entities": "화면 상단 양 끝에 2대의 헬리콥터(일부), 2개의 탐조등 불빛, 다수의 소형 쿼드코프터 드론들이 존재함.",
        "hard_violations": [
         "레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
        ],
        "physics": "헬리콥터와 드론들 모두 양력을 받아 허공에 떠 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "단일 헬리콥터 조명과 우측을 향해 빽빽하게 밀집된 드론 떼를 프롬프트와 레퍼런스에 맞게 정확히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "레퍼런스와 설정에 없는 두 번째 헬리콥터와 교차하는 탐조등 불빛을 임의로 복제하여 추가한 점이 치명적인 오류입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "헬리콥터 탐조등이 우측 하단을 향해 빛을 비추고 있으며, 드론들의 전면 축이 화면 우측을 향해 비행하고 있음.",
        "built_space": "화면 최하단에 난민촌의 구조물 지붕들이 배치되어 있음.",
        "entities": "1대의 헬리콥터, 1개의 탐조등 불빛, 벌떼처럼 빽빽하게 밀집된 다수의 소형 쿼드코프터 드론 실루엣이 보임.",
        "hard_violations": [],
        "physics": "헬리콥터와 드론들 모두 로터의 회전을 통한 양력으로 공중에 안정적으로 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 상단 양쪽에서 두 개의 탐조등이 중앙 하단을 향해 교차하며 빛을 비추고 있으며, 드론들은 대체로 우측이나 정면을 향함.",
        "built_space": "화면 하단에 넓은 면적의 난민촌 건물과 골목들이 위치해 있음.",
        "entities": "화면 상단 양 끝에 2대의 헬리콥터(일부), 2개의 탐조등 불빛, 다수의 소형 쿼드코프터 드론들이 존재함.",
        "hard_violations": [
         "레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
        ],
        "physics": "헬리콥터와 드론들 모두 양력을 받아 허공에 떠 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "하단의 정착촌이 화면 약 3분의 1을 차지해 하늘과 드론 군집 중심의 구도를 약화하며, 드론들도 화면 오른쪽으로 전진하기보다 카메라 쪽을 향해 떠 있는 모습이다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "어두운 하늘을 넓게 채운 군집과 대각선 서치라이트가 요구 구도 및 참조에 더 충실하지만, 드론들의 오른쪽 전진 방향은 불명확하고 우상단 한 기가 유독 크게 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "좌상단 서치라이트는 오른쪽 아래로, 우상단 서치라이트는 왼쪽 아래로 향해 중앙 군집과 그 아래 공간을 비춘다. 드론들은 대체로 좌우로 펼쳐진 팔과 정면의 하부 장치를 보여 카메라 쪽을 향한 자세로 읽힌다. 일부 사선 기체는 있으나 군집 전체의 전방 축이 화면 오른쪽을 향하지 않으며, 두 광선에도 단일한 진행 방향이 없다.",
        "built_space": "화면 하단 약 3분의 1에 골함석 지붕, 방수포, 창문과 통로가 있는 임시 주거 건물 수십 채가 보인다. 상단 양쪽에는 잘린 헬리콥터 동체 일부와 서치라이트가 각각 하나씩 있다. 참조 사진에는 정착촌 건축물이 보이지 않아 이 건물들의 정확한 장소 일치 여부는 검증할 수 없다. 사람이나 반사는 없으며, 하늘 중심이어야 할 화면에서 건축물의 비중이 크다.",
        "entities": "소형 다중회전익 드론 수십 기, 헬리콥터 일부 두 대, 서치라이트 두 줄기, 어두운 밤하늘과 정착촌이 보인다. 드론의 몸체와 하부 장치는 실제 기계 형태로 표현됐고 빛 앞에서 실루엣이 드러난다. 사람, 얼굴, 읽을 수 있는 글자는 없다. 교회 지하실과 그 안의 인물·물건은 이 하늘 장면의 프레임 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "드론들은 회전익과 일부 회전 흐림이 보여 공중 체류를 지탱할 추진 장치가 있다. 수평에 가까운 자세는 호버링으로 가능하지만 오른쪽으로 전진하는 기울기는 뚜렷하지 않다. 잘린 헬리콥터의 로터 전체는 확인하기 어렵지만 기체 일부만 보이는 구도 자체가 무지지 부유를 뜻하지는 않는다. 건물은 지면에 놓여 있고 광선은 상단의 실제 광원에서 퍼진다."
       },
       {
        "label": "B",
        "direction": "좌상단 헬리콥터의 서치라이트가 화면 오른쪽 아래로 뻗어 중앙의 드론 무리를 비춘다. 군집 분포는 광선의 대각선 흐름을 따르지만, 개별 기체 대부분은 하부 카메라와 좌우 대칭에 가까운 팔을 정면으로 보여 오른쪽을 향한 전방 축이 확인되지 않는다. 우상단의 큰 드론 역시 주로 카메라 쪽 아래를 드러낸다.",
        "built_space": "정착촌 지붕은 화면 맨 아래의 좁은 띠에만 있고, 나머지 대부분은 드론과 구름 낀 밤하늘이다. 좌상단에 일부 잘린 헬리콥터 한 대와 서치라이트 하나가 보인다. 참조의 헬리콥터·광선·하늘 관계를 유지하면서 하늘을 주된 공간으로 삼는다. 참조에 건축 세부가 없어 하단 지붕의 정확한 일치는 검증할 수 없으며, 사람이나 반사는 없다.",
        "entities": "소형 다중회전익 드론 수십 기가 화면 전반에 분산되고, 헬리콥터 한 대와 밝은 서치라이트, 어두운 구름, 하단의 임시 주거 지붕들이 보인다. 드론은 참조처럼 팔, 회전익, 하부 장치를 갖춘 기계로 읽힌다. 다만 우상단 한 기는 다른 기체보다 현저히 커서 개별 드론을 강조하지 말라는 요구에 덜 맞는다. 사람, 얼굴, 판독 가능한 글자는 없고 지하실의 인물과 소품은 적절히 화면 밖에 있다.",
        "hard_violations": [],
        "physics": "드론의 회전익과 회전 흐림, 헬리콥터의 주회전익이 보여 각 항공기의 부양 수단을 확인할 수 있다. 드론의 거의 수평인 자세는 물리적으로 가능한 비행 또는 호버링이지만 요구된 오른쪽 전진 동작을 명확히 보여주지는 않는다. 지붕은 건물 위에 놓여 있고 서치라이트는 헬리콥터 하부에서 시작해 대기 중으로 확산된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "하단의 정착촌이 화면 약 3분의 1을 차지해 하늘과 드론 군집 중심의 구도를 약화하며, 드론들도 화면 오른쪽으로 전진하기보다 카메라 쪽을 향해 떠 있는 모습이다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "어두운 하늘을 넓게 채운 군집과 대각선 서치라이트가 요구 구도 및 참조에 더 충실하지만, 드론들의 오른쪽 전진 방향은 불명확하고 우상단 한 기가 유독 크게 보인다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "좌상단 서치라이트는 오른쪽 아래로, 우상단 서치라이트는 왼쪽 아래로 향해 중앙 군집과 그 아래 공간을 비춘다. 드론들은 대체로 좌우로 펼쳐진 팔과 정면의 하부 장치를 보여 카메라 쪽을 향한 자세로 읽힌다. 일부 사선 기체는 있으나 군집 전체의 전방 축이 화면 오른쪽을 향하지 않으며, 두 광선에도 단일한 진행 방향이 없다.",
        "built_space": "화면 하단 약 3분의 1에 골함석 지붕, 방수포, 창문과 통로가 있는 임시 주거 건물 수십 채가 보인다. 상단 양쪽에는 잘린 헬리콥터 동체 일부와 서치라이트가 각각 하나씩 있다. 참조 사진에는 정착촌 건축물이 보이지 않아 이 건물들의 정확한 장소 일치 여부는 검증할 수 없다. 사람이나 반사는 없으며, 하늘 중심이어야 할 화면에서 건축물의 비중이 크다.",
        "entities": "소형 다중회전익 드론 수십 기, 헬리콥터 일부 두 대, 서치라이트 두 줄기, 어두운 밤하늘과 정착촌이 보인다. 드론의 몸체와 하부 장치는 실제 기계 형태로 표현됐고 빛 앞에서 실루엣이 드러난다. 사람, 얼굴, 읽을 수 있는 글자는 없다. 교회 지하실과 그 안의 인물·물건은 이 하늘 장면의 프레임 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "드론들은 회전익과 일부 회전 흐림이 보여 공중 체류를 지탱할 추진 장치가 있다. 수평에 가까운 자세는 호버링으로 가능하지만 오른쪽으로 전진하는 기울기는 뚜렷하지 않다. 잘린 헬리콥터의 로터 전체는 확인하기 어렵지만 기체 일부만 보이는 구도 자체가 무지지 부유를 뜻하지는 않는다. 건물은 지면에 놓여 있고 광선은 상단의 실제 광원에서 퍼진다."
       },
       {
        "label": "A",
        "direction": "좌상단 헬리콥터의 서치라이트가 화면 오른쪽 아래로 뻗어 중앙의 드론 무리를 비춘다. 군집 분포는 광선의 대각선 흐름을 따르지만, 개별 기체 대부분은 하부 카메라와 좌우 대칭에 가까운 팔을 정면으로 보여 오른쪽을 향한 전방 축이 확인되지 않는다. 우상단의 큰 드론 역시 주로 카메라 쪽 아래를 드러낸다.",
        "built_space": "정착촌 지붕은 화면 맨 아래의 좁은 띠에만 있고, 나머지 대부분은 드론과 구름 낀 밤하늘이다. 좌상단에 일부 잘린 헬리콥터 한 대와 서치라이트 하나가 보인다. 참조의 헬리콥터·광선·하늘 관계를 유지하면서 하늘을 주된 공간으로 삼는다. 참조에 건축 세부가 없어 하단 지붕의 정확한 일치는 검증할 수 없으며, 사람이나 반사는 없다.",
        "entities": "소형 다중회전익 드론 수십 기가 화면 전반에 분산되고, 헬리콥터 한 대와 밝은 서치라이트, 어두운 구름, 하단의 임시 주거 지붕들이 보인다. 드론은 참조처럼 팔, 회전익, 하부 장치를 갖춘 기계로 읽힌다. 다만 우상단 한 기는 다른 기체보다 현저히 커서 개별 드론을 강조하지 말라는 요구에 덜 맞는다. 사람, 얼굴, 판독 가능한 글자는 없고 지하실의 인물과 소품은 적절히 화면 밖에 있다.",
        "hard_violations": [],
        "physics": "드론의 회전익과 회전 흐림, 헬리콥터의 주회전익이 보여 각 항공기의 부양 수단을 확인할 수 있다. 드론의 거의 수평인 자세는 물리적으로 가능한 비행 또는 호버링이지만 요구된 오른쪽 전진 동작을 명확히 보여주지는 않는다. 지붕은 건물 위에 놓여 있고 서치라이트는 헬리콥터 하부에서 시작해 대기 중으로 확산된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.143
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.893
   },
   "violations": {
    "B": [
     "[gemini-pro] 레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 893
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "단일 헬리콥터 조명과 우측을 향해 빽빽하게 밀집된 드론 떼를 프롬프트와 레퍼런스에 맞게 정확히 구현했습니다."
   },
   {
    "label": "B",
    "score": 893,
    "verdict_ko": "레퍼런스와 설정에 없는 두 번째 헬리콥터와 교차하는 탐조등 불빛을 임의로 복제하여 추가한 점이 치명적인 오류입니다.  ★위반: [gemini-pro] 레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L183B01.png",
    "asset_id": "b50a2636-aacd-45fa-8152-e123b37e1b11",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-32da-731a-bdd7-f55bd9082167",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S29sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:36:48.950919+00:00",
  "fingerprint": "0df4cd58968d7973cc869fd23b287a49a48d67b9f32c63b3690620a212cc7ddc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S29sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S29sh12_sel.png",
  "source_sha256": "d3655d86403ed2185c74bf9024ed511d1ee019d3182c1d374a3690e5f07df0d1",
  "file": "S29sh12_cine.png",
  "staged_sha256": "f07f8b48e7650b56c17a35555d0d8e1807e6dfe6564b91434c33bc0ae11a072f",
  "latency_ms": 17487
 },
 "S30sh5::signage": {
  "fp": "02a1ddaf0c2a122a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S30sh5": {
  "input_fingerprint": "b043e9bdce48ee78",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑 틈새로 눈부신 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업.\n\nLOCATION (lock): Beside the small organ and old radio inside the church basement prayer room, as lanterns brighten around the robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Brilliant blue light bursts through the chest armor gaps, retaining readable metal edges around the brightest emission.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small organ and the old broken radio remain in the basement, along with the cross, desk and chairs. Light is emerging from Charlie's chest; his old coat and hat, worn metal body and faded chest logo persist.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑 틈새로 눈부신 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업.\n\nLOCATION (lock): Beside the small organ and old radio inside the church basement prayer room, as lanterns brighten around the robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Brilliant blue light bursts through the chest armor gaps, retaining readable metal edges around the brightest emission.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small organ and the old broken radio remain in the basement, along with the cross, desk and chairs. Light is emerging from Charlie's chest; his old coat and hat, worn metal body and faded chest logo persist.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑 틈새로 눈부신 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업.\n\nLOCATION (lock): Beside the small organ and old radio inside the church basement prayer room, as lanterns brighten around the robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Brilliant blue light bursts through the chest armor gaps, retaining readable metal edges around the brightest emission.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small organ and the old broken radio remain in the basement, along with the cross, desk and chairs. Light is emerging from Charlie's chest; his old coat and hat, worn metal body and faded chest logo persist.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 찰리의 가슴을 정면에서 주시하며, 푸른 빛이 가슴에서 바깥으로 뿜어져 나옴.",
    "built_space": "배경에 피아노(오르간)가 배치되어 있으며, 지하실의 분위기가 나타남.",
    "entities": "찰리의 샌드 베이지색 금속 장갑, 트렌치코트, 벨트 디테일이 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "자연스럽게 서 있는 자세를 유지하고 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 찰리의 가슴을 향하며, 장갑 틈새로 푸른 빛이 발산됨.",
    "built_space": "배경 좌측에 십자가, 우측에 피아노가 보이며 지하실의 형태를 갖춤.",
    "entities": "찰리의 금속 장갑과 코트가 레퍼런스와 일치하나, 중앙부 묘사가 약간 다름.",
    "hard_violations": [],
    "physics": "안정적으로 서 있는 자세를 보여줌."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴 장갑 틈새로 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업을 사실적이고 강렬한 빛 반사로 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "클로즈업과 배경 요소는 좋으나, 빛이 강하게 뿜어져 나오기보다는 단순한 LED 조명선처럼 묘사되어 극적인 느낌이 부족합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 찰리의 가슴을 정면에서 주시하며, 푸른 빛이 가슴에서 바깥으로 뿜어져 나옴.",
        "built_space": "배경에 피아노(오르간)가 배치되어 있으며, 지하실의 분위기가 나타남.",
        "entities": "찰리의 샌드 베이지색 금속 장갑, 트렌치코트, 벨트 디테일이 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 자세를 유지하고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 찰리의 가슴을 향하며, 장갑 틈새로 푸른 빛이 발산됨.",
        "built_space": "배경 좌측에 십자가, 우측에 피아노가 보이며 지하실의 형태를 갖춤.",
        "entities": "찰리의 금속 장갑과 코트가 레퍼런스와 일치하나, 중앙부 묘사가 약간 다름.",
        "hard_violations": [],
        "physics": "안정적으로 서 있는 자세를 보여줌."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴 장갑 틈새로 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업을 사실적이고 강렬한 빛 반사로 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "클로즈업과 배경 요소는 좋으나, 빛이 강하게 뿜어져 나오기보다는 단순한 LED 조명선처럼 묘사되어 극적인 느낌이 부족합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 찰리의 가슴을 정면에서 주시하며, 푸른 빛이 가슴에서 바깥으로 뿜어져 나옴.",
        "built_space": "배경에 피아노(오르간)가 배치되어 있으며, 지하실의 분위기가 나타남.",
        "entities": "찰리의 샌드 베이지색 금속 장갑, 트렌치코트, 벨트 디테일이 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 자세를 유지하고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 찰리의 가슴을 향하며, 장갑 틈새로 푸른 빛이 발산됨.",
        "built_space": "배경 좌측에 십자가, 우측에 피아노가 보이며 지하실의 형태를 갖춤.",
        "entities": "찰리의 금속 장갑과 코트가 레퍼런스와 일치하나, 중앙부 묘사가 약간 다름.",
        "hard_violations": [],
        "physics": "안정적으로 서 있는 자세를 보여줌."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "가슴을 더 밀착해 잡고 원형 중심부보다 장갑판 틈새의 눈부신 청색 방출을 강조하여, 지시된 순간과 발광 위치를 가장 정확하게 구현했다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "가슴 클로즈업과 오르간 주변 공간은 충실하지만, 강한 원형 중심광이 시선을 지배해 장갑 틈새에서 빛이 뿜어지는 순간이라는 핵심이 A보다 덜 선명하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "푸른빛이 원형 가슴 부품 둘레와 가로·세로 장갑 이음새에서 바깥으로 새어 나와 코트 안쪽과 금속 모서리를 밝힌다. 원형 부품 중심은 비교적 어두워 발광의 출처가 틈새로 읽힌다. 턱 아래만 보여 시선 방향은 확인할 수 없으며, 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "왼쪽 배경에 벽 십자가 하나와 낡은 회벽이 보이고, 오른쪽 아래에는 건반이 있는 목제 오르간 일부가 하나 보인다. 찰리의 상체가 전경 대부분을 차지하고 배경 설비는 일부만 드러난다. 이전 장면과 오르간의 화면상 좌우 관계는 다르지만 촬영 각도가 열려 있어 이것만으로 구조 모순을 확정할 수 없다. 라디오·책상·의자·등불은 식별되지 않으며 반사상이나 중복 설비는 없다.",
        "entities": "찰리 한 개체의 몸통, 양어깨 일부, 팔 일부와 흰 마스크형 턱이 보인다. 모래색 각진 장갑, 중앙 원형 가슴 부품, 긁히고 닳은 금속, 올리브갈색 낡은 코트와 허리띠가 참조와 부합한다. 모자와 눈, 짧은 다리는 프레임 밖이다. 흐려진 가슴 표식은 뚜렷하게 확인되지 않으며 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "장갑판과 원형 부품은 몸통에 결합되어 있고 팔은 어깨와 팔꿈치 관절로 이어진다. 코트는 어깨에 걸리고 허리띠가 몸통을 감싸며 자연스럽게 접힌다. 하체의 바닥 접촉은 크롭 밖이지만 공중에 뜬 자세를 나타내는 단서는 없다. 틈새의 광원이 인접한 금속과 천에 청색 반사광을 만들어 물리적 재질감이 유지된다."
       },
       {
        "label": "B",
        "direction": "가슴 장갑 틈새와 중앙 원형 부품 모두에서 푸른빛이 카메라 쪽으로 방출된다. 특히 원형 중심이 희게 타오르므로 시선이 틈새보다 중심 광원으로 먼저 모인다. 얼굴은 턱 부분만 보여 눈의 목표는 확인할 수 없고, 무기나 방향성 있는 신체 이동은 없다.",
        "built_space": "왼쪽 배경에 목제 오르간 하나가 보이며 건반, 펼친 악보, 상단의 액자 하나와 책 더미, 꽃병 형태의 소품이 확인된다. 이는 이전 장면의 오르간 주변 구성과 잘 연결된다. 찰리는 오르간 앞쪽 전경을 차지하고 배경 크기도 과장되지 않았다. 십자가·라디오·책상·의자는 이 크롭에서 확인되지 않으며 불가능한 반사나 중복 설비는 없다.",
        "entities": "찰리 한 개체의 상체와 양팔 일부가 보인다. 흰 마스크형 턱, 모래색 어깨 및 가슴 장갑, 중앙 원형 부품, 낡은 코트와 버클 허리띠가 참조의 외형을 유지한다. 모자와 눈, 다리는 프레임 밖이며 가슴의 희미한 표식은 명확히 판별되지 않는다. 다른 인물은 없고 배경 악보도 읽을 수 있는 문자로 드러나지 않는다.",
        "hard_violations": [],
        "physics": "어깨·팔·가슴의 기계 부품이 관절과 몸통에 연결되어 있으며 따로 떠 있는 물체는 없다. 코트는 어깨에서 내려와 허리띠에 잡혀 있고 소매에는 팔 자세에 따른 주름이 보인다. 발은 프레임 밖이므로 접지를 직접 확인할 수 없지만 떠 있거나 지지 없이 기울어진 몸은 아니다. 강한 청색광이 장갑 가장자리와 코트 안쪽에 반사되며 금속 경계도 대체로 읽힌다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "가슴을 더 밀착해 잡고 원형 중심부보다 장갑판 틈새의 눈부신 청색 방출을 강조하여, 지시된 순간과 발광 위치를 가장 정확하게 구현했다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "가슴 클로즈업과 오르간 주변 공간은 충실하지만, 강한 원형 중심광이 시선을 지배해 장갑 틈새에서 빛이 뿜어지는 순간이라는 핵심이 A보다 덜 선명하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "푸른빛이 원형 가슴 부품 둘레와 가로·세로 장갑 이음새에서 바깥으로 새어 나와 코트 안쪽과 금속 모서리를 밝힌다. 원형 부품 중심은 비교적 어두워 발광의 출처가 틈새로 읽힌다. 턱 아래만 보여 시선 방향은 확인할 수 없으며, 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "왼쪽 배경에 벽 십자가 하나와 낡은 회벽이 보이고, 오른쪽 아래에는 건반이 있는 목제 오르간 일부가 하나 보인다. 찰리의 상체가 전경 대부분을 차지하고 배경 설비는 일부만 드러난다. 이전 장면과 오르간의 화면상 좌우 관계는 다르지만 촬영 각도가 열려 있어 이것만으로 구조 모순을 확정할 수 없다. 라디오·책상·의자·등불은 식별되지 않으며 반사상이나 중복 설비는 없다.",
        "entities": "찰리 한 개체의 몸통, 양어깨 일부, 팔 일부와 흰 마스크형 턱이 보인다. 모래색 각진 장갑, 중앙 원형 가슴 부품, 긁히고 닳은 금속, 올리브갈색 낡은 코트와 허리띠가 참조와 부합한다. 모자와 눈, 짧은 다리는 프레임 밖이다. 흐려진 가슴 표식은 뚜렷하게 확인되지 않으며 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "장갑판과 원형 부품은 몸통에 결합되어 있고 팔은 어깨와 팔꿈치 관절로 이어진다. 코트는 어깨에 걸리고 허리띠가 몸통을 감싸며 자연스럽게 접힌다. 하체의 바닥 접촉은 크롭 밖이지만 공중에 뜬 자세를 나타내는 단서는 없다. 틈새의 광원이 인접한 금속과 천에 청색 반사광을 만들어 물리적 재질감이 유지된다."
       },
       {
        "label": "A",
        "direction": "가슴 장갑 틈새와 중앙 원형 부품 모두에서 푸른빛이 카메라 쪽으로 방출된다. 특히 원형 중심이 희게 타오르므로 시선이 틈새보다 중심 광원으로 먼저 모인다. 얼굴은 턱 부분만 보여 눈의 목표는 확인할 수 없고, 무기나 방향성 있는 신체 이동은 없다.",
        "built_space": "왼쪽 배경에 목제 오르간 하나가 보이며 건반, 펼친 악보, 상단의 액자 하나와 책 더미, 꽃병 형태의 소품이 확인된다. 이는 이전 장면의 오르간 주변 구성과 잘 연결된다. 찰리는 오르간 앞쪽 전경을 차지하고 배경 크기도 과장되지 않았다. 십자가·라디오·책상·의자는 이 크롭에서 확인되지 않으며 불가능한 반사나 중복 설비는 없다.",
        "entities": "찰리 한 개체의 상체와 양팔 일부가 보인다. 흰 마스크형 턱, 모래색 어깨 및 가슴 장갑, 중앙 원형 부품, 낡은 코트와 버클 허리띠가 참조의 외형을 유지한다. 모자와 눈, 다리는 프레임 밖이며 가슴의 희미한 표식은 명확히 판별되지 않는다. 다른 인물은 없고 배경 악보도 읽을 수 있는 문자로 드러나지 않는다.",
        "hard_violations": [],
        "physics": "어깨·팔·가슴의 기계 부품이 관절과 몸통에 연결되어 있으며 따로 떠 있는 물체는 없다. 코트는 어깨에서 내려와 허리띠에 잡혀 있고 소매에는 팔 자세에 따른 주름이 보인다. 발은 프레임 밖이므로 접지를 직접 확인할 수 없지만 떠 있거나 지지 없이 기울어진 몸은 아니다. 강한 청색광이 장갑 가장자리와 코트 안쪽에 반사되며 금속 경계도 대체로 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.714
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.714
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1714
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "가슴 장갑 틈새로 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업을 사실적이고 강렬한 빛 반사로 잘 구현했습니다."
   },
   {
    "label": "B",
    "score": 1714,
    "verdict_ko": "클로즈업과 배경 요소는 좋으나, 빛이 강하게 뿜어져 나오기보다는 단순한 LED 조명선처럼 묘사되어 극적인 느낌이 부족합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6_sel.png",
    "asset_id": "badf6883-3504-4df6-a6ab-784e4bb3b0cb",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-3485-70ce-bff6-2d08683f9e0d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S29sh6"
  }
 },
 "S30sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:19:09.723529+00:00",
  "fingerprint": "4889ade3bc6df999faa9cf7ed7ba793d32f35401364ded4af1734a745f357cae",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S30sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S30sh5_sel.png",
  "source_sha256": "3715dd98806d6f5ce3ba2a50098f73270124a8d5cfe7c5ee7e2eb85b6296b677",
  "file": "S30sh5_cine.png",
  "staged_sha256": "bac5d97f3cf4608d49dc7a32995dc5a81218dbae2a5d504b9ced99209859c145",
  "latency_ms": 9031
 },
 "S30sh10::signage": {
  "fp": "8c35ec84c871a895",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S30sh10": {
  "input_fingerprint": "1bd7a682a05a4e89",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 몸통을 향해 양팔을 뻗어 덥석 끌어안은 채 밀착된 구도환의 환한 상체.\n\nLOCATION (lock): Inside the sparse church basement prayer room, near the foot of the stairs and the organ, under brightened lantern light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The now-bright lantern illumination gives the embrace clear facial and bodily separation with restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same basement prayer-room surfaces and its cross, small desk, chairs, and organ. Exclude the upstairs doorway and exterior surveillance equipment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns are now brightly lit, and the previously broken radio is operating. Charlie retains his old coat and hat beside the small organ. 구도환: He has entered the basement and holds an excited, forward-leaning embrace posture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 몸통을 향해 양팔을 뻗어 덥석 끌어안은 채 밀착된 구도환의 환한 상체.\n\nLOCATION (lock): Inside the sparse church basement prayer room, near the foot of the stairs and the organ, under brightened lantern light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The now-bright lantern illumination gives the embrace clear facial and bodily separation with restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same basement prayer-room surfaces and its cross, small desk, chairs, and organ. Exclude the upstairs doorway and exterior surveillance equipment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns are now brightly lit, and the previously broken radio is operating. Charlie retains his old coat and hat beside the small organ. 구도환: He has entered the basement and holds an excited, forward-leaning embrace posture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 몸통을 향해 양팔을 뻗어 덥석 끌어안은 채 밀착된 구도환의 환한 상체.\n\nLOCATION (lock): Inside the sparse church basement prayer room, near the foot of the stairs and the organ, under brightened lantern light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The now-bright lantern illumination gives the embrace clear facial and bodily separation with restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same basement prayer-room surfaces and its cross, small desk, chairs, and organ. Exclude the upstairs doorway and exterior surveillance equipment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns are now brightly lit, and the previously broken radio is operating. Charlie retains his old coat and hat beside the small organ. 구도환: He has entered the basement and holds an excited, forward-leaning embrace posture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "구도환은 찰리의 얼굴을 향해 시선을 고정하고 웃고 있으며, 찰리는 구도환을 내려다봄.",
    "built_space": "오르간, 의자, 랜턴이 배치되었으나, 지시된 계단, 책상, 십자가는 보이지 않음.",
    "entities": "찰리는 이전 샷의 파란색 발광 가슴을 포함하여 외형이 일치함. 하지만 구도환은 레퍼런스와 전혀 다른 갈색 코트를 입고 있음.",
    "hard_violations": [],
    "physics": "구도환은 두 발로 서서 양팔을 뻗어 찰리의 가슴에 손을 얹고 있으며 지지 상태가 자연스러움."
   },
   {
    "label": "B",
    "direction": "구도환은 찰리의 품에 안겨 미소 짓고 있으며, 찰리는 고개를 숙여 구도환을 바라봄.",
    "built_space": "오르간, 계단, 십자가, 책상, 랜턴이 올바르게 배치됨. 단, 제외 지시된 위층 출입구가 계단 위로 묘사됨.",
    "entities": "구도환의 얼굴과 복장(녹색 코트, 밧줄 벨트 등)이 레퍼런스와 정확히 일치함. 찰리의 외형도 일치하나 가슴의 파란색 발광이 누락됨. 작동 중인 라디오가 책상 위에 있음.",
    "hard_violations": [],
    "physics": "구도환은 찰리에게 몸을 기대어 밀착된 포옹을 하고 있으며, 찰리의 오른손이 구도환의 등을 감싸며 물리적으로 안정된 상태임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "제외 지시된 위층 출입구가 노출되고 찰리의 가슴 발광이 누락되었으나, 구도환의 완벽한 레퍼런스 일치도와 필수 소품(라디오, 십자가 등)의 배치, 밀착된 포옹 연출이 매우 훌륭합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 가슴 발광은 잘 구현되었으나, 구도환의 복장 레퍼런스를 완전히 위반하였고 요구된 배경 소품과 밀착된 포옹 디테일이 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "구도환은 찰리의 품에 안겨 미소 짓고 있으며, 찰리는 고개를 숙여 구도환을 바라봄.",
        "built_space": "오르간, 계단, 십자가, 책상, 랜턴이 올바르게 배치됨. 단, 제외 지시된 위층 출입구가 계단 위로 묘사됨.",
        "entities": "구도환의 얼굴과 복장(녹색 코트, 밧줄 벨트 등)이 레퍼런스와 정확히 일치함. 찰리의 외형도 일치하나 가슴의 파란색 발광이 누락됨. 작동 중인 라디오가 책상 위에 있음.",
        "hard_violations": [],
        "physics": "구도환은 찰리에게 몸을 기대어 밀착된 포옹을 하고 있으며, 찰리의 오른손이 구도환의 등을 감싸며 물리적으로 안정된 상태임."
       },
       {
        "label": "A",
        "direction": "구도환은 찰리의 얼굴을 향해 시선을 고정하고 웃고 있으며, 찰리는 구도환을 내려다봄.",
        "built_space": "오르간, 의자, 랜턴이 배치되었으나, 지시된 계단, 책상, 십자가는 보이지 않음.",
        "entities": "찰리는 이전 샷의 파란색 발광 가슴을 포함하여 외형이 일치함. 하지만 구도환은 레퍼런스와 전혀 다른 갈색 코트를 입고 있음.",
        "hard_violations": [],
        "physics": "구도환은 두 발로 서서 양팔을 뻗어 찰리의 가슴에 손을 얹고 있으며 지지 상태가 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "제외 지시된 위층 출입구가 노출되고 찰리의 가슴 발광이 누락되었으나, 구도환의 완벽한 레퍼런스 일치도와 필수 소품(라디오, 십자가 등)의 배치, 밀착된 포옹 연출이 매우 훌륭합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 가슴 발광은 잘 구현되었으나, 구도환의 복장 레퍼런스를 완전히 위반하였고 요구된 배경 소품과 밀착된 포옹 디테일이 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "구도환은 찰리의 품에 안겨 미소 짓고 있으며, 찰리는 고개를 숙여 구도환을 바라봄.",
        "built_space": "오르간, 계단, 십자가, 책상, 랜턴이 올바르게 배치됨. 단, 제외 지시된 위층 출입구가 계단 위로 묘사됨.",
        "entities": "구도환의 얼굴과 복장(녹색 코트, 밧줄 벨트 등)이 레퍼런스와 정확히 일치함. 찰리의 외형도 일치하나 가슴의 파란색 발광이 누락됨. 작동 중인 라디오가 책상 위에 있음.",
        "hard_violations": [],
        "physics": "구도환은 찰리에게 몸을 기대어 밀착된 포옹을 하고 있으며, 찰리의 오른손이 구도환의 등을 감싸며 물리적으로 안정된 상태임."
       },
       {
        "label": "A",
        "direction": "구도환은 찰리의 얼굴을 향해 시선을 고정하고 웃고 있으며, 찰리는 구도환을 내려다봄.",
        "built_space": "오르간, 의자, 랜턴이 배치되었으나, 지시된 계단, 책상, 십자가는 보이지 않음.",
        "entities": "찰리는 이전 샷의 파란색 발광 가슴을 포함하여 외형이 일치함. 하지만 구도환은 레퍼런스와 전혀 다른 갈색 코트를 입고 있음.",
        "hard_violations": [],
        "physics": "구도환은 두 발로 서서 양팔을 뻗어 찰리의 가슴에 손을 얹고 있으며 지지 상태가 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "구도환이 웃으며 찰리의 몸통에 실제로 밀착해 끌어안은 순간과 참조 의상을 잘 구현했지만, 상체 중심 구도보다 다소 넓고 찰리의 가슴 발광이 사라졌다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "상체 중심 촬영과 찰리의 발광·재질은 충실하지만, 구도환의 몸통이 떨어져 있어 완료된 포옹보다 다가가는 순간으로 보이며 얼굴·머리·외투도 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "구도환은 찰리 쪽으로 몸을 기울이고 웃는 얼굴을 오른쪽으로 향한다. 가까운 팔은 찰리의 가슴을 가로질러 반대편 위팔을 감싸며, 다른 팔은 찰리 뒤로 가려진다. 찰리는 얼굴을 구도환 쪽으로 내려 향한다. 포옹의 대상과 접촉 방향이 일치한다.",
        "built_space": "왼쪽에 목제 오르간 한 대와 벤치 하나, 중앙 뒤에 위로 이어지는 계단 한 줄, 오른쪽에 작은 책상 하나와 벽 십자가 하나가 보인다. 등은 상부에 매달린 것 하나와 책상 위 하나이며, 책상에는 라디오 한 대가 놓였다. 두 인물은 계단 아래 오르간 옆에 서 있다. 낡은 벽과 목제 악기는 참조 장소와 대체로 이어지며, 중복 설비나 불가능한 반사는 없다.",
        "entities": "등장하는 것은 구도환과 찰리뿐이다. 구도환은 검은 머리에 회색이 섞인 한국인 중년 남성으로 보이며, 웃는 얼굴과 후드 달린 낡은 외투·끈 허리띠가 참조에 가깝다. 찰리는 흰 기계식 얼굴, 베이지 장갑판, 긴 금속 팔, 낡은 코트와 모자를 유지한다. 다만 이전 장면의 강한 청색 가슴 발광은 보이지 않는다. 라디오는 존재하지만 작동 여부는 정지 화면으로 확인하기 어렵다.",
        "hard_violations": [],
        "physics": "구도환의 가슴과 어깨가 찰리 몸통에 붙고, 보이는 손바닥은 찰리 위팔에 닿아 있다. 찰리의 반대쪽 금속 손은 구도환의 등 뒤를 받치는 위치에 보인다. 발은 프레임 밖이지만 두 몸통 모두 서 있는 하체로 자연스럽게 이어지며 부유를 암시하지 않는다. 책상 위 등과 라디오는 상판에 지지되고, 상부 등은 프레임 위쪽으로 이어지는 매달림 구조로 보인다."
       },
       {
        "label": "B",
        "direction": "구도환은 찰리 얼굴을 올려다보며 양팔을 몸통 쪽으로 뻗는다. 가까운 손은 찰리 가슴 옆 코트에 닿고, 먼 팔은 반대쪽 몸통 뒤로 향한다. 찰리는 구도환을 내려다본다. 팔과 시선의 대상은 맞지만, 구도환의 가슴은 찰리에게 붙지 않아 덥석 끌어안은 상태보다는 포옹 직전으로 읽힌다.",
        "built_space": "왼쪽 배경에 목제 오르간 한 대, 그 앞 벤치 하나, 더 왼쪽에 의자 하나가 보인다. 오르간 위에는 등 하나와 액자·꽃병이 놓여 있다. 작은 책상, 라디오, 계단과 벽 십자가는 이 구도에서 보이지 않는다. 낡은 벽과 오르간 표면은 이전 장면과 유사하며, 설비 중복이나 광학적으로 불가능한 반사는 없다.",
        "entities": "구도환과 찰리만 등장한다. 구도환은 검은 머리의 한국인 중년 남성으로 보이지만 참조보다 짧고 내려온 머리, 다른 인상의 얼굴을 갖고 있다. 외투도 참조의 후드·끈 여밈 외투가 아니라 갈색 재킷형 코트로 바뀌었다. 찰리는 모자, 흰 기계식 얼굴, 베이지 장갑판과 낡은 코트를 유지하며 청색 가슴 발광도 이어진다. 짧은 다리는 프레임 밖이라 평가 대상이 아니다.",
        "hard_violations": [],
        "physics": "구도환은 서 있는 하체에서 상체를 앞으로 기울이고, 보이는 손바닥으로 찰리의 코트와 가슴 측면을 짚는다. 팔의 연결과 접촉은 가능하지만 두 몸통 사이에 간격이 남는다. 찰리의 팔은 어깨에서 자연스럽게 내려오며, 부유하거나 지지 없이 놓인 신체는 없다. 등과 배경 소품은 오르간 상판에, 의자와 벤치는 바닥에 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "구도환이 웃으며 찰리의 몸통에 실제로 밀착해 끌어안은 순간과 참조 의상을 잘 구현했지만, 상체 중심 구도보다 다소 넓고 찰리의 가슴 발광이 사라졌다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "상체 중심 촬영과 찰리의 발광·재질은 충실하지만, 구도환의 몸통이 떨어져 있어 완료된 포옹보다 다가가는 순간으로 보이며 얼굴·머리·외투도 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "구도환은 찰리 쪽으로 몸을 기울이고 웃는 얼굴을 오른쪽으로 향한다. 가까운 팔은 찰리의 가슴을 가로질러 반대편 위팔을 감싸며, 다른 팔은 찰리 뒤로 가려진다. 찰리는 얼굴을 구도환 쪽으로 내려 향한다. 포옹의 대상과 접촉 방향이 일치한다.",
        "built_space": "왼쪽에 목제 오르간 한 대와 벤치 하나, 중앙 뒤에 위로 이어지는 계단 한 줄, 오른쪽에 작은 책상 하나와 벽 십자가 하나가 보인다. 등은 상부에 매달린 것 하나와 책상 위 하나이며, 책상에는 라디오 한 대가 놓였다. 두 인물은 계단 아래 오르간 옆에 서 있다. 낡은 벽과 목제 악기는 참조 장소와 대체로 이어지며, 중복 설비나 불가능한 반사는 없다.",
        "entities": "등장하는 것은 구도환과 찰리뿐이다. 구도환은 검은 머리에 회색이 섞인 한국인 중년 남성으로 보이며, 웃는 얼굴과 후드 달린 낡은 외투·끈 허리띠가 참조에 가깝다. 찰리는 흰 기계식 얼굴, 베이지 장갑판, 긴 금속 팔, 낡은 코트와 모자를 유지한다. 다만 이전 장면의 강한 청색 가슴 발광은 보이지 않는다. 라디오는 존재하지만 작동 여부는 정지 화면으로 확인하기 어렵다.",
        "hard_violations": [],
        "physics": "구도환의 가슴과 어깨가 찰리 몸통에 붙고, 보이는 손바닥은 찰리 위팔에 닿아 있다. 찰리의 반대쪽 금속 손은 구도환의 등 뒤를 받치는 위치에 보인다. 발은 프레임 밖이지만 두 몸통 모두 서 있는 하체로 자연스럽게 이어지며 부유를 암시하지 않는다. 책상 위 등과 라디오는 상판에 지지되고, 상부 등은 프레임 위쪽으로 이어지는 매달림 구조로 보인다."
       },
       {
        "label": "A",
        "direction": "구도환은 찰리 얼굴을 올려다보며 양팔을 몸통 쪽으로 뻗는다. 가까운 손은 찰리 가슴 옆 코트에 닿고, 먼 팔은 반대쪽 몸통 뒤로 향한다. 찰리는 구도환을 내려다본다. 팔과 시선의 대상은 맞지만, 구도환의 가슴은 찰리에게 붙지 않아 덥석 끌어안은 상태보다는 포옹 직전으로 읽힌다.",
        "built_space": "왼쪽 배경에 목제 오르간 한 대, 그 앞 벤치 하나, 더 왼쪽에 의자 하나가 보인다. 오르간 위에는 등 하나와 액자·꽃병이 놓여 있다. 작은 책상, 라디오, 계단과 벽 십자가는 이 구도에서 보이지 않는다. 낡은 벽과 오르간 표면은 이전 장면과 유사하며, 설비 중복이나 광학적으로 불가능한 반사는 없다.",
        "entities": "구도환과 찰리만 등장한다. 구도환은 검은 머리의 한국인 중년 남성으로 보이지만 참조보다 짧고 내려온 머리, 다른 인상의 얼굴을 갖고 있다. 외투도 참조의 후드·끈 여밈 외투가 아니라 갈색 재킷형 코트로 바뀌었다. 찰리는 모자, 흰 기계식 얼굴, 베이지 장갑판과 낡은 코트를 유지하며 청색 가슴 발광도 이어진다. 짧은 다리는 프레임 밖이라 평가 대상이 아니다.",
        "hard_violations": [],
        "physics": "구도환은 서 있는 하체에서 상체를 앞으로 기울이고, 보이는 손바닥으로 찰리의 코트와 가슴 측면을 짚는다. 팔의 연결과 접촉은 가능하지만 두 몸통 사이에 간격이 남는다. 찰리의 팔은 어깨에서 자연스럽게 내려오며, 부유하거나 지지 없이 놓인 신체는 없다. 등과 배경 소품은 오르간 상판에, 의자와 벤치는 바닥에 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.321,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.321,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "제외 지시된 위층 출입구가 노출되고 찰리의 가슴 발광이 누락되었으나, 구도환의 완벽한 레퍼런스 일치도와 필수 소품(라디오, 십자가 등)의 배치, 밀착된 포옹 연출이 매우 훌륭합니다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "찰리의 가슴 발광은 잘 구현되었으나, 구도환의 복장 레퍼런스를 완전히 위반하였고 요구된 배경 소품과 밀착된 포옹 디테일이 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S30sh5_sel.png",
    "asset_id": "a9d6a42a-8a4a-491a-9537-e91f945910fd",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1023014>",
    "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-3627-7e62-bed4-3f40d30cb127",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S30sh5"
  }
 },
 "S30sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:20:32.805719+00:00",
  "fingerprint": "45dc78b6c6cb6dec6b602fce9942ea1fb191803fdb3df68d3e4c00fc85a9298d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S30sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S30sh10_sel.png",
  "source_sha256": "15c73a8ff41242a580a6485c113977d97d919c82ba35de55fb7dfcefef95acaa",
  "file": "S30sh10_cine.png",
  "staged_sha256": "7ad1ab78bdf8f51b715b49ddf22cdc1f8e9329289ca800d19a8a213263b6fae2",
  "latency_ms": 9979
 },
 "S30sh17::signage": {
  "fp": "b36907ac3be0935c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S30sh17": {
  "input_fingerprint": "ffeaedc0d77f8fc8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사라는 말에 충격을 받은 듯 놀란 눈으로 구도환을 일제히 쳐다보는 앰버와 신부의 상체.\n\nLOCATION (lock): In the lantern-lit gathering area of the church basement prayer room, near the small table and organ. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright lantern illumination and controlled contrast, letting widened eyes carry the shock without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 구도환 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's fixed surfaces and sparse furnishings, including the small organ. Exclude the upstairs entrance and outdoor searchlights or drones.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit and the old radio remains functional, with the organ and other sparse furnishings unchanged. Charlie still wears the old coat and hat. 앰버: She remains in the basement, startled, with the outer garment used as a mouth covering. 신부: He stands in the basement wearing his clerical collar, visibly startled. 구도환: He remains in the basement and shakes his head in denial.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사라는 말에 충격을 받은 듯 놀란 눈으로 구도환을 일제히 쳐다보는 앰버와 신부의 상체.\n\nLOCATION (lock): In the lantern-lit gathering area of the church basement prayer room, near the small table and organ. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright lantern illumination and controlled contrast, letting widened eyes carry the shock without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 구도환 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's fixed surfaces and sparse furnishings, including the small organ. Exclude the upstairs entrance and outdoor searchlights or drones.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit and the old radio remains functional, with the organ and other sparse furnishings unchanged. Charlie still wears the old coat and hat. 앰버: She remains in the basement, startled, with the outer garment used as a mouth covering. 신부: He stands in the basement wearing his clerical collar, visibly startled. 구도환: He remains in the basement and shakes his head in denial.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사라는 말에 충격을 받은 듯 놀란 눈으로 구도환을 일제히 쳐다보는 앰버와 신부의 상체.\n\nLOCATION (lock): In the lantern-lit gathering area of the church basement prayer room, near the small table and organ. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright lantern illumination and controlled contrast, letting widened eyes carry the shock without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 구도환 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's fixed surfaces and sparse furnishings, including the small organ. Exclude the upstairs entrance and outdoor searchlights or drones.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit and the old radio remains functional, with the organ and other sparse furnishings unchanged. Charlie still wears the old coat and hat. 앰버: She remains in the basement, startled, with the outer garment used as a mouth covering. 신부: He stands in the basement wearing his clerical collar, visibly startled. 구도환: He remains in the basement and shakes his head in denial.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버와 신부 모두 눈을 크게 뜨고 카메라 왼쪽의 화면 밖 지점을 바라본다. 같은 상대를 보는 반응으로 읽히지만, 구도환이 보이지 않아 시선의 도착 대상을 직접 확인할 수는 없다. 무기나 이동 동작은 없다.",
    "built_space": "왼쪽에 계단 출입구 하나, 중앙 뒤에 오르간 하나와 의자 하나, 오른쪽에 작은 탁자 하나와 라디오 하나가 보인다. 등은 천장에 매달린 것 하나와 탁자 위에서 일부 가려진 것 하나이며, 오른쪽 벽에는 십자가 하나가 있다. 두 사람은 가구 앞에 서 있다. 낡은 벽과 밝은 등불은 이어지지만, 참조에서 계단 왼쪽에 있던 오르간이 여기서는 계단 오른쪽에 놓여 공간 연속성이 약하다. 반사는 없다.",
    "entities": "금발의 어린 여자아이와 주름진 얼굴의 한국인 노년 남성이 보이며, 앰버와 신부의 나이·성별·외형에 대체로 맞는다. 앰버의 남색 상의와 갈색 작업복은 참조와 연결되지만 머리의 고글은 없다. 입과 코를 가린 것은 겉옷을 들어 올린 형태보다 별도의 목 감싸개로 보인다. 신부는 검은 성직복과 흰 칼라를 착용하고 목걸이는 손에 일부 가려져 있다. 구도환과 찰리는 화면에 없다. 라디오는 존재하지만 작동 여부는 확인할 수 없으며, 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "두 사람의 몸통은 프레임 아래로 자연스럽게 이어지며, 발이 잘렸다는 이유로 부유한다고 볼 근거는 없다. 신부의 두 손은 가슴 앞에서 서로 맞잡혀 있다. 앰버의 얼굴 가림 천은 머리와 목을 감싸 고정된 것으로 보이므로 떠 있는 물체는 아니지만, 겉옷을 입에 대는 행동과는 다르다. 천장 등은 위쪽 고정부에 매달리고 라디오는 탁자 위에 놓여 있다."
   },
   {
    "label": "A",
    "direction": "앰버는 왼쪽 전경의 구도환 얼굴을 향해 올려다보고, 신부도 같은 남자의 얼굴 쪽으로 눈과 고개를 돌린다. 두 사람의 놀란 시선이 실제로 구도환에게 모인다. 구도환 역시 두 사람 쪽을 향하지만, 이 정지 화면만으로 고개를 저어 부인하는 동작까지 분명하게 확인되지는 않는다.",
    "built_space": "중앙 뒤에 오르간 하나, 그 위에 성화 액자 하나와 책 더미가 있고, 오른쪽에 계단 출입구 하나가 보인다. 오르간이 계단 왼쪽에 있는 관계는 참조와 맞는다. 천장 등 하나가 상단에 보이며 작은 탁자와 라디오는 이 구도에서 보이지 않는다. 앰버와 신부는 오르간 앞에, 구도환은 카메라 가까운 왼쪽에 서 있어 가구와 충돌하지 않는다. 벽의 거친 재질과 밝은 등불 조명이 유지되며 반사는 없다.",
    "entities": "앰버는 금발에 둥근 얼굴과 큰 눈을 가진 어린 여자아이로, 참조의 외형에 대체로 부합한다. 고글은 없지만 겉옷을 직접 끌어올려 입을 가리고 있다. 신부는 주름진 한국인 노년 남성이며 검은 성직복, 흰 칼라, 십자가 목걸이가 확인된다. 왼쪽 남성의 검고 희끗한 머리와 낡은 후드 겉옷은 이전 장면의 구도환과 연결되지만 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵다. 찰리나 다른 인물은 없고 읽을 수 있는 글자도 없다.",
    "hard_violations": [],
    "physics": "앰버는 양손으로 겉옷의 가장자리를 움켜쥐고 입까지 들어 올려 천의 지지와 주름이 자연스럽다. 신부와 구도환은 서 있는 몸통이 프레임 아래로 이어지며 비정상적인 부유나 관절 꺾임은 보이지 않는다. 십자가는 목걸이 줄에 매달려 있고 천장 등은 프레임 위로 이어지는 지지부에 걸려 있다. 오르간 위 액자와 책도 받침 면에 놓여 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 사람의 놀란 상체를 담은 미디엄 숏은 맞지만, 오르간과 계단의 위치 관계가 참조와 다르고 앰버의 입 가림도 겉옷보다 별도 천으로 보인다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "앰버와 신부가 구도환을 함께 바라보는 시선, 겉옷으로 입을 가린 행동, 지하실의 공간 관계를 가장 충실하게 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 신부 모두 눈을 크게 뜨고 카메라 왼쪽의 화면 밖 지점을 바라본다. 같은 상대를 보는 반응으로 읽히지만, 구도환이 보이지 않아 시선의 도착 대상을 직접 확인할 수는 없다. 무기나 이동 동작은 없다.",
        "built_space": "왼쪽에 계단 출입구 하나, 중앙 뒤에 오르간 하나와 의자 하나, 오른쪽에 작은 탁자 하나와 라디오 하나가 보인다. 등은 천장에 매달린 것 하나와 탁자 위에서 일부 가려진 것 하나이며, 오른쪽 벽에는 십자가 하나가 있다. 두 사람은 가구 앞에 서 있다. 낡은 벽과 밝은 등불은 이어지지만, 참조에서 계단 왼쪽에 있던 오르간이 여기서는 계단 오른쪽에 놓여 공간 연속성이 약하다. 반사는 없다.",
        "entities": "금발의 어린 여자아이와 주름진 얼굴의 한국인 노년 남성이 보이며, 앰버와 신부의 나이·성별·외형에 대체로 맞는다. 앰버의 남색 상의와 갈색 작업복은 참조와 연결되지만 머리의 고글은 없다. 입과 코를 가린 것은 겉옷을 들어 올린 형태보다 별도의 목 감싸개로 보인다. 신부는 검은 성직복과 흰 칼라를 착용하고 목걸이는 손에 일부 가려져 있다. 구도환과 찰리는 화면에 없다. 라디오는 존재하지만 작동 여부는 확인할 수 없으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 프레임 아래로 자연스럽게 이어지며, 발이 잘렸다는 이유로 부유한다고 볼 근거는 없다. 신부의 두 손은 가슴 앞에서 서로 맞잡혀 있다. 앰버의 얼굴 가림 천은 머리와 목을 감싸 고정된 것으로 보이므로 떠 있는 물체는 아니지만, 겉옷을 입에 대는 행동과는 다르다. 천장 등은 위쪽 고정부에 매달리고 라디오는 탁자 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "앰버는 왼쪽 전경의 구도환 얼굴을 향해 올려다보고, 신부도 같은 남자의 얼굴 쪽으로 눈과 고개를 돌린다. 두 사람의 놀란 시선이 실제로 구도환에게 모인다. 구도환 역시 두 사람 쪽을 향하지만, 이 정지 화면만으로 고개를 저어 부인하는 동작까지 분명하게 확인되지는 않는다.",
        "built_space": "중앙 뒤에 오르간 하나, 그 위에 성화 액자 하나와 책 더미가 있고, 오른쪽에 계단 출입구 하나가 보인다. 오르간이 계단 왼쪽에 있는 관계는 참조와 맞는다. 천장 등 하나가 상단에 보이며 작은 탁자와 라디오는 이 구도에서 보이지 않는다. 앰버와 신부는 오르간 앞에, 구도환은 카메라 가까운 왼쪽에 서 있어 가구와 충돌하지 않는다. 벽의 거친 재질과 밝은 등불 조명이 유지되며 반사는 없다.",
        "entities": "앰버는 금발에 둥근 얼굴과 큰 눈을 가진 어린 여자아이로, 참조의 외형에 대체로 부합한다. 고글은 없지만 겉옷을 직접 끌어올려 입을 가리고 있다. 신부는 주름진 한국인 노년 남성이며 검은 성직복, 흰 칼라, 십자가 목걸이가 확인된다. 왼쪽 남성의 검고 희끗한 머리와 낡은 후드 겉옷은 이전 장면의 구도환과 연결되지만 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵다. 찰리나 다른 인물은 없고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "앰버는 양손으로 겉옷의 가장자리를 움켜쥐고 입까지 들어 올려 천의 지지와 주름이 자연스럽다. 신부와 구도환은 서 있는 몸통이 프레임 아래로 이어지며 비정상적인 부유나 관절 꺾임은 보이지 않는다. 십자가는 목걸이 줄에 매달려 있고 천장 등은 프레임 위로 이어지는 지지부에 걸려 있다. 오르간 위 액자와 책도 받침 면에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "두 사람의 놀란 상체를 담은 미디엄 숏은 맞지만, 오르간과 계단의 위치 관계가 참조와 다르고 앰버의 입 가림도 겉옷보다 별도 천으로 보인다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "앰버와 신부가 구도환을 함께 바라보는 시선, 겉옷으로 입을 가린 행동, 지하실의 공간 관계를 가장 충실하게 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버와 신부 모두 눈을 크게 뜨고 카메라 왼쪽의 화면 밖 지점을 바라본다. 같은 상대를 보는 반응으로 읽히지만, 구도환이 보이지 않아 시선의 도착 대상을 직접 확인할 수는 없다. 무기나 이동 동작은 없다.",
        "built_space": "왼쪽에 계단 출입구 하나, 중앙 뒤에 오르간 하나와 의자 하나, 오른쪽에 작은 탁자 하나와 라디오 하나가 보인다. 등은 천장에 매달린 것 하나와 탁자 위에서 일부 가려진 것 하나이며, 오른쪽 벽에는 십자가 하나가 있다. 두 사람은 가구 앞에 서 있다. 낡은 벽과 밝은 등불은 이어지지만, 참조에서 계단 왼쪽에 있던 오르간이 여기서는 계단 오른쪽에 놓여 공간 연속성이 약하다. 반사는 없다.",
        "entities": "금발의 어린 여자아이와 주름진 얼굴의 한국인 노년 남성이 보이며, 앰버와 신부의 나이·성별·외형에 대체로 맞는다. 앰버의 남색 상의와 갈색 작업복은 참조와 연결되지만 머리의 고글은 없다. 입과 코를 가린 것은 겉옷을 들어 올린 형태보다 별도의 목 감싸개로 보인다. 신부는 검은 성직복과 흰 칼라를 착용하고 목걸이는 손에 일부 가려져 있다. 구도환과 찰리는 화면에 없다. 라디오는 존재하지만 작동 여부는 확인할 수 없으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 프레임 아래로 자연스럽게 이어지며, 발이 잘렸다는 이유로 부유한다고 볼 근거는 없다. 신부의 두 손은 가슴 앞에서 서로 맞잡혀 있다. 앰버의 얼굴 가림 천은 머리와 목을 감싸 고정된 것으로 보이므로 떠 있는 물체는 아니지만, 겉옷을 입에 대는 행동과는 다르다. 천장 등은 위쪽 고정부에 매달리고 라디오는 탁자 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "앰버는 왼쪽 전경의 구도환 얼굴을 향해 올려다보고, 신부도 같은 남자의 얼굴 쪽으로 눈과 고개를 돌린다. 두 사람의 놀란 시선이 실제로 구도환에게 모인다. 구도환 역시 두 사람 쪽을 향하지만, 이 정지 화면만으로 고개를 저어 부인하는 동작까지 분명하게 확인되지는 않는다.",
        "built_space": "중앙 뒤에 오르간 하나, 그 위에 성화 액자 하나와 책 더미가 있고, 오른쪽에 계단 출입구 하나가 보인다. 오르간이 계단 왼쪽에 있는 관계는 참조와 맞는다. 천장 등 하나가 상단에 보이며 작은 탁자와 라디오는 이 구도에서 보이지 않는다. 앰버와 신부는 오르간 앞에, 구도환은 카메라 가까운 왼쪽에 서 있어 가구와 충돌하지 않는다. 벽의 거친 재질과 밝은 등불 조명이 유지되며 반사는 없다.",
        "entities": "앰버는 금발에 둥근 얼굴과 큰 눈을 가진 어린 여자아이로, 참조의 외형에 대체로 부합한다. 고글은 없지만 겉옷을 직접 끌어올려 입을 가리고 있다. 신부는 주름진 한국인 노년 남성이며 검은 성직복, 흰 칼라, 십자가 목걸이가 확인된다. 왼쪽 남성의 검고 희끗한 머리와 낡은 후드 겉옷은 이전 장면의 구도환과 연결되지만 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵다. 찰리나 다른 인물은 없고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "앰버는 양손으로 겉옷의 가장자리를 움켜쥐고 입까지 들어 올려 천의 지지와 주름이 자연스럽다. 신부와 구도환은 서 있는 몸통이 프레임 아래로 이어지며 비정상적인 부유나 관절 꺾임은 보이지 않는다. 십자가는 목걸이 줄에 매달려 있고 천장 등은 프레임 위로 이어지는 지지부에 걸려 있다. 오르간 위 액자와 책도 받침 면에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 6,
   "A": 9
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 6,
    "verdict_ko": "두 사람의 놀란 상체를 담은 미디엄 숏은 맞지만, 오르간과 계단의 위치 관계가 참조와 다르고 앰버의 입 가림도 겉옷보다 별도 천으로 보인다."
   },
   {
    "label": "A",
    "score": 9,
    "verdict_ko": "앰버와 신부가 구도환을 함께 바라보는 시선, 겉옷으로 입을 가린 행동, 지하실의 공간 관계를 가장 충실하게 구현했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 구도환 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S30sh10_sel.png",
    "asset_id": "43c1d96b-64a7-4103-912e-1405125aac03",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-37d8-7f3e-94d5-ea0a54aff998",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S30sh10"
  },
  "staged_characters_added": [
   "C20"
  ]
 },
 "S30sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:22:01.652086+00:00",
  "fingerprint": "45ff159fa35d6ed78b732a80555c2a29c31e8f68fec13c5657da95426ce233d8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S30sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S30sh17_sel.png",
  "source_sha256": "893d43e14ddcbf6b85b445b9974ffdd6ab5515d64dca5143d9c07fee7f7e0e6a",
  "file": "S30sh17_cine.png",
  "staged_sha256": "61aa6a0eaa675378febcf57a781b903beb90f6778ec508b1a4cb6d7e6fb7b1ff",
  "latency_ms": 10900
 },
 "S31sh2::signage": {
  "fp": "72f50f3f97715f02",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::c1fe96fc846da4bd": {
  "subjects": [],
  "subject_text": "인천 난민촌 관리사무소 3층 감금실\n문고리가 달린 잠금식 출입문과 의자가 있는 폐쇄된 방. 출입문은 바깥 복도와 연결돼 있다.",
  "identity": "canonical",
  "scope_id": "L189",
  "scope_role": "location_interior",
  "scope_sha": "e56880f19b804921"
 },
 "S31sh2::bgfirst_bg": {
  "input_fingerprint": "22441b1de99bf4a4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2__bgfirst_bg.png",
  "asset_id": "6d0e961f-cfc1-4ccb-921e-40cb43da0593",
  "input_asset_ids": [
   "c9bf8684-260f-4c7f-bb67-b37f8e71578e",
   "d31b4d4c-5488-4514-8d64-f7f3972f4a12"
  ]
 },
 "S31sh2": {
  "input_fingerprint": "9647b7892ae1bd04",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management office is brightly lit, with the interrogation taking place on the third floor. 미연: Her face is badly swollen from the beating, and she remains in the interrogation room. 박철진: He is seated in the interrogation room after dismissing the subordinate from the beating.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management office is brightly lit, with the interrogation taking place on the third floor. 미연: Her face is badly swollen from the beating, and she remains in the interrogation room. 박철진: He is seated in the interrogation room after dismissing the subordinate from the beating.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management office is brightly lit, with the interrogation taking place on the third floor. 미연: Her face is badly swollen from the beating, and she remains in the interrogation room. 박철진: He is seated in the interrogation room after dismissing the subordinate from the beating.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2__bgfirst_bg.png",
     "asset_id": "6d0e961f-cfc1-4ccb-921e-40cb43da0593",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S31sh2.png",
     "asset_id": "c9bf8684-260f-4c7f-bb67-b37f8e71578e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L189B01.png",
     "asset_id": "d31b4d4c-5488-4514-8d64-f7f3972f4a12",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연과 박철진이 서로의 얼굴을 직접 마주보고 있음.",
    "built_space": "레퍼런스와 동일한 형태의 구금실 내부로, 두 개의 의자가 정확히 배치되어 있으며 열린 문과 복도 구조도 일치함.",
    "entities": "미연은 부은 얼굴과 레퍼런스의 체크 셔츠, 앞치마를 정확히 착용함. 박철진 역시 지정된 검은색 복장과 외모를 유지함.",
    "hard_violations": [],
    "physics": "두 사람 모두 의자에 안정적으로 앉아 체중을 지탱하고 있으며, 두 발이 바닥에 자연스럽게 닿아 있음."
   },
   {
    "label": "B",
    "direction": "박철진이 미연을 향해 몸을 기울이고 시선을 맞추고 있음.",
    "built_space": "구금실의 기본적인 벽면과 문의 위치는 비슷하나, 두 인물 뒤쪽으로 레퍼런스에 존재하지 않는 책상이 추가되어 있음.",
    "entities": "박철진의 외모와 의상은 일치하나, 미연은 레퍼런스에 없는 짙은 회색의 작업복(또는 죄수복)을 입고 있어 복장 지시를 위반함.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스 장소에 존재하지 않는 사물(책상 및 서류)이 임의로 추가됨",
     "[gpt-high] 지정 장소의 참조에 없고 장면에서도 요구하지 않은 책상과 그 위의 종이·필기구를 추가했다."
    ],
    "physics": "인물들 모두 의자에 앉아 무게를 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 미디엄 샷과 의자 크롭 프레이밍을 따르지 않고 풀 샷에 가깝게 렌더링되었으나, 두 인물의 외모와 의상, 장소 레퍼런스는 매우 충실하게 재현되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 프레이밍을 어겼을 뿐만 아니라, 장소에 없는 사물을 임의로 추가하고 인물의 의상 레퍼런스를 무시하는 등 다수의 위반 사항이 확인됩니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연과 박철진이 서로의 얼굴을 직접 마주보고 있음.",
        "built_space": "레퍼런스와 동일한 형태의 구금실 내부로, 두 개의 의자가 정확히 배치되어 있으며 열린 문과 복도 구조도 일치함.",
        "entities": "미연은 부은 얼굴과 레퍼런스의 체크 셔츠, 앞치마를 정확히 착용함. 박철진 역시 지정된 검은색 복장과 외모를 유지함.",
        "hard_violations": [],
        "physics": "두 사람 모두 의자에 안정적으로 앉아 체중을 지탱하고 있으며, 두 발이 바닥에 자연스럽게 닿아 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 미연을 향해 몸을 기울이고 시선을 맞추고 있음.",
        "built_space": "구금실의 기본적인 벽면과 문의 위치는 비슷하나, 두 인물 뒤쪽으로 레퍼런스에 존재하지 않는 책상이 추가되어 있음.",
        "entities": "박철진의 외모와 의상은 일치하나, 미연은 레퍼런스에 없는 짙은 회색의 작업복(또는 죄수복)을 입고 있어 복장 지시를 위반함.",
        "hard_violations": [
         "레퍼런스 장소에 존재하지 않는 사물(책상 및 서류)이 임의로 추가됨"
        ],
        "physics": "인물들 모두 의자에 앉아 무게를 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 미디엄 샷과 의자 크롭 프레이밍을 따르지 않고 풀 샷에 가깝게 렌더링되었으나, 두 인물의 외모와 의상, 장소 레퍼런스는 매우 충실하게 재현되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 프레이밍을 어겼을 뿐만 아니라, 장소에 없는 사물을 임의로 추가하고 인물의 의상 레퍼런스를 무시하는 등 다수의 위반 사항이 확인됩니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연과 박철진이 서로의 얼굴을 직접 마주보고 있음.",
        "built_space": "레퍼런스와 동일한 형태의 구금실 내부로, 두 개의 의자가 정확히 배치되어 있으며 열린 문과 복도 구조도 일치함.",
        "entities": "미연은 부은 얼굴과 레퍼런스의 체크 셔츠, 앞치마를 정확히 착용함. 박철진 역시 지정된 검은색 복장과 외모를 유지함.",
        "hard_violations": [],
        "physics": "두 사람 모두 의자에 안정적으로 앉아 체중을 지탱하고 있으며, 두 발이 바닥에 자연스럽게 닿아 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 미연을 향해 몸을 기울이고 시선을 맞추고 있음.",
        "built_space": "구금실의 기본적인 벽면과 문의 위치는 비슷하나, 두 인물 뒤쪽으로 레퍼런스에 존재하지 않는 책상이 추가되어 있음.",
        "entities": "박철진의 외모와 의상은 일치하나, 미연은 레퍼런스에 없는 짙은 회색의 작업복(또는 죄수복)을 입고 있어 복장 지시를 위반함.",
        "hard_violations": [
         "레퍼런스 장소에 존재하지 않는 사물(책상 및 서류)이 임의로 추가됨"
        ],
        "physics": "인물들 모두 의자에 앉아 무게를 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "바짝 다가앉은 거리와 거만한 표정은 잘 드러나지만, 없는 책상과 문구류를 추가했고 상체 중심보다 넓은 구도와 미연의 바뀐 복장이 지시에서 벗어난다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "장소와 인물 복장을 더 충실히 유지하고 박철진의 상체도 크게 잡았지만, 의자 사이가 넓어 ‘바짝 당겨 앉은’ 압박감이 부족하고 의자 노출도 과하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 박철진은 왼쪽 미연의 얼굴을 바라보며 상체를 앞으로 기울인다. 미연도 박철진 쪽으로 얼굴과 시선을 향한다. 서로의 무릎이 거의 맞닿아 대면 심문의 방향과 침범하는 거리가 명확하다.",
        "built_space": "목재 좌판과 검은 금속 골조의 의자 두 개에 한 사람씩 마주 앉았다. 오른쪽 열린 회색 문 하나, 밝고 어두운 회색으로 나뉜 벽, 회색 바닥, 천장의 직사각 조명 하나와 환기구 하나 및 작은 원형 장치가 보인다. 주요 건축 요소는 참조와 유사하지만, 두 사람 뒤 중앙에 참조에 없는 책상 하나가 추가되었다. 박철진의 의자는 옆 가장자리 일부가 아니라 좌판과 등받이 상당 부분이 드러나며, 인물도 무릎 아래까지 보여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "중년의 한국인 남녀로 읽히는 두 사람만 보인다. 박철진의 짧은 검은 머리, 얼굴, 검은 제복과 붉은 완장은 참조와 대체로 맞는다. 미연은 얼굴에 멍과 눈 주변 부기가 있지만, 참조의 짧은 곱슬머리 대신 묶은 긴 머리이며 체크 셔츠와 갈색 앞치마 대신 검은 작업복을 입었다. 모자도 없다. 추가된 책상에는 종이와 필기구가 보인다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "지정 장소의 참조에 없고 장면에서도 요구하지 않은 책상과 그 위의 종이·필기구를 추가했다."
        ],
        "physics": "두 사람 모두 엉덩이가 각자의 의자 좌판에 놓이고 등받이는 몸 뒤에 있다. 박철진은 양손을 허벅지에 대고 앞으로 숙여 체중을 지탱하며, 미연의 손은 무릎 위에 놓인다. 의자 다리는 바닥으로 이어지고 책상 위 물건도 상판에 놓여 있다. 발 일부가 화면 밖인 것은 지지 불가능의 증거가 아니며, 떠 있는 신체나 물건은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "오른쪽 박철진은 미연의 얼굴을 똑바로 응시하고 상체를 그녀 쪽으로 숙인다. 왼쪽 미연은 눈꺼풀을 내린 채 얼굴을 박철진 쪽으로 돌리고 있다. 대면 방향은 맞지만 두 사람의 얼굴과 무릎 사이에 상당한 간격이 남아 있어 의자를 바짝 당겼다는 관계는 약하다.",
        "built_space": "마주 보는 목재·금속 의자 두 개에 각각 한 사람이 앉아 있고, 뒤에는 열린 회색 문 하나와 복도가 보인다. 두 가지 회색의 벽, 회색 바닥, 천장 조명 하나, 환기구 하나와 원형 장치가 참조의 공간 구성을 유지한다. 추가 책상이나 인물은 없다. 박철진의 상체는 A보다 크게 잡혔지만 좌판과 등받이가 넓게 드러나므로 의자 옆 가장자리만 부분적으로 보이라는 조건에는 덜 정확하다. 외부가 보이지 않아 밤 여부는 직접 확인할 수 없으며 밝은 실내조명은 맞는다.",
        "entities": "중년의 한국인 남녀로 읽히는 두 사람만 등장한다. 미연의 검은 곱슬머리, 체크 셔츠, 갈색 앞치마와 올리브색 바지는 참조에 가깝고, 볼과 눈 주변의 멍과 부기가 보인다. 참조의 모자는 없다. 박철진의 얼굴과 짧은 검은 머리, 검은 제복, 붉은 완장과 허리 장비는 대체로 일치한다. 읽을 수 있는 문구나 자막은 없다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 각자의 좌판에 실려 있고 등받이는 뒤쪽에 놓인다. 박철진은 양손을 허벅지에 얹고 골반에서 앞으로 기울여 앉아 있으며, 미연의 팔과 손은 무릎 가까이에 놓인다. 무릎은 자연스럽게 굽혀지고 의자 다리는 바닥으로 이어진다. 허리 장비는 벨트에 부착되어 있으며, 지지 없이 떠 있는 신체나 물건은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "바짝 다가앉은 거리와 거만한 표정은 잘 드러나지만, 없는 책상과 문구류를 추가했고 상체 중심보다 넓은 구도와 미연의 바뀐 복장이 지시에서 벗어난다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "장소와 인물 복장을 더 충실히 유지하고 박철진의 상체도 크게 잡았지만, 의자 사이가 넓어 ‘바짝 당겨 앉은’ 압박감이 부족하고 의자 노출도 과하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 박철진은 왼쪽 미연의 얼굴을 바라보며 상체를 앞으로 기울인다. 미연도 박철진 쪽으로 얼굴과 시선을 향한다. 서로의 무릎이 거의 맞닿아 대면 심문의 방향과 침범하는 거리가 명확하다.",
        "built_space": "목재 좌판과 검은 금속 골조의 의자 두 개에 한 사람씩 마주 앉았다. 오른쪽 열린 회색 문 하나, 밝고 어두운 회색으로 나뉜 벽, 회색 바닥, 천장의 직사각 조명 하나와 환기구 하나 및 작은 원형 장치가 보인다. 주요 건축 요소는 참조와 유사하지만, 두 사람 뒤 중앙에 참조에 없는 책상 하나가 추가되었다. 박철진의 의자는 옆 가장자리 일부가 아니라 좌판과 등받이 상당 부분이 드러나며, 인물도 무릎 아래까지 보여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "중년의 한국인 남녀로 읽히는 두 사람만 보인다. 박철진의 짧은 검은 머리, 얼굴, 검은 제복과 붉은 완장은 참조와 대체로 맞는다. 미연은 얼굴에 멍과 눈 주변 부기가 있지만, 참조의 짧은 곱슬머리 대신 묶은 긴 머리이며 체크 셔츠와 갈색 앞치마 대신 검은 작업복을 입었다. 모자도 없다. 추가된 책상에는 종이와 필기구가 보인다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "지정 장소의 참조에 없고 장면에서도 요구하지 않은 책상과 그 위의 종이·필기구를 추가했다."
        ],
        "physics": "두 사람 모두 엉덩이가 각자의 의자 좌판에 놓이고 등받이는 몸 뒤에 있다. 박철진은 양손을 허벅지에 대고 앞으로 숙여 체중을 지탱하며, 미연의 손은 무릎 위에 놓인다. 의자 다리는 바닥으로 이어지고 책상 위 물건도 상판에 놓여 있다. 발 일부가 화면 밖인 것은 지지 불가능의 증거가 아니며, 떠 있는 신체나 물건은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "오른쪽 박철진은 미연의 얼굴을 똑바로 응시하고 상체를 그녀 쪽으로 숙인다. 왼쪽 미연은 눈꺼풀을 내린 채 얼굴을 박철진 쪽으로 돌리고 있다. 대면 방향은 맞지만 두 사람의 얼굴과 무릎 사이에 상당한 간격이 남아 있어 의자를 바짝 당겼다는 관계는 약하다.",
        "built_space": "마주 보는 목재·금속 의자 두 개에 각각 한 사람이 앉아 있고, 뒤에는 열린 회색 문 하나와 복도가 보인다. 두 가지 회색의 벽, 회색 바닥, 천장 조명 하나, 환기구 하나와 원형 장치가 참조의 공간 구성을 유지한다. 추가 책상이나 인물은 없다. 박철진의 상체는 A보다 크게 잡혔지만 좌판과 등받이가 넓게 드러나므로 의자 옆 가장자리만 부분적으로 보이라는 조건에는 덜 정확하다. 외부가 보이지 않아 밤 여부는 직접 확인할 수 없으며 밝은 실내조명은 맞는다.",
        "entities": "중년의 한국인 남녀로 읽히는 두 사람만 등장한다. 미연의 검은 곱슬머리, 체크 셔츠, 갈색 앞치마와 올리브색 바지는 참조에 가깝고, 볼과 눈 주변의 멍과 부기가 보인다. 참조의 모자는 없다. 박철진의 얼굴과 짧은 검은 머리, 검은 제복, 붉은 완장과 허리 장비는 대체로 일치한다. 읽을 수 있는 문구나 자막은 없다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 각자의 좌판에 실려 있고 등받이는 뒤쪽에 놓인다. 박철진은 양손을 허벅지에 얹고 골반에서 앞으로 기울여 앉아 있으며, 미연의 팔과 손은 무릎 가까이에 놓인다. 무릎은 자연스럽게 굽혀지고 의자 다리는 바닥으로 이어진다. 허리 장비는 벨트에 부착되어 있으며, 지지 없이 떠 있는 신체나 물건은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 레퍼런스 장소에 존재하지 않는 사물(책상 및 서류)이 임의로 추가됨",
     "[gpt-high] 지정 장소의 참조에 없고 장면에서도 요구하지 않은 책상과 그 위의 종이·필기구를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 미디엄 샷과 의자 크롭 프레이밍을 따르지 않고 풀 샷에 가깝게 렌더링되었으나, 두 인물의 외모와 의상, 장소 레퍼런스는 매우 충실하게 재현되었습니다."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "지정된 프레이밍을 어겼을 뿐만 아니라, 장소에 없는 사물을 임의로 추가하고 인물의 의상 레퍼런스를 무시하는 등 다수의 위반 사항이 확인됩니다.  ★위반: [gemini-pro] 레퍼런스 장소에 존재하지 않는 사물(책상 및 서류)이 임의로 추가됨 / [gpt-high] 지정 장소의 참조에 없고 장면에서도 요구하지 않은 책상과 그 위의 종이·필기구를 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L189B01.png",
    "asset_id": "d31b4d4c-5488-4514-8d64-f7f3972f4a12",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-398d-7b05-8d86-ac672889b5bc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2__bgfirst_bg.png",
   "bg_asset_id": "6d0e961f-cfc1-4ccb-921e-40cb43da0593",
   "bg_record_key": "S31sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S31sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:23:03.672816+00:00",
  "fingerprint": "5a8528faa9c14f3165de446c49f017758af9dafd88c91da67530fc8123759b50",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S31sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S31sh2_sel.png",
  "source_sha256": "cbe51d80bb7424af58fa3ce195afcb70f1919b1a5d54b2a71ef89d135ec8489b",
  "file": "S31sh2_cine.png",
  "staged_sha256": "ddcf9b40f850e5c7150870d837b3ea726e650cda6abb6503070634b32600f5ba",
  "latency_ms": 9313
 },
 "S31sh8::signage": {
  "fp": "193145b90fc8ec5f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S31sh8": {
  "input_fingerprint": "15fba1138b520bd0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수하 1의 귓속말에 놀라 눈이 커진 채 자리에서 반쯤 일어선 박철진의 엉거주춤한 자세.\n\nLOCATION (lock): Beside the interrogation chair inside the illuminated third-floor room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Partly vacated as he interrupts his rise) — The seat is visible obliquely below and behind his hips; used as The visible separation between hips and seat proves that he has not fully stood.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination consistent with the questioning shot, preserving facial detail through the incomplete rise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same confinement-room surfaces, interrogation chair, and nighttime interior lighting. Exclude equipment or furniture from the separate militia office.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, and its door has been opened for the report. 박철진: He remains at the interrogation position and reacts angrily to the report. 수하 1: He has entered the interrogation room and is delivering a whispered report, wearing a militia uniform and armband.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수하 1의 귓속말에 놀라 눈이 커진 채 자리에서 반쯤 일어선 박철진의 엉거주춤한 자세.\n\nLOCATION (lock): Beside the interrogation chair inside the illuminated third-floor room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Partly vacated as he interrupts his rise) — The seat is visible obliquely below and behind his hips; used as The visible separation between hips and seat proves that he has not fully stood.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination consistent with the questioning shot, preserving facial detail through the incomplete rise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same confinement-room surfaces, interrogation chair, and nighttime interior lighting. Exclude equipment or furniture from the separate militia office.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, and its door has been opened for the report. 박철진: He remains at the interrogation position and reacts angrily to the report. 수하 1: He has entered the interrogation room and is delivering a whispered report, wearing a militia uniform and armband.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수하 1의 귓속말에 놀라 눈이 커진 채 자리에서 반쯤 일어선 박철진의 엉거주춤한 자세.\n\nLOCATION (lock): Beside the interrogation chair inside the illuminated third-floor room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Partly vacated as he interrupts his rise) — The seat is visible obliquely below and behind his hips; used as The visible separation between hips and seat proves that he has not fully stood.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination consistent with the questioning shot, preserving facial detail through the incomplete rise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same confinement-room surfaces, interrogation chair, and nighttime interior lighting. Exclude equipment or furniture from the separate militia office.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, and its door has been opened for the report. 박철진: He remains at the interrogation position and reacts angrily to the report. 수하 1: He has entered the interrogation room and is delivering a whispered report, wearing a militia uniform and armband.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "오른쪽 남성(박철진)이 왼쪽 남성(수하 1)의 귀를 향해 시선과 입을 맞추고 있으며, 왼쪽 남성은 정면을 향해 놀란 눈빛을 보임.",
    "built_space": "취조실 내부. 뒤쪽 벽과 열린 문이 보임. 왼쪽 남성 뒤편에 의자가 하나 있고, 오른쪽 끝에 다른 의자의 일부가 보임.",
    "entities": "왼쪽 남성은 수하 1의 얼굴이며 민병대 완장을 착용함. 오른쪽 남성은 박철진의 얼굴과 복장임. 귓속말을 하는 자와 놀라는 자의 역할이 완전히 반대로 적용됨.",
    "hard_violations": [
     "[gemini-pro] 오른쪽 남성(박철진)이 의자가 없음에도 공중에 떠서 앉아 있는 물리적으로 불가능한 자세(투명 의자 포즈)를 취하고 있음."
    ],
    "physics": "왼쪽 남성은 의자 위에서 엉거주춤 떠 있으며 두 발이 바닥에 닿아 있음. 오른쪽 남성은 받쳐주는 의자가 전혀 없으나 기마 자세로 공중에 떠서 체중을 지탱하는 불가능한 포즈임."
   },
   {
    "label": "B",
    "direction": "오른쪽 남성이 왼쪽 남성(박철진)의 귀에 대고 귓속말을 하고, 왼쪽 남성은 놀란 표정으로 시선을 밖으로 향함.",
    "built_space": "취조실 내부. 열린 문과 벽이 보임. 지정된 의자가 인물의 엉덩이 아래나 뒤가 아닌, 두 인물 사이의 엉뚱한 공간(왼쪽 남성 다리 뒤)에 놓여 있음.",
    "entities": "왼쪽 남성은 박철진의 얼굴과 완장 복장과 일치함. 오른쪽 남성은 측면만 보여 정확한 신원 파악이 어려우나 수하 1의 역할을 수행 중임.",
    "hard_violations": [
     "[gemini-pro] 왼쪽 남성의 왼손이 지탱할 표면이 전혀 없는 화면 아래쪽 허공(보이지 않는 탁자)을 평평하게 짚고 있는 물리적 위반."
    ],
    "physics": "왼쪽 남성은 자리에서 반쯤 일어선 자세이나, 왼손이 아무것도 없는 허공의 표면을 짚고 체중을 싣고 있음. 오른쪽 남성은 몸을 기울이고 있으며 다리 지탱은 보이지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "박철진이 놀라 일어나는 핵심 상황과 인물 배치는 구성했으나, 왼손이 허공의 투명한 표면을 짚고 있는 물리적 오류가 발생했고 의자 위치가 요구사항을 벗어남."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "박철진과 수하 1의 역할이 완전히 뒤바뀌어 프롬프트의 지시를 정면으로 위반했으며, 오른쪽 인물이 투명 의자에 앉은 듯한 불가능한 자세를 보임."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 남성(박철진)이 왼쪽 남성(수하 1)의 귀를 향해 시선과 입을 맞추고 있으며, 왼쪽 남성은 정면을 향해 놀란 눈빛을 보임.",
        "built_space": "취조실 내부. 뒤쪽 벽과 열린 문이 보임. 왼쪽 남성 뒤편에 의자가 하나 있고, 오른쪽 끝에 다른 의자의 일부가 보임.",
        "entities": "왼쪽 남성은 수하 1의 얼굴이며 민병대 완장을 착용함. 오른쪽 남성은 박철진의 얼굴과 복장임. 귓속말을 하는 자와 놀라는 자의 역할이 완전히 반대로 적용됨.",
        "hard_violations": [
         "오른쪽 남성(박철진)이 의자가 없음에도 공중에 떠서 앉아 있는 물리적으로 불가능한 자세(투명 의자 포즈)를 취하고 있음."
        ],
        "physics": "왼쪽 남성은 의자 위에서 엉거주춤 떠 있으며 두 발이 바닥에 닿아 있음. 오른쪽 남성은 받쳐주는 의자가 전혀 없으나 기마 자세로 공중에 떠서 체중을 지탱하는 불가능한 포즈임."
       },
       {
        "label": "B",
        "direction": "오른쪽 남성이 왼쪽 남성(박철진)의 귀에 대고 귓속말을 하고, 왼쪽 남성은 놀란 표정으로 시선을 밖으로 향함.",
        "built_space": "취조실 내부. 열린 문과 벽이 보임. 지정된 의자가 인물의 엉덩이 아래나 뒤가 아닌, 두 인물 사이의 엉뚱한 공간(왼쪽 남성 다리 뒤)에 놓여 있음.",
        "entities": "왼쪽 남성은 박철진의 얼굴과 완장 복장과 일치함. 오른쪽 남성은 측면만 보여 정확한 신원 파악이 어려우나 수하 1의 역할을 수행 중임.",
        "hard_violations": [
         "왼쪽 남성의 왼손이 지탱할 표면이 전혀 없는 화면 아래쪽 허공(보이지 않는 탁자)을 평평하게 짚고 있는 물리적 위반."
        ],
        "physics": "왼쪽 남성은 자리에서 반쯤 일어선 자세이나, 왼손이 아무것도 없는 허공의 표면을 짚고 체중을 싣고 있음. 오른쪽 남성은 몸을 기울이고 있으며 다리 지탱은 보이지 않음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "박철진이 놀라 일어나는 핵심 상황과 인물 배치는 구성했으나, 왼손이 허공의 투명한 표면을 짚고 있는 물리적 오류가 발생했고 의자 위치가 요구사항을 벗어남."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "박철진과 수하 1의 역할이 완전히 뒤바뀌어 프롬프트의 지시를 정면으로 위반했으며, 오른쪽 인물이 투명 의자에 앉은 듯한 불가능한 자세를 보임."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 남성(박철진)이 왼쪽 남성(수하 1)의 귀를 향해 시선과 입을 맞추고 있으며, 왼쪽 남성은 정면을 향해 놀란 눈빛을 보임.",
        "built_space": "취조실 내부. 뒤쪽 벽과 열린 문이 보임. 왼쪽 남성 뒤편에 의자가 하나 있고, 오른쪽 끝에 다른 의자의 일부가 보임.",
        "entities": "왼쪽 남성은 수하 1의 얼굴이며 민병대 완장을 착용함. 오른쪽 남성은 박철진의 얼굴과 복장임. 귓속말을 하는 자와 놀라는 자의 역할이 완전히 반대로 적용됨.",
        "hard_violations": [
         "오른쪽 남성(박철진)이 의자가 없음에도 공중에 떠서 앉아 있는 물리적으로 불가능한 자세(투명 의자 포즈)를 취하고 있음."
        ],
        "physics": "왼쪽 남성은 의자 위에서 엉거주춤 떠 있으며 두 발이 바닥에 닿아 있음. 오른쪽 남성은 받쳐주는 의자가 전혀 없으나 기마 자세로 공중에 떠서 체중을 지탱하는 불가능한 포즈임."
       },
       {
        "label": "B",
        "direction": "오른쪽 남성이 왼쪽 남성(박철진)의 귀에 대고 귓속말을 하고, 왼쪽 남성은 놀란 표정으로 시선을 밖으로 향함.",
        "built_space": "취조실 내부. 열린 문과 벽이 보임. 지정된 의자가 인물의 엉덩이 아래나 뒤가 아닌, 두 인물 사이의 엉뚱한 공간(왼쪽 남성 다리 뒤)에 놓여 있음.",
        "entities": "왼쪽 남성은 박철진의 얼굴과 완장 복장과 일치함. 오른쪽 남성은 측면만 보여 정확한 신원 파악이 어려우나 수하 1의 역할을 수행 중임.",
        "hard_violations": [
         "왼쪽 남성의 왼손이 지탱할 표면이 전혀 없는 화면 아래쪽 허공(보이지 않는 탁자)을 평평하게 짚고 있는 물리적 위반."
        ],
        "physics": "왼쪽 남성은 자리에서 반쯤 일어선 자세이나, 왼손이 아무것도 없는 허공의 표면을 짚고 체중을 싣고 있음. 오른쪽 남성은 몸을 기울이고 있으며 다리 지탱은 보이지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "귓속말에 놀란 표정과 가까운 인물 크기는 맞지만, 좌판이 거의 잘려 반쯤 일어선 상태를 입증하는 핵심 구도가 부족하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "다소 넓은 구도이지만 좌판과 굽힌 무릎을 함께 보여 중단된 기립 동작이 더 명확하고, 심문실 구조와 수하의 완장도 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 수하는 입을 박철진의 화면 오른쪽 귀 가까이 대고 속삭인다. 박철진은 눈을 크게 뜨고 화면 오른쪽 앞쪽의 프레임 밖을 바라본다. 귓속말의 전달 방향은 맞으며, 무기나 겨누는 물체는 없다.",
        "built_space": "밝은 상부와 회색 하부로 나뉜 벽, 콘크리트 바닥, 오른쪽의 열린 문 한 개가 보인다. 천장에는 직사각형 조명 두 개와 사각 환기구 한 개가 보인다. 박철진 뒤에는 목재 등받이 의자 한 개가 있으나 좌판은 하단에서 거의 잘린다. 하단 왼쪽에도 손이 닿는 목재 면 일부가 보이지만 그 전체 형태는 확인되지 않는다. 참조와 방의 재료는 유사하지만 좌판과 엉덩이의 관계를 보여야 한다는 구도 조건은 약하다.",
        "entities": "짧은 검은 머리의 한국인 성인 남성으로 읽히는 두 사람만 등장한다. 박철진은 검은 민병대 복장, 허리 장비, 붉은 완장을 착용하지만 참조보다 얼굴이 둥글고 젊어 보인다. 수하도 검은 제복과 허리 장비를 착용하며, 옆얼굴만 보여 신원 일치 판단에는 한계가 있다. 수하의 완장은 보이지 않는다. 박철진 완장의 검은 도형은 참조의 단색 완장과 다르다. 이전 장면의 여성이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "박철진은 상체를 숙이고 무릎을 굽혔으며, 화면 왼쪽 손바닥을 하단 목재 면에 짚어 지지한다. 다른 손은 허리 장비 근처에 있다. 발은 화면 밖이지만 다리와 손의 자세로 일어서다 멈춘 동작은 가능하다. 수하는 다리를 아래로 내린 채 상체만 기울인다. 공중에 떠 있는 몸이나 장비는 없다. 다만 의자 좌판이 잘려 실제 기립 높이는 판단하기 어렵다."
       },
       {
        "label": "B",
        "direction": "오른쪽 수하는 손을 입 옆에 모으고 박철진의 화면 오른쪽 귀에 입을 가까이 대어 보고한다. 박철진은 눈과 입을 벌린 채 화면 오른쪽 앞쪽을 바라본다. 수하의 말이 향하는 대상이 명확하며, 카메라를 향한 의도적인 응시는 아니다.",
        "built_space": "참조와 같은 두 색의 벽, 콘크리트 바닥, 오른쪽의 열린 문 한 개가 보인다. 천장에는 직사각형 조명 한 개, 사각 환기구 한 개, 작은 원형 부속 한 개가 보인다. 왼쪽에는 목재 좌판과 등받이의 의자 한 개, 오른쪽 가장자리에는 다른 의자 일부가 보여 참조의 두 의자와 양립한다. 박철진 뒤아래 좌판이 비스듬히 드러나며, 엉덩이가 좌판 앞쪽 위로 이동한 상태를 확인할 수 있다. 인물의 무릎까지 들어와 요청한 미디엄 숏보다 약간 넓다.",
        "entities": "박철진과 수하로 읽히는 짧은 검은 머리의 한국인 성인 남성 두 명만 보인다. 두 사람 모두 검은 민병대 제복과 붉은 완장을 착용하고 허리 장비를 지닌다. 박철진은 참조보다 다소 젊고 둥근 얼굴로 표현되며, 수하는 측면이라 얼굴 전체의 일치를 확정하기 어렵다. 박철진 완장의 검은 도형은 참조와 다르다. 참조의 여성, 별도 사무실 가구, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "박철진은 양 무릎을 굽히고 손을 각각 허벅지와 무릎에 대어 앞으로 체중을 옮긴다. 엉덩이와 좌판의 간격은 크지 않지만 완전히 선 자세가 아니라 일어나던 중 멈춘 자세로 읽힌다. 수하는 무릎을 굽혀 귀 높이를 맞추고 한 손을 자기 허벅지에 댄다. 발 접지는 하단 밖이지만 두 사람의 다리와 몸통 정렬은 가능한 자세이며, 허리 장비는 벨트에 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "귓속말에 놀란 표정과 가까운 인물 크기는 맞지만, 좌판이 거의 잘려 반쯤 일어선 상태를 입증하는 핵심 구도가 부족하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "다소 넓은 구도이지만 좌판과 굽힌 무릎을 함께 보여 중단된 기립 동작이 더 명확하고, 심문실 구조와 수하의 완장도 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 수하는 입을 박철진의 화면 오른쪽 귀 가까이 대고 속삭인다. 박철진은 눈을 크게 뜨고 화면 오른쪽 앞쪽의 프레임 밖을 바라본다. 귓속말의 전달 방향은 맞으며, 무기나 겨누는 물체는 없다.",
        "built_space": "밝은 상부와 회색 하부로 나뉜 벽, 콘크리트 바닥, 오른쪽의 열린 문 한 개가 보인다. 천장에는 직사각형 조명 두 개와 사각 환기구 한 개가 보인다. 박철진 뒤에는 목재 등받이 의자 한 개가 있으나 좌판은 하단에서 거의 잘린다. 하단 왼쪽에도 손이 닿는 목재 면 일부가 보이지만 그 전체 형태는 확인되지 않는다. 참조와 방의 재료는 유사하지만 좌판과 엉덩이의 관계를 보여야 한다는 구도 조건은 약하다.",
        "entities": "짧은 검은 머리의 한국인 성인 남성으로 읽히는 두 사람만 등장한다. 박철진은 검은 민병대 복장, 허리 장비, 붉은 완장을 착용하지만 참조보다 얼굴이 둥글고 젊어 보인다. 수하도 검은 제복과 허리 장비를 착용하며, 옆얼굴만 보여 신원 일치 판단에는 한계가 있다. 수하의 완장은 보이지 않는다. 박철진 완장의 검은 도형은 참조의 단색 완장과 다르다. 이전 장면의 여성이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "박철진은 상체를 숙이고 무릎을 굽혔으며, 화면 왼쪽 손바닥을 하단 목재 면에 짚어 지지한다. 다른 손은 허리 장비 근처에 있다. 발은 화면 밖이지만 다리와 손의 자세로 일어서다 멈춘 동작은 가능하다. 수하는 다리를 아래로 내린 채 상체만 기울인다. 공중에 떠 있는 몸이나 장비는 없다. 다만 의자 좌판이 잘려 실제 기립 높이는 판단하기 어렵다."
       },
       {
        "label": "A",
        "direction": "오른쪽 수하는 손을 입 옆에 모으고 박철진의 화면 오른쪽 귀에 입을 가까이 대어 보고한다. 박철진은 눈과 입을 벌린 채 화면 오른쪽 앞쪽을 바라본다. 수하의 말이 향하는 대상이 명확하며, 카메라를 향한 의도적인 응시는 아니다.",
        "built_space": "참조와 같은 두 색의 벽, 콘크리트 바닥, 오른쪽의 열린 문 한 개가 보인다. 천장에는 직사각형 조명 한 개, 사각 환기구 한 개, 작은 원형 부속 한 개가 보인다. 왼쪽에는 목재 좌판과 등받이의 의자 한 개, 오른쪽 가장자리에는 다른 의자 일부가 보여 참조의 두 의자와 양립한다. 박철진 뒤아래 좌판이 비스듬히 드러나며, 엉덩이가 좌판 앞쪽 위로 이동한 상태를 확인할 수 있다. 인물의 무릎까지 들어와 요청한 미디엄 숏보다 약간 넓다.",
        "entities": "박철진과 수하로 읽히는 짧은 검은 머리의 한국인 성인 남성 두 명만 보인다. 두 사람 모두 검은 민병대 제복과 붉은 완장을 착용하고 허리 장비를 지닌다. 박철진은 참조보다 다소 젊고 둥근 얼굴로 표현되며, 수하는 측면이라 얼굴 전체의 일치를 확정하기 어렵다. 박철진 완장의 검은 도형은 참조와 다르다. 참조의 여성, 별도 사무실 가구, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "박철진은 양 무릎을 굽히고 손을 각각 허벅지와 무릎에 대어 앞으로 체중을 옮긴다. 엉덩이와 좌판의 간격은 크지 않지만 완전히 선 자세가 아니라 일어나던 중 멈춘 자세로 읽힌다. 수하는 무릎을 굽혀 귀 높이를 맞추고 한 손을 자기 허벅지에 댄다. 발 접지는 하단 밖이지만 두 사람의 다리와 몸통 정렬은 가능한 자세이며, 허리 장비는 벨트에 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 오른쪽 남성(박철진)이 의자가 없음에도 공중에 떠서 앉아 있는 물리적으로 불가능한 자세(투명 의자 포즈)를 취하고 있음."
    ],
    "B": [
     "[gemini-pro] 왼쪽 남성의 왼손이 지탱할 표면이 전혀 없는 화면 아래쪽 허공(보이지 않는 탁자)을 평평하게 짚고 있는 물리적 위반."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1500,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "박철진이 놀라 일어나는 핵심 상황과 인물 배치는 구성했으나, 왼손이 허공의 투명한 표면을 짚고 있는 물리적 오류가 발생했고 의자 위치가 요구사항을 벗어남.  ★위반: [gemini-pro] 왼쪽 남성의 왼손이 지탱할 표면이 전혀 없는 화면 아래쪽 허공(보이지 않는 탁자)을 평평하게 짚고 있는 물리적 위반."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "박철진과 수하 1의 역할이 완전히 뒤바뀌어 프롬프트의 지시를 정면으로 위반했으며, 오른쪽 인물이 투명 의자에 앉은 듯한 불가능한 자세를 보임.  ★위반: [gemini-pro] 오른쪽 남성(박철진)이 의자가 없음에도 공중에 떠서 앉아 있는 물리적으로 불가능한 자세(투명 의자 포즈)를 취하고 있음."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2_sel.png",
    "asset_id": "481c048a-d274-48ab-9500-1507d3701c08",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1192369>",
    "asset_id": "a969dd53-900f-4bf4-b80a-e38b77b03fbf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-3cca-7515-b439-7b59451958df",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S31sh2"
  },
  "staged_characters_added": [
   "C17"
  ]
 },
 "S31sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:24:40.189263+00:00",
  "fingerprint": "3ec36e3bbbe23ef67c0828b5f2a7a71115df2c5caaf81ca23fab5586d0762cbb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S31sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S31sh8_sel.png",
  "source_sha256": "d6d2f3c1c62ba21b11e09f522115f038b4f8bdd4013fbcab55edb27d549fedb4",
  "file": "S31sh8_cine.png",
  "staged_sha256": "8ccf313008c42678f17434f25230519ba39e666d3902e4ac16bcdeedf944e963",
  "latency_ms": 13114
 },
 "S31sh9::signage": {
  "fp": "42982951c83501c8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S31sh9": {
  "input_fingerprint": "ed138f2b7ca206bc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 성당이라는 말에 충격을 받은 듯 두 눈이 동그라진 미연의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the detainee's seat inside the illuminated third-floor interrogation room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 미연 alone in the close frame in the middle-left of the frame, foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office illumination keeps the swelling and suddenly widened eyes legible without glamorizing the injury.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same confinement-room surfaces, chair, and interior lighting. Exclude the command-office desk, speakerphone, and wall-mounted knife from the separate office scene.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, with the door opened during the report. 미연: Her face remains badly swollen from the beating, and she is visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 성당이라는 말에 충격을 받은 듯 두 눈이 동그라진 미연의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the detainee's seat inside the illuminated third-floor interrogation room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 미연 alone in the close frame in the middle-left of the frame, foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office illumination keeps the swelling and suddenly widened eyes legible without glamorizing the injury.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same confinement-room surfaces, chair, and interior lighting. Exclude the command-office desk, speakerphone, and wall-mounted knife from the separate office scene.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, with the door opened during the report. 미연: Her face remains badly swollen from the beating, and she is visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 성당이라는 말에 충격을 받은 듯 두 눈이 동그라진 미연의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the detainee's seat inside the illuminated third-floor interrogation room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 미연 alone in the close frame in the middle-left of the frame, foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office illumination keeps the swelling and suddenly widened eyes legible without glamorizing the injury.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same confinement-room surfaces, chair, and interior lighting. Exclude the command-office desk, speakerphone, and wall-mounted knife from the separate office scene.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, with the door opened during the report. 미연: Her face remains badly swollen from the beating, and she is visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 우측 아래를 향하며 두 눈을 크게 뜸.",
    "built_space": "취조실 벽면. 피사체 뒤에 둥근 금속 프레임의 의자가 있으나 레퍼런스와 형태가 다름.",
    "entities": "미연(체크 셔츠, 앞치마 착용). 심하게 부어오른 얼굴(badly swollen) 묘사가 턱없이 부족함.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스와 다른 형태의 의자(둥근 금속 프레임) 배치로 장소 고정(LOCATION lock) 위반",
     "[gemini-pro] 피사체를 정중앙에 배치하여 '화면 중앙 좌측' 구도 지시 위반"
    ],
    "physics": "자연스러운 자세로 몸을 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 좌측 허공을 향하며 충격받은 듯 눈을 크게 뜸.",
    "built_space": "취조실 투톤 벽면. 레퍼런스와 일치하는 형태의 의자가 피사체와 자연스럽게 배치됨.",
    "entities": "미연(체크 셔츠 착용, 앞치마는 보이지 않음). 폭행으로 심하게 부어오른 얼굴과 상처가 지시대로 잘 표현됨.",
    "hard_violations": [],
    "physics": "의자에 앉아 상체를 안정적으로 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '화면 중앙 좌측(middle-left)' 구도를 무시하고 피사체를 중앙에 배치했으며, 얼굴의 심한 붓기가 묘사되지 않았고 레퍼런스와 다른 형태의 의자가 등장함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 '화면 중앙 좌측' 구도를 정확히 준수하고 심하게 부어오른 얼굴과 충격받은 표정을 훌륭하게 구현했으나 앞치마가 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 우측 아래를 향하며 두 눈을 크게 뜸.",
        "built_space": "취조실 벽면. 피사체 뒤에 둥근 금속 프레임의 의자가 있으나 레퍼런스와 형태가 다름.",
        "entities": "미연(체크 셔츠, 앞치마 착용). 심하게 부어오른 얼굴(badly swollen) 묘사가 턱없이 부족함.",
        "hard_violations": [
         "레퍼런스와 다른 형태의 의자(둥근 금속 프레임) 배치로 장소 고정(LOCATION lock) 위반",
         "피사체를 정중앙에 배치하여 '화면 중앙 좌측' 구도 지시 위반"
        ],
        "physics": "자연스러운 자세로 몸을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 좌측 허공을 향하며 충격받은 듯 눈을 크게 뜸.",
        "built_space": "취조실 투톤 벽면. 레퍼런스와 일치하는 형태의 의자가 피사체와 자연스럽게 배치됨.",
        "entities": "미연(체크 셔츠 착용, 앞치마는 보이지 않음). 폭행으로 심하게 부어오른 얼굴과 상처가 지시대로 잘 표현됨.",
        "hard_violations": [],
        "physics": "의자에 앉아 상체를 안정적으로 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '화면 중앙 좌측(middle-left)' 구도를 무시하고 피사체를 중앙에 배치했으며, 얼굴의 심한 붓기가 묘사되지 않았고 레퍼런스와 다른 형태의 의자가 등장함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 '화면 중앙 좌측' 구도를 정확히 준수하고 심하게 부어오른 얼굴과 충격받은 표정을 훌륭하게 구현했으나 앞치마가 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 우측 아래를 향하며 두 눈을 크게 뜸.",
        "built_space": "취조실 벽면. 피사체 뒤에 둥근 금속 프레임의 의자가 있으나 레퍼런스와 형태가 다름.",
        "entities": "미연(체크 셔츠, 앞치마 착용). 심하게 부어오른 얼굴(badly swollen) 묘사가 턱없이 부족함.",
        "hard_violations": [
         "레퍼런스와 다른 형태의 의자(둥근 금속 프레임) 배치로 장소 고정(LOCATION lock) 위반",
         "피사체를 정중앙에 배치하여 '화면 중앙 좌측' 구도 지시 위반"
        ],
        "physics": "자연스러운 자세로 몸을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 좌측 허공을 향하며 충격받은 듯 눈을 크게 뜸.",
        "built_space": "취조실 투톤 벽면. 레퍼런스와 일치하는 형태의 의자가 피사체와 자연스럽게 배치됨.",
        "entities": "미연(체크 셔츠 착용, 앞치마는 보이지 않음). 폭행으로 심하게 부어오른 얼굴과 상처가 지시대로 잘 표현됨.",
        "hard_violations": [],
        "physics": "의자에 앉아 상체를 안정적으로 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "중앙 왼쪽 배치와 기존 의자는 맞지만, 허리까지 보이는 넓은 구도가 핵심인 얼굴 클로즈업을 놓치며 앞치마도 빠졌다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "얼굴을 더 크게 담고 갑자기 동그래진 눈과 굳은 반응을 선명히 구현해 우세하지만, 중앙 오른쪽 배치와 달라진 의자, 충분히 좁지 않은 구도는 불일치한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "몸은 화면 오른쪽으로 비스듬히 앉아 있고 얼굴은 카메라 쪽으로 돌아와 있다. 두 눈은 화면 왼쪽의 프레임 밖을 향하며, 바라보는 상대는 보이지 않는다. 특정 시선 대상이 명시되지 않은 충격 반응으로는 가능하다. 무기나 방향성 소품은 없다.",
        "built_space": "밝은 상부와 회색 하부로 나뉜 벽 두 면, 모서리 하나, 걸레받이와 회색 바닥이 보인다. 여성 뒤에는 갈색 목재 등판과 검은 금속 틀의 의자 하나가 있어 이전 장면의 의자와 잘 맞는다. 여성의 등 뒤에 등받이가 놓인 배치도 자연스럽다. 문과 천장 조명은 프레임 밖이므로 개폐 상태나 개수를 확인할 수 없다. 반사는 없다.",
        "entities": "중년 한국인 여성으로 보이는 인물 한 명만 있으며 검은 단발과 얼굴 윤곽, 갈색 계열 체크 셔츠가 인물 참조와 대체로 맞는다. 참조의 모자는 없고, 앞치마가 보여야 할 상체에는 셔츠만 있다. 양 볼과 눈 밑의 붉은 멍은 보이지만 심하게 부어오른 얼굴이라는 조건에 비하면 부기가 약하다. 눈은 정상적인 해부학적 형태로 커져 있고 입이 살짝 벌어져 있다. 금지된 사무실 소품이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "몸통은 의자에 앉은 자세로 수직에 가깝게 유지되며 등받이가 몸 뒤에 보인다. 실제 좌판과 골반 접촉은 하단 크롭에 가려져 있지만, 의자와 몸의 관계에 부유나 불가능한 지지는 없다. 팔은 몸 옆으로 내려가며 손은 프레임 밖이다. 공중에 뜬 물체나 운동 중인 신체는 없다."
       },
       {
        "label": "B",
        "direction": "몸과 얼굴은 거의 정면이고 두 눈은 렌즈보다 화면 왼쪽의 가까운 프레임 밖 지점을 향한다. 대상 인물은 보이지 않는다. 눈꺼풀이 크게 열리고 입이 벌어진 모습은 말을 듣고 놀란 순간으로 읽힌다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "두 색으로 칠한 벽 두 면과 모서리 하나, 걸레받이, 회색 바닥이 이어진다. 오른쪽 끝에는 열린 출입구 일부가 보이고 벽에는 스위치판 하나가 있다. 등 뒤의 목재 등받이는 하나로 이어져 보이나, 넓고 둥근 밝은 회색 금속 틀은 참조의 좁고 검은 의자 틀과 다르다. 인물은 등받이 앞에 자연스럽게 자리한다. 거울이나 반사는 없다.",
        "entities": "검은 단발의 중년 한국인 여성 한 명으로, 얼굴과 체크 셔츠가 참조에 대체로 부합한다. 갈색 앞치마와 가슴 주머니의 연필 및 끈도 유지된다. 모자는 없다. 눈 밑과 볼에 붉은 변색 및 약한 부기가 있지만 심한 구타 후의 부종은 충분히 드러나지 않는다. 크게 뜬 눈의 홍채와 동공은 정상이며 추가 인물, 금지된 사무실 소품, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "상체는 등받이 앞에서 곧게 유지되고 어깨와 목의 연결도 자연스럽다. 좌판과 골반은 프레임 밖이지만 앉은 자세와 의자 위치가 모순되지 않는다. 앞치마는 어깨끈으로 지지되고 연필과 끈은 주머니에 걸쳐 있다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "중앙 왼쪽 배치와 기존 의자는 맞지만, 허리까지 보이는 넓은 구도가 핵심인 얼굴 클로즈업을 놓치며 앞치마도 빠졌다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "얼굴을 더 크게 담고 갑자기 동그래진 눈과 굳은 반응을 선명히 구현해 우세하지만, 중앙 오른쪽 배치와 달라진 의자, 충분히 좁지 않은 구도는 불일치한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "몸은 화면 오른쪽으로 비스듬히 앉아 있고 얼굴은 카메라 쪽으로 돌아와 있다. 두 눈은 화면 왼쪽의 프레임 밖을 향하며, 바라보는 상대는 보이지 않는다. 특정 시선 대상이 명시되지 않은 충격 반응으로는 가능하다. 무기나 방향성 소품은 없다.",
        "built_space": "밝은 상부와 회색 하부로 나뉜 벽 두 면, 모서리 하나, 걸레받이와 회색 바닥이 보인다. 여성 뒤에는 갈색 목재 등판과 검은 금속 틀의 의자 하나가 있어 이전 장면의 의자와 잘 맞는다. 여성의 등 뒤에 등받이가 놓인 배치도 자연스럽다. 문과 천장 조명은 프레임 밖이므로 개폐 상태나 개수를 확인할 수 없다. 반사는 없다.",
        "entities": "중년 한국인 여성으로 보이는 인물 한 명만 있으며 검은 단발과 얼굴 윤곽, 갈색 계열 체크 셔츠가 인물 참조와 대체로 맞는다. 참조의 모자는 없고, 앞치마가 보여야 할 상체에는 셔츠만 있다. 양 볼과 눈 밑의 붉은 멍은 보이지만 심하게 부어오른 얼굴이라는 조건에 비하면 부기가 약하다. 눈은 정상적인 해부학적 형태로 커져 있고 입이 살짝 벌어져 있다. 금지된 사무실 소품이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "몸통은 의자에 앉은 자세로 수직에 가깝게 유지되며 등받이가 몸 뒤에 보인다. 실제 좌판과 골반 접촉은 하단 크롭에 가려져 있지만, 의자와 몸의 관계에 부유나 불가능한 지지는 없다. 팔은 몸 옆으로 내려가며 손은 프레임 밖이다. 공중에 뜬 물체나 운동 중인 신체는 없다."
       },
       {
        "label": "A",
        "direction": "몸과 얼굴은 거의 정면이고 두 눈은 렌즈보다 화면 왼쪽의 가까운 프레임 밖 지점을 향한다. 대상 인물은 보이지 않는다. 눈꺼풀이 크게 열리고 입이 벌어진 모습은 말을 듣고 놀란 순간으로 읽힌다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "두 색으로 칠한 벽 두 면과 모서리 하나, 걸레받이, 회색 바닥이 이어진다. 오른쪽 끝에는 열린 출입구 일부가 보이고 벽에는 스위치판 하나가 있다. 등 뒤의 목재 등받이는 하나로 이어져 보이나, 넓고 둥근 밝은 회색 금속 틀은 참조의 좁고 검은 의자 틀과 다르다. 인물은 등받이 앞에 자연스럽게 자리한다. 거울이나 반사는 없다.",
        "entities": "검은 단발의 중년 한국인 여성 한 명으로, 얼굴과 체크 셔츠가 참조에 대체로 부합한다. 갈색 앞치마와 가슴 주머니의 연필 및 끈도 유지된다. 모자는 없다. 눈 밑과 볼에 붉은 변색 및 약한 부기가 있지만 심한 구타 후의 부종은 충분히 드러나지 않는다. 크게 뜬 눈의 홍채와 동공은 정상이며 추가 인물, 금지된 사무실 소품, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "상체는 등받이 앞에서 곧게 유지되고 어깨와 목의 연결도 자연스럽다. 좌판과 골반은 프레임 밖이지만 앉은 자세와 의자 위치가 모순되지 않는다. 앞치마는 어깨끈으로 지지되고 연필과 끈은 주머니에 걸쳐 있다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.833
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.833
   },
   "violations": {
    "A": [
     "[gemini-pro] 레퍼런스와 다른 형태의 의자(둥근 금속 프레임) 배치로 장소 고정(LOCATION lock) 위반",
     "[gemini-pro] 피사체를 정중앙에 배치하여 '화면 중앙 좌측' 구도 지시 위반"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1833
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "지정된 '화면 중앙 좌측(middle-left)' 구도를 무시하고 피사체를 중앙에 배치했으며, 얼굴의 심한 붓기가 묘사되지 않았고 레퍼런스와 다른 형태의 의자가 등장함.  ★위반: [gemini-pro] 레퍼런스와 다른 형태의 의자(둥근 금속 프레임) 배치로 장소 고정(LOCATION lock) 위반 / [gemini-pro] 피사체를 정중앙에 배치하여 '화면 중앙 좌측' 구도 지시 위반"
   },
   {
    "label": "B",
    "score": 1833,
    "verdict_ko": "지정된 '화면 중앙 좌측' 구도를 정확히 준수하고 심하게 부어오른 얼굴과 충격받은 표정을 훌륭하게 구현했으나 앞치마가 누락됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh8_sel.png",
    "asset_id": "06d8f17f-86e1-4963-9c29-07aea9507d98",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-3e77-7f7b-b4c7-a166884f87bc",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S31sh8"
  }
 },
 "S31sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:25:46.718255+00:00",
  "fingerprint": "a9f90395e970a9930fc78c9b81c55dde99e720ed077f721c51201a7140d8152f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S31sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S31sh9_sel.png",
  "source_sha256": "a20032972c86dc6f96e0a5f253f50a57b36b1fefd82e9cd1f94a63bbe284917e",
  "file": "S31sh9_cine.png",
  "staged_sha256": "2cdcadc066d15e2dccae8c8ff0cc3d9d017c4faa48129f799974586ba23334b9",
  "latency_ms": 9926
 },
 "S32sh5::signage": {
  "fp": "920671277827e0c3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S32sh5": {
  "input_fingerprint": "29d68fa4df68ed78",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환을 향해 고개를 한쪽으로 강하게 꺾은 채 굳은 표정으로 거절하는 현우의 단호한 얼굴.\n\nLOCATION (lock): Near the table and wall inside the church basement prayer room, in the lantern light established earlier. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Basement wall (Behind 현우, supporting his lean); used as A narrow visible area anchors his resistant posture without introducing additional detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the basement, with controlled contrast keeping his firm expression readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's walls and sparse fixed furnishings, including the cross and small organ. Exclude the upstairs entrance and exterior surveillance lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit, with the small table, chairs, organ and cross unchanged. Charlie now holds the functioning radio and still wears his old coat and hat. 현우: He stands against the wall with folded arms, facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환을 향해 고개를 한쪽으로 강하게 꺾은 채 굳은 표정으로 거절하는 현우의 단호한 얼굴.\n\nLOCATION (lock): Near the table and wall inside the church basement prayer room, in the lantern light established earlier. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Basement wall (Behind 현우, supporting his lean); used as A narrow visible area anchors his resistant posture without introducing additional detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the basement, with controlled contrast keeping his firm expression readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's walls and sparse fixed furnishings, including the cross and small organ. Exclude the upstairs entrance and exterior surveillance lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit, with the small table, chairs, organ and cross unchanged. Charlie now holds the functioning radio and still wears his old coat and hat. 현우: He stands against the wall with folded arms, facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환을 향해 고개를 한쪽으로 강하게 꺾은 채 굳은 표정으로 거절하는 현우의 단호한 얼굴.\n\nLOCATION (lock): Near the table and wall inside the church basement prayer room, in the lantern light established earlier. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Basement wall (Behind 현우, supporting his lean); used as A narrow visible area anchors his resistant posture without introducing additional detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the basement, with controlled contrast keeping his firm expression readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's walls and sparse fixed furnishings, including the cross and small organ. Exclude the upstairs entrance and exterior surveillance lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit, with the small table, chairs, organ and cross unchanged. Charlie now holds the functioning radio and still wears his old coat and hat. 현우: He stands against the wall with folded arms, facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우가 화면 좌측 전경에 있는 인물을 향해 시선을 두고 있음.",
    "built_space": "지하실 벽면에 현우가 기대어 있으며, 화면에 2개의 랜턴이 보임.",
    "entities": "현우(회색 셔츠, 팔짱, 얼굴 상처)의 모습은 일치하나, 목록에 없는 인물(전경의 남성)이 포함됨.",
    "hard_violations": [
     "[gemini-pro] 허용되지 않은 인물 추가 (PEOPLE 목록에 없는 전경의 남성)",
     "[gpt-high] 현우만 등장하도록 제한했는데 왼쪽 전경에 별도의 성인 남성을 추가했다.",
     "[gpt-high] 벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
    ],
    "physics": "현우가 벽에 기대어 무릎을 세우고 앉아 있음."
   },
   {
    "label": "B",
    "direction": "현우의 시선과 고개가 화면 좌측 밖을 향해 꺾여 있음.",
    "built_space": "지하실 벽면에 현우가 기대어 있고, 배경에 오르간과 랜턴이 배치됨.",
    "entities": "현우(회색 셔츠, 팔짱, 얼굴 상처)가 단독으로 등장하며 조건과 일치함.",
    "hard_violations": [
     "[gpt-high] 벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
    ],
    "physics": "현우가 벽에 기대어 무릎을 세운 채 앉아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "PEOPLE 제약을 어기고 전경에 허용되지 않은 인물을 추가한 결정적인 위반이 있으며, 현우가 서 있는 대신 앉아 있습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "허용된 인물만 화면에 담아내고 배경과 단호한 표정을 잘 살렸으나, 요구된 선 자세가 아닌 앉은 자세로 묘사된 점이 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우가 화면 좌측 전경에 있는 인물을 향해 시선을 두고 있음.",
        "built_space": "지하실 벽면에 현우가 기대어 있으며, 화면에 2개의 랜턴이 보임.",
        "entities": "현우(회색 셔츠, 팔짱, 얼굴 상처)의 모습은 일치하나, 목록에 없는 인물(전경의 남성)이 포함됨.",
        "hard_violations": [
         "허용되지 않은 인물 추가 (PEOPLE 목록에 없는 전경의 남성)"
        ],
        "physics": "현우가 벽에 기대어 무릎을 세우고 앉아 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선과 고개가 화면 좌측 밖을 향해 꺾여 있음.",
        "built_space": "지하실 벽면에 현우가 기대어 있고, 배경에 오르간과 랜턴이 배치됨.",
        "entities": "현우(회색 셔츠, 팔짱, 얼굴 상처)가 단독으로 등장하며 조건과 일치함.",
        "hard_violations": [],
        "physics": "현우가 벽에 기대어 무릎을 세운 채 앉아 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "PEOPLE 제약을 어기고 전경에 허용되지 않은 인물을 추가한 결정적인 위반이 있으며, 현우가 서 있는 대신 앉아 있습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "허용된 인물만 화면에 담아내고 배경과 단호한 표정을 잘 살렸으나, 요구된 선 자세가 아닌 앉은 자세로 묘사된 점이 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우가 화면 좌측 전경에 있는 인물을 향해 시선을 두고 있음.",
        "built_space": "지하실 벽면에 현우가 기대어 있으며, 화면에 2개의 랜턴이 보임.",
        "entities": "현우(회색 셔츠, 팔짱, 얼굴 상처)의 모습은 일치하나, 목록에 없는 인물(전경의 남성)이 포함됨.",
        "hard_violations": [
         "허용되지 않은 인물 추가 (PEOPLE 목록에 없는 전경의 남성)"
        ],
        "physics": "현우가 벽에 기대어 무릎을 세우고 앉아 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선과 고개가 화면 좌측 밖을 향해 꺾여 있음.",
        "built_space": "지하실 벽면에 현우가 기대어 있고, 배경에 오르간과 랜턴이 배치됨.",
        "entities": "현우(회색 셔츠, 팔짱, 얼굴 상처)가 단독으로 등장하며 조건과 일치함.",
        "hard_violations": [],
        "physics": "현우가 벽에 기대어 무릎을 세운 채 앉아 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "불필요한 인물 없이 현우의 외형과 공간 재질은 살렸지만, 서 있는 얼굴 클로즈업 대신 무릎까지 보이는 앉은 자세이며 강하게 고개를 꺾는 거절 동작도 부족하다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "현우의 단호한 시선은 보이나, 허용되지 않은 상대 남성을 크게 추가했고 서 있는 얼굴 클로즈업을 앉은 자세의 대면 구도로 바꿨다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽 바깥을 향한다. 구도환을 화면 밖 상대라고 해석할 수 있지만 실제 시선 대상은 보이지 않는다. 고개는 거의 세운 채 살짝 돌렸으며, 한쪽으로 강하게 꺾은 순간은 드러나지 않는다.",
        "built_space": "오른쪽 거친 벽에 현우의 뒤통수와 등·어깨가 닿는다. 왼쪽에는 목재 건반 악기 한 대, 그 앞 의자 일부 하나, 악기 위 켜진 랜턴 하나와 액자 일부가 보이고 중앙 뒤에는 푸른 커튼이 있다. 이전 장소의 재질과 요소는 유사하지만 배경이 넓게 노출되어 좁은 벽 영역만 남기는 얼굴 클로즈업과 다르다. 올라온 무릎과 낮게 기대는 몸통은 서 있기보다 앉아 있는 배치로 읽힌다.",
        "entities": "현우로 보이는 앳된 동아시아계 남성 한 명만 있다. 헝클어진 검은 머리, 회색 셔츠, 마른 체격은 인물 참조와 대체로 맞으며 얼굴의 멍과 찰과상, 겉옷을 벗은 상태, 팔짱이 보인다. 한국계 미국인이라는 국적은 외형만으로 확인할 수 없다. 다리 부상과 신발 속 카드는 이 구도에서 확인되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
        ],
        "physics": "등과 머리는 벽에 기대고, 접힌 팔과 손은 서로의 팔에 자연스럽게 얹혀 있다. 무릎은 굽힌 다리에 연결되어 있으며 비행하거나 떠 있는 신체는 보이지 않는다. 엉덩이와 바닥 접점은 화면 밖이라 직접 확인되지 않는다. 랜턴은 악기 상판 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽 전경의 남성 얼굴을 올려다보며 그 남성도 현우를 향한다. 거절의 상대를 향한 시선 관계는 명확하지만, 고개는 거의 수직이고 강하게 한쪽으로 꺾이지 않았다.",
        "built_space": "현우는 오른쪽 벽에 머리와 등·어깨를 붙이고 무릎을 세워 앉아 보인다. 왼쪽 전경의 남성이 화면 상당 부분을 차지한다. 랜턴은 위쪽 중앙에 하나, 오른쪽 가장자리에 일부 잘린 하나로 총 두 개가 보인다. 거친 벽의 재질은 장소 참조와 유사하지만, 좁은 벽만 배경으로 둔 얼굴 클로즈업이 아니라 상대를 포함한 넓은 대면 구도다. 악기·십자가·탁자는 이 화면에 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 검은 헝클어진 머리와 회색 셔츠, 얼굴 상처, 겉옷 없는 상태와 팔짱이 참조 및 상태 지시와 대체로 맞는다. 그러나 왼쪽에 짧은 검은 머리와 어두운 옷의 성인 남성이 추가되어 현우만 허용한 인물 제한을 어긴다. 다리 부상과 숨긴 카드는 확인할 수 없고 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우만 등장하도록 제한했는데 왼쪽 전경에 별도의 성인 남성을 추가했다.",
         "벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
        ],
        "physics": "현우의 머리와 등은 벽에 접촉하고 팔짱 낀 손은 반대쪽 팔에 놓여 있다. 굽힌 무릎과 몸통의 연결은 자연스러우며 엉덩이와 발의 지지점은 화면 밖이다. 상대 남성도 하체가 잘렸을 뿐 공중에 떠 있는 모습은 아니다. 랜턴의 상부 고정점은 프레임에 잘려 있어 부착 방식은 확인되지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "불필요한 인물 없이 현우의 외형과 공간 재질은 살렸지만, 서 있는 얼굴 클로즈업 대신 무릎까지 보이는 앉은 자세이며 강하게 고개를 꺾는 거절 동작도 부족하다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "현우의 단호한 시선은 보이나, 허용되지 않은 상대 남성을 크게 추가했고 서 있는 얼굴 클로즈업을 앉은 자세의 대면 구도로 바꿨다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽 바깥을 향한다. 구도환을 화면 밖 상대라고 해석할 수 있지만 실제 시선 대상은 보이지 않는다. 고개는 거의 세운 채 살짝 돌렸으며, 한쪽으로 강하게 꺾은 순간은 드러나지 않는다.",
        "built_space": "오른쪽 거친 벽에 현우의 뒤통수와 등·어깨가 닿는다. 왼쪽에는 목재 건반 악기 한 대, 그 앞 의자 일부 하나, 악기 위 켜진 랜턴 하나와 액자 일부가 보이고 중앙 뒤에는 푸른 커튼이 있다. 이전 장소의 재질과 요소는 유사하지만 배경이 넓게 노출되어 좁은 벽 영역만 남기는 얼굴 클로즈업과 다르다. 올라온 무릎과 낮게 기대는 몸통은 서 있기보다 앉아 있는 배치로 읽힌다.",
        "entities": "현우로 보이는 앳된 동아시아계 남성 한 명만 있다. 헝클어진 검은 머리, 회색 셔츠, 마른 체격은 인물 참조와 대체로 맞으며 얼굴의 멍과 찰과상, 겉옷을 벗은 상태, 팔짱이 보인다. 한국계 미국인이라는 국적은 외형만으로 확인할 수 없다. 다리 부상과 신발 속 카드는 이 구도에서 확인되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
        ],
        "physics": "등과 머리는 벽에 기대고, 접힌 팔과 손은 서로의 팔에 자연스럽게 얹혀 있다. 무릎은 굽힌 다리에 연결되어 있으며 비행하거나 떠 있는 신체는 보이지 않는다. 엉덩이와 바닥 접점은 화면 밖이라 직접 확인되지 않는다. 랜턴은 악기 상판 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽 전경의 남성 얼굴을 올려다보며 그 남성도 현우를 향한다. 거절의 상대를 향한 시선 관계는 명확하지만, 고개는 거의 수직이고 강하게 한쪽으로 꺾이지 않았다.",
        "built_space": "현우는 오른쪽 벽에 머리와 등·어깨를 붙이고 무릎을 세워 앉아 보인다. 왼쪽 전경의 남성이 화면 상당 부분을 차지한다. 랜턴은 위쪽 중앙에 하나, 오른쪽 가장자리에 일부 잘린 하나로 총 두 개가 보인다. 거친 벽의 재질은 장소 참조와 유사하지만, 좁은 벽만 배경으로 둔 얼굴 클로즈업이 아니라 상대를 포함한 넓은 대면 구도다. 악기·십자가·탁자는 이 화면에 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 검은 헝클어진 머리와 회색 셔츠, 얼굴 상처, 겉옷 없는 상태와 팔짱이 참조 및 상태 지시와 대체로 맞는다. 그러나 왼쪽에 짧은 검은 머리와 어두운 옷의 성인 남성이 추가되어 현우만 허용한 인물 제한을 어긴다. 다리 부상과 숨긴 카드는 확인할 수 없고 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우만 등장하도록 제한했는데 왼쪽 전경에 별도의 성인 남성을 추가했다.",
         "벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
        ],
        "physics": "현우의 머리와 등은 벽에 접촉하고 팔짱 낀 손은 반대쪽 팔에 놓여 있다. 굽힌 무릎과 몸통의 연결은 자연스러우며 엉덩이와 발의 지지점은 화면 밖이다. 상대 남성도 하체가 잘렸을 뿐 공중에 떠 있는 모습은 아니다. 랜턴의 상부 고정점은 프레임에 잘려 있어 부착 방식은 확인되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.833,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.583,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 허용되지 않은 인물 추가 (PEOPLE 목록에 없는 전경의 남성)",
     "[gpt-high] 현우만 등장하도록 제한했는데 왼쪽 전경에 별도의 성인 남성을 추가했다.",
     "[gpt-high] 벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
    ],
    "B": [
     "[gpt-high] 벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 583,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 583,
    "verdict_ko": "PEOPLE 제약을 어기고 전경에 허용되지 않은 인물을 추가한 결정적인 위반이 있으며, 현우가 서 있는 대신 앉아 있습니다.  ★위반: [gemini-pro] 허용되지 않은 인물 추가 (PEOPLE 목록에 없는 전경의 남성) / [gpt-high] 현우만 등장하도록 제한했는데 왼쪽 전경에 별도의 성인 남성을 추가했다. / [gpt-high] 벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "허용된 인물만 화면에 담아내고 배경과 단호한 표정을 잘 살렸으나, 요구된 선 자세가 아닌 앉은 자세로 묘사된 점이 아쉽습니다.  ★위반: [gpt-high] 벽에 서 있어야 하는 현우를 무릎을 세우고 벽에 기대 앉은 자세로 배치했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S30sh17_sel.png",
    "asset_id": "df30a9f4-6336-4813-9f6b-f8c22a412e5c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-401d-7dd2-a1b8-bad76e073ecb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S30sh17"
  }
 },
 "S32sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:26:48.951928+00:00",
  "fingerprint": "1930a6e15dbc08c9a4f309e8b4fd4258622d719671d1be6600d1d8fcdfd1c91c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S32sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S32sh5_sel.png",
  "source_sha256": "def9e20d5a38f53c03e3276ed38f0a72a4c2f3f23c19b201ab4096ed5f2d3693",
  "file": "S32sh5_cine.png",
  "staged_sha256": "6c02202013adfe6350a6e36bbe55e858a7d10f825e51d6954f01de1f0f6890e0",
  "latency_ms": 10924
 },
 "S32sh8::signage": {
  "fp": "9854909145f5a1d7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S32sh8": {
  "input_fingerprint": "1ec9ffd79031c25b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낡은 손전등을 꽉 쥔 채 지하실 구석의 비밀 통로 쪽을 검지손가락으로 가리키는 신부의 역동적인 상체.\n\nLOCATION (lock): At the concealed passage entrance in a corner of the church basement prayer room, lit by lanterns and a carried flashlight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 신부 in the middle-left of the frame, midground, points to secret passage approach; secret passage approach in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Secret passage (The escape route indicated by 신부) — The approach to the passage is seen obliquely beyond his pointing hand; its continuation lies outside the frame; used as Provides a spatial destination for the gesture without specifying an unsupported door mechanism; Old flashlight (Gripped tightly by 신부) — Seen side-on near his lower torso; used as A small practical object supporting the urgency of departure, kept subordinate to the pointing hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral basement ambient illumination and readable hand contours without assuming that the flashlight has been switched on.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement furnishings and lit lanterns remain unchanged, with a secret exit passage available. Charlie retains the radio, old coat and hat. 신부: He has returned with the warning and carries a flashlight as he leads toward the secret passage, still wearing his clerical collar.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 신부 right now, so 신부's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 신부: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낡은 손전등을 꽉 쥔 채 지하실 구석의 비밀 통로 쪽을 검지손가락으로 가리키는 신부의 역동적인 상체.\n\nLOCATION (lock): At the concealed passage entrance in a corner of the church basement prayer room, lit by lanterns and a carried flashlight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 신부 in the middle-left of the frame, midground, points to secret passage approach; secret passage approach in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Secret passage (The escape route indicated by 신부) — The approach to the passage is seen obliquely beyond his pointing hand; its continuation lies outside the frame; used as Provides a spatial destination for the gesture without specifying an unsupported door mechanism; Old flashlight (Gripped tightly by 신부) — Seen side-on near his lower torso; used as A small practical object supporting the urgency of departure, kept subordinate to the pointing hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral basement ambient illumination and readable hand contours without assuming that the flashlight has been switched on.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement furnishings and lit lanterns remain unchanged, with a secret exit passage available. Charlie retains the radio, old coat and hat. 신부: He has returned with the warning and carries a flashlight as he leads toward the secret passage, still wearing his clerical collar.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 신부 right now, so 신부's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 신부: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낡은 손전등을 꽉 쥔 채 지하실 구석의 비밀 통로 쪽을 검지손가락으로 가리키는 신부의 역동적인 상체.\n\nLOCATION (lock): At the concealed passage entrance in a corner of the church basement prayer room, lit by lanterns and a carried flashlight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 신부 in the middle-left of the frame, midground, points to secret passage approach; secret passage approach in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Secret passage (The escape route indicated by 신부) — The approach to the passage is seen obliquely beyond his pointing hand; its continuation lies outside the frame; used as Provides a spatial destination for the gesture without specifying an unsupported door mechanism; Old flashlight (Gripped tightly by 신부) — Seen side-on near his lower torso; used as A small practical object supporting the urgency of departure, kept subordinate to the pointing hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral basement ambient illumination and readable hand contours without assuming that the flashlight has been switched on.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement furnishings and lit lanterns remain unchanged, with a secret exit passage available. Charlie retains the radio, old coat and hat. 신부: He has returned with the warning and carries a flashlight as he leads toward the secret passage, still wearing his clerical collar.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 신부 right now, so 신부's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 신부: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "사제가 오른쪽 통로를 가리키며 시선도 같은 곳을 향함.",
    "built_space": "왼쪽 피아노와 그 위의 랜턴 배치가 이전 샷과 일치함. 오른쪽에 커튼이 있는 통로 위치.",
    "entities": "사제의 얼굴, 복장, 십자가, 손전등 모두 프롬프트 및 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "오른손으로 손전등을 단단히 쥐고, 왼팔을 뻗어 가리키는 자세가 해부학적으로 올바르게 지지됨."
   },
   {
    "label": "B",
    "direction": "사제가 오른쪽 통로를 가리키며 시선도 그곳을 향함.",
    "built_space": "피아노의 배치 방향이 다르며, 이전 샷의 랜턴이 벽과 허공에 두 개로 복제되어 배치됨.",
    "entities": "사제의 외모와 복장은 레퍼런스와 잘 일치하며 손전등을 들고 있음.",
    "hard_violations": [
     "[gemini-pro] 이전 샷의 단일 랜턴 픽스처를 2개로 복제하여 임의의 위치에 배치함 (구조물 일관성 위반)"
    ],
    "physics": "오른손으로 손전등을 잡고 왼손으로 지시하는 자세가 바닥에 서서 유지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물 묘사와 프레이밍이 우수하며, 이전 샷의 피아노와 랜턴 위치 등 공간적 연속성을 완벽히 유지했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물은 잘 표현되었으나, 지정된 단일 랜턴을 복제하여 잘못된 위치에 배치한 점이 구조물 위반입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사제가 오른쪽 통로를 가리키며 시선도 같은 곳을 향함.",
        "built_space": "왼쪽 피아노와 그 위의 랜턴 배치가 이전 샷과 일치함. 오른쪽에 커튼이 있는 통로 위치.",
        "entities": "사제의 얼굴, 복장, 십자가, 손전등 모두 프롬프트 및 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "오른손으로 손전등을 단단히 쥐고, 왼팔을 뻗어 가리키는 자세가 해부학적으로 올바르게 지지됨."
       },
       {
        "label": "B",
        "direction": "사제가 오른쪽 통로를 가리키며 시선도 그곳을 향함.",
        "built_space": "피아노의 배치 방향이 다르며, 이전 샷의 랜턴이 벽과 허공에 두 개로 복제되어 배치됨.",
        "entities": "사제의 외모와 복장은 레퍼런스와 잘 일치하며 손전등을 들고 있음.",
        "hard_violations": [
         "이전 샷의 단일 랜턴 픽스처를 2개로 복제하여 임의의 위치에 배치함 (구조물 일관성 위반)"
        ],
        "physics": "오른손으로 손전등을 잡고 왼손으로 지시하는 자세가 바닥에 서서 유지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물 묘사와 프레이밍이 우수하며, 이전 샷의 피아노와 랜턴 위치 등 공간적 연속성을 완벽히 유지했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물은 잘 표현되었으나, 지정된 단일 랜턴을 복제하여 잘못된 위치에 배치한 점이 구조물 위반입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "사제가 오른쪽 통로를 가리키며 시선도 같은 곳을 향함.",
        "built_space": "왼쪽 피아노와 그 위의 랜턴 배치가 이전 샷과 일치함. 오른쪽에 커튼이 있는 통로 위치.",
        "entities": "사제의 얼굴, 복장, 십자가, 손전등 모두 프롬프트 및 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "오른손으로 손전등을 단단히 쥐고, 왼팔을 뻗어 가리키는 자세가 해부학적으로 올바르게 지지됨."
       },
       {
        "label": "B",
        "direction": "사제가 오른쪽 통로를 가리키며 시선도 그곳을 향함.",
        "built_space": "피아노의 배치 방향이 다르며, 이전 샷의 랜턴이 벽과 허공에 두 개로 복제되어 배치됨.",
        "entities": "사제의 외모와 복장은 레퍼런스와 잘 일치하며 손전등을 들고 있음.",
        "hard_violations": [
         "이전 샷의 단일 랜턴 픽스처를 2개로 복제하여 임의의 위치에 배치함 (구조물 일관성 위반)"
        ],
        "physics": "오른손으로 손전등을 잡고 왼손으로 지시하는 자세가 바닥에 서서 유지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "중앙 왼쪽 신부의 역동적인 상체, 통로를 향한 검지와 시선, 허리 가까이 쥔 낡은 손전등이 핵심 지시를 충족하지만 등불 배치와 조명색의 연속성은 덜 정확하다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "미디엄 숏과 통로를 가리키는 손은 맞지만 시선이 반대편을 향하고, 새로 드러난 묶인 커튼과 켜진 손전등이 장소·조명 지시의 충실도를 낮춘다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 검지는 화면 오른쪽 뒤의 어두운 통로 입구를 향하고, 눈과 얼굴도 같은 방향을 향한다. 지시 대상이 손 너머에 보인다. 하복부 가까이 쥔 손전등의 렌즈도 오른쪽을 향하지만 뚜렷한 투사광은 보이지 않는다.",
        "built_space": "중앙 왼쪽에 허리 부근까지 보이는 신부가 있고, 오른쪽 배경에는 두꺼운 벽체로 둘러싸인 입구 하나와 안쪽 단차가 비스듬히 보인다. 왼쪽 뒤에는 피아노 한 대가 일부 가려져 있다. 등불은 왼쪽 위, 피아노 위쪽, 통로 안, 오른쪽 가장자리에 총 네 개 보이며 피아노 쪽 것은 어둡다. 낡은 회벽과 목재 피아노는 장소 참조와 통하지만, 참조에서 피아노 위에 켜져 있던 등불 및 비교적 중립적인 벽색과의 연속성은 불완전하다. 반사상은 없다.",
        "entities": "인물은 신부 한 명뿐이며, 주름진 얼굴과 짧은 회흑색 머리의 한국인 60대 남성이라는 설정 및 인물 참조에 대체로 부합한다. 흰 성직자 칼라, 때 묻은 검은 셔츠, 십자가 목걸이가 보인다. 손에 든 물건은 마모된 금속제 손전등 한 개로, 몸통 아래에서 측면으로 보인다. 이전 장면의 청년이나 다른 인물은 없고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "손전등은 손가락으로 몸통을 감싸 쥐어 지지하고 있다. 반대쪽 팔은 어깨에서 자연스럽게 뻗으며 검지만 펴져 있어 실제 지시 동작으로 가능하다. 몸통의 전방 기울기도 무리하지 않다. 발은 프레임 밖이므로 접지를 직접 확인할 수 없지만 공중에 떠 있는 자세는 아니다. 목걸이는 목에 걸려 있고, 등불은 가구에 놓이거나 상단 걸이로 매달린 형태다."
       },
       {
        "label": "B",
        "direction": "뻗은 검지는 오른쪽 뒤 커튼 옆의 통로 접근부를 향한다. 그러나 얼굴과 눈은 화면 왼쪽의 프레임 밖을 향하여 지시 대상과 시선이 갈린다. 손전등은 오른쪽으로 향하고 희미한 빛을 투사하는 것으로 보인다.",
        "built_space": "신부는 중앙 왼쪽에 허리 부근까지 보이며, 오른쪽 뒤에는 계단형 단차가 있는 좁은 통로 입구 하나가 보인다. 입구 오른쪽에는 묶어 젖힌 큰 커튼이 있다. 왼쪽에는 목재 피아노 한 대와 그 앞 의자 일부가 보인다. 등불은 피아노 왼쪽 끝의 어두운 것, 피아노 위의 켜진 것, 오른쪽 위에 걸린 것까지 세 개다. 피아노와 그 위 등불은 이전 장소의 인상을 잘 이어가지만, 묶인 커튼은 참조에서 확인되지 않는 공간 세부다. 반사상은 없다.",
        "entities": "신부 한 명만 등장하며, 연령대와 한국인 남성 외형, 회흑색 머리 및 얼굴 특징은 참조에 대체로 맞는다. 흰 성직자 칼라와 검은 셔츠, 십자가 목걸이가 있으나 셔츠의 낡고 오염된 정도는 참조보다 약하다. 하복부 앞에 금속제 손전등 한 개를 쥐고 있으며 손에는 반지도 보인다. 피아노 위 종이는 있으나 글자는 판독되지 않고, 다른 인물이나 자막은 없다.",
        "hard_violations": [],
        "physics": "손전등은 손바닥과 굽힌 손가락으로 확실히 지지된다. 반대 팔을 뒤쪽 통로로 뻗고 고개를 왼쪽으로 돌리는 자세는 물리적으로 가능하지만, 시선까지 통로로 이끄는 동작은 아니다. 하체는 잘려 있어 발의 접촉은 보이지 않는다. 커튼은 위에서 내려와 옆으로 묶여 있고, 피아노 위 물건들은 상판에 놓여 있어 지지가 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "중앙 왼쪽 신부의 역동적인 상체, 통로를 향한 검지와 시선, 허리 가까이 쥔 낡은 손전등이 핵심 지시를 충족하지만 등불 배치와 조명색의 연속성은 덜 정확하다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "미디엄 숏과 통로를 가리키는 손은 맞지만 시선이 반대편을 향하고, 새로 드러난 묶인 커튼과 켜진 손전등이 장소·조명 지시의 충실도를 낮춘다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 검지는 화면 오른쪽 뒤의 어두운 통로 입구를 향하고, 눈과 얼굴도 같은 방향을 향한다. 지시 대상이 손 너머에 보인다. 하복부 가까이 쥔 손전등의 렌즈도 오른쪽을 향하지만 뚜렷한 투사광은 보이지 않는다.",
        "built_space": "중앙 왼쪽에 허리 부근까지 보이는 신부가 있고, 오른쪽 배경에는 두꺼운 벽체로 둘러싸인 입구 하나와 안쪽 단차가 비스듬히 보인다. 왼쪽 뒤에는 피아노 한 대가 일부 가려져 있다. 등불은 왼쪽 위, 피아노 위쪽, 통로 안, 오른쪽 가장자리에 총 네 개 보이며 피아노 쪽 것은 어둡다. 낡은 회벽과 목재 피아노는 장소 참조와 통하지만, 참조에서 피아노 위에 켜져 있던 등불 및 비교적 중립적인 벽색과의 연속성은 불완전하다. 반사상은 없다.",
        "entities": "인물은 신부 한 명뿐이며, 주름진 얼굴과 짧은 회흑색 머리의 한국인 60대 남성이라는 설정 및 인물 참조에 대체로 부합한다. 흰 성직자 칼라, 때 묻은 검은 셔츠, 십자가 목걸이가 보인다. 손에 든 물건은 마모된 금속제 손전등 한 개로, 몸통 아래에서 측면으로 보인다. 이전 장면의 청년이나 다른 인물은 없고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "손전등은 손가락으로 몸통을 감싸 쥐어 지지하고 있다. 반대쪽 팔은 어깨에서 자연스럽게 뻗으며 검지만 펴져 있어 실제 지시 동작으로 가능하다. 몸통의 전방 기울기도 무리하지 않다. 발은 프레임 밖이므로 접지를 직접 확인할 수 없지만 공중에 떠 있는 자세는 아니다. 목걸이는 목에 걸려 있고, 등불은 가구에 놓이거나 상단 걸이로 매달린 형태다."
       },
       {
        "label": "A",
        "direction": "뻗은 검지는 오른쪽 뒤 커튼 옆의 통로 접근부를 향한다. 그러나 얼굴과 눈은 화면 왼쪽의 프레임 밖을 향하여 지시 대상과 시선이 갈린다. 손전등은 오른쪽으로 향하고 희미한 빛을 투사하는 것으로 보인다.",
        "built_space": "신부는 중앙 왼쪽에 허리 부근까지 보이며, 오른쪽 뒤에는 계단형 단차가 있는 좁은 통로 입구 하나가 보인다. 입구 오른쪽에는 묶어 젖힌 큰 커튼이 있다. 왼쪽에는 목재 피아노 한 대와 그 앞 의자 일부가 보인다. 등불은 피아노 왼쪽 끝의 어두운 것, 피아노 위의 켜진 것, 오른쪽 위에 걸린 것까지 세 개다. 피아노와 그 위 등불은 이전 장소의 인상을 잘 이어가지만, 묶인 커튼은 참조에서 확인되지 않는 공간 세부다. 반사상은 없다.",
        "entities": "신부 한 명만 등장하며, 연령대와 한국인 남성 외형, 회흑색 머리 및 얼굴 특징은 참조에 대체로 맞는다. 흰 성직자 칼라와 검은 셔츠, 십자가 목걸이가 있으나 셔츠의 낡고 오염된 정도는 참조보다 약하다. 하복부 앞에 금속제 손전등 한 개를 쥐고 있으며 손에는 반지도 보인다. 피아노 위 종이는 있으나 글자는 판독되지 않고, 다른 인물이나 자막은 없다.",
        "hard_violations": [],
        "physics": "손전등은 손바닥과 굽힌 손가락으로 확실히 지지된다. 반대 팔을 뒤쪽 통로로 뻗고 고개를 왼쪽으로 돌리는 자세는 물리적으로 가능하지만, 시선까지 통로로 이끄는 동작은 아니다. 하체는 잘려 있어 발의 접촉은 보이지 않는다. 커튼은 위에서 내려와 옆으로 묶여 있고, 피아노 위 물건들은 상판에 놓여 있어 지지가 자연스럽다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 이전 샷의 단일 랜턴 픽스처를 2개로 복제하여 임의의 위치에 배치함 (구조물 일관성 위반)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "인물 묘사와 프레이밍이 우수하며, 이전 샷의 피아노와 랜턴 위치 등 공간적 연속성을 완벽히 유지했습니다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "인물은 잘 표현되었으나, 지정된 단일 랜턴을 복제하여 잘못된 위치에 배치한 점이 구조물 위반입니다.  ★위반: [gemini-pro] 이전 샷의 단일 랜턴 픽스처를 2개로 복제하여 임의의 위치에 배치함 (구조물 일관성 위반)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S32sh5_sel.png",
    "asset_id": "e28fe58e-7ac1-4d1c-b4d1-8286f08cedcc",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-41c2-7d6a-938b-ea29b5701739",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S32sh5"
  }
 },
 "S32sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:28:19.124328+00:00",
  "fingerprint": "9363ce11a9c0755abe944b99c00f1563e7a634a528cc5f86cd107d618cf8f559",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S32sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S32sh8_sel.png",
  "source_sha256": "e8f362aa20382a8a32c54df5520e2c16374f7b276ba3f7d12f7c480660c6db4b",
  "file": "S32sh8_cine.png",
  "staged_sha256": "b32a01d4bee57f3d7d4b0bdebce563ca62bc041029073e9ce6051dfb85ba4382",
  "latency_ms": 11559
 },
 "S32sh9::signage": {
  "fp": "9d486fc277c0b7f6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S32sh9": {
  "input_fingerprint": "e8fdcadead5c1f55",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 통로 쪽으로 향하려는 신부의 팔을 다급하게 꽉 붙잡은 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): Just before the hidden passage entrance inside the church basement, within the prayer room's lantern-lit area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding neutral ambient illumination, using controlled contrast to make the gripping fingers and arrested arm clearly legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement corner and the concealed passage entrance visible in the reference. Exclude the upstairs front door and outdoor searchlights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The lit basement and its secret exit passage remain unchanged. Charlie still has the radio and wears the old coat and hat. 현우: He is at the departure point, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 신부: He pauses on his way toward the secret passage, holding the flashlight and wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 통로 쪽으로 향하려는 신부의 팔을 다급하게 꽉 붙잡은 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): Just before the hidden passage entrance inside the church basement, within the prayer room's lantern-lit area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding neutral ambient illumination, using controlled contrast to make the gripping fingers and arrested arm clearly legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement corner and the concealed passage entrance visible in the reference. Exclude the upstairs front door and outdoor searchlights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The lit basement and its secret exit passage remain unchanged. Charlie still has the radio and wears the old coat and hat. 현우: He is at the departure point, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 신부: He pauses on his way toward the secret passage, holding the flashlight and wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 통로 쪽으로 향하려는 신부의 팔을 다급하게 꽉 붙잡은 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): Just before the hidden passage entrance inside the church basement, within the prayer room's lantern-lit area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding neutral ambient illumination, using controlled contrast to make the gripping fingers and arrested arm clearly legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement corner and the concealed passage entrance visible in the reference. Exclude the upstairs front door and outdoor searchlights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The lit basement and its secret exit passage remain unchanged. Charlie still has the radio and wears the old coat and hat. 현우: He is at the departure point, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 신부: He pauses on his way toward the secret passage, holding the flashlight and wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 손이 신부의 팔을 향해 뻗어 꽉 붙잡음.",
    "built_space": "지하실 배경에 피아노와 랜턴이 알맞게 배치됨.",
    "entities": "신부(검은 셔츠, 로만칼라)와 현우(회색 셔츠, 멍든 손) 모두 일치함.",
    "hard_violations": [],
    "physics": "현우의 손이 신부의 팔을 자연스럽게 쥐고 있음."
   },
   {
    "label": "B",
    "direction": "검은 소매를 입은 손이 손전등을 든 맨팔을 붙잡음.",
    "built_space": "지하실 내부로 피아노와 커튼 통로가 보임.",
    "entities": "복장과 소품은 등장하나 신체의 주인이 완전히 뒤섞임.",
    "hard_violations": [
     "[gemini-pro] 신부의 몸통에서 이어진 검은 소매 팔에 현우의 멍든 손이 달려 있고 화면 밖에서 들어온 맨팔이 신부의 노인 손과 손전등을 쥐고 있는 등 해부학적으로 불가능한 신체 결합",
     "[gpt-high] 신부의 소매 위에 얹힌 붙잡는 손이 손목이나 전완으로 이어지지 않고 독립된 신체 조각처럼 끝나 있어, 현우의 팔로 지지되는 정상적인 해부학적 연결이 성립하지 않는다."
    ],
    "physics": "신체 구조와 연결이 엇갈려 물리적으로 성립할 수 없는 자세임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물과 행동 묘사는 정확하나 인서트 클로즈업 프레이밍을 어기고 미디엄 샷으로 넓게 렌더링됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프레이밍은 지시문에 더 가깝지만, 두 인물의 팔과 복장, 피부가 해부학적으로 불가능하게 결합된 치명적 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 손이 신부의 팔을 향해 뻗어 꽉 붙잡음.",
        "built_space": "지하실 배경에 피아노와 랜턴이 알맞게 배치됨.",
        "entities": "신부(검은 셔츠, 로만칼라)와 현우(회색 셔츠, 멍든 손) 모두 일치함.",
        "hard_violations": [],
        "physics": "현우의 손이 신부의 팔을 자연스럽게 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "검은 소매를 입은 손이 손전등을 든 맨팔을 붙잡음.",
        "built_space": "지하실 내부로 피아노와 커튼 통로가 보임.",
        "entities": "복장과 소품은 등장하나 신체의 주인이 완전히 뒤섞임.",
        "hard_violations": [
         "신부의 몸통에서 이어진 검은 소매 팔에 현우의 멍든 손이 달려 있고 화면 밖에서 들어온 맨팔이 신부의 노인 손과 손전등을 쥐고 있는 등 해부학적으로 불가능한 신체 결합"
        ],
        "physics": "신체 구조와 연결이 엇갈려 물리적으로 성립할 수 없는 자세임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물과 행동 묘사는 정확하나 인서트 클로즈업 프레이밍을 어기고 미디엄 샷으로 넓게 렌더링됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프레이밍은 지시문에 더 가깝지만, 두 인물의 팔과 복장, 피부가 해부학적으로 불가능하게 결합된 치명적 오류가 있음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 손이 신부의 팔을 향해 뻗어 꽉 붙잡음.",
        "built_space": "지하실 배경에 피아노와 랜턴이 알맞게 배치됨.",
        "entities": "신부(검은 셔츠, 로만칼라)와 현우(회색 셔츠, 멍든 손) 모두 일치함.",
        "hard_violations": [],
        "physics": "현우의 손이 신부의 팔을 자연스럽게 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "검은 소매를 입은 손이 손전등을 든 맨팔을 붙잡음.",
        "built_space": "지하실 내부로 피아노와 커튼 통로가 보임.",
        "entities": "복장과 소품은 등장하나 신체의 주인이 완전히 뒤섞임.",
        "hard_violations": [
         "신부의 몸통에서 이어진 검은 소매 팔에 현우의 멍든 손이 달려 있고 화면 밖에서 들어온 맨팔이 신부의 노인 손과 손전등을 쥐고 있는 등 해부학적으로 불가능한 신체 결합"
        ],
        "physics": "신체 구조와 연결이 엇갈려 물리적으로 성립할 수 없는 자세임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "붙잡는 손에 손목·팔 연결이 없어 현우의 동작으로 성립하지 않으며, 손 디테일보다 신부의 상체와 손전등을 넓게 보여준다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우가 신부의 팔을 꽉 붙잡는 접촉과 신체 연결은 명확하지만, 인서트보다 넓은 구도이며 현우의 회색 긴팔 셔츠가 반팔 티셔츠로 바뀌었다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 손가락들이 신부의 손목 가까운 검은 소매를 아래로 감싸고 있다. 신부가 쥔 손전등은 화면 오른쪽 벽을 향한다. 눈은 잘려 시선을 확인할 수 없으며, 뒤쪽 커튼 입구로 이동하려던 방향도 분명하지 않다.",
        "built_space": "왼쪽에 피아노 한 대와 의자 한 개, 그 위쪽에 탁상 랜턴 한 개, 오른쪽 위에 벽걸이 랜턴 한 개가 보인다. 뒤에는 묶인 커튼과 어두운 입구 하나가 있어 참조의 지하실 재료와 주요 배치를 대체로 유지한다. 다만 상체와 주변 가구까지 넓게 들어와 손 디테일 인서트의 화면 점유율을 충족하지 못한다.",
        "entities": "신부의 검은 성직자 셔츠, 흰 칼라, 나이 든 손과 황동색 손전등은 참조에 부합한다. 붙잡는 손에는 상처가 있지만 현우의 팔이나 회색 셔츠로 이어지는 부분이 보이지 않아 인물 귀속이 불명확하다. 얼굴 대부분과 하체는 프레임 밖이므로 신원 세부, 다리 부상, 신발 속 카드는 평가할 수 없다. 읽을 수 있는 글씨는 없다.",
        "hard_violations": [
         "신부의 소매 위에 얹힌 붙잡는 손이 손목이나 전완으로 이어지지 않고 독립된 신체 조각처럼 끝나 있어, 현우의 팔로 지지되는 정상적인 해부학적 연결이 성립하지 않는다."
        ],
        "physics": "신부의 손전등은 손가락으로 감싸 쥐어 지지되며 신부의 팔도 상체에 연결된다. 반면 붙잡는 손은 소매와 접촉하기만 할 뿐 손목 뒤로 이어지는 팔이 없고, 힘을 전달할 현우의 신체 연결이 없다. 다른 한 손은 신부의 반대쪽 소매에 이어져 아래로 내려와 있다."
       },
       {
        "label": "B",
        "direction": "오른쪽 현우가 왼쪽으로 팔을 뻗어 신부의 팔꿈치 위 소매를 움켜쥔다. 신부의 턱은 현우 쪽으로 돌아가 있지만 두 사람의 눈은 프레임 밖이다. 비밀 통로는 오른쪽 뒤에 보이며, 신부가 그쪽으로 향하던 순간인지는 몸통과 팔만으로 확실하지 않다.",
        "built_space": "뒤쪽에 피아노 한 대의 일부와 그 위 랜턴 한 개가 보이고, 오른쪽에는 어두운 통로 입구와 커튼 일부가 있다. 낡은 회벽과 가구는 참조 장소와 대체로 일치하며 고정물 중복은 없다. 두 사람의 상체와 얼굴 아래쪽까지 포함해 요구된 손 디테일 인서트보다 넓다.",
        "entities": "검은 셔츠, 흰 성직자 칼라와 목걸이를 착용한 신부, 젊은 남성의 팔과 상처 난 손이 보인다. 보이는 얼굴 일부는 두 인물의 연령대와 대체로 맞지만 정확한 얼굴·머리 일치는 확인할 수 없다. 현우는 참조의 회색 긴팔 단추 셔츠 대신 회색 반팔 티셔츠를 입었다. 손전등과 하체는 프레임 밖이므로 부재로 감점하지 않는다. 추가 인물이나 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "현우의 손은 손목, 전완, 굽힌 팔꿈치, 상완을 거쳐 오른쪽 몸통에 자연스럽게 연결된다. 손가락이 신부의 소매를 강하게 모아 쥐어 천에 당김 주름이 생기며, 신부의 팔도 어깨에서 아래로 이어져 붙잡혀 멈추는 동작이 물리적으로 가능하다. 떠 있는 물체나 분리된 신체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "붙잡는 손에 손목·팔 연결이 없어 현우의 동작으로 성립하지 않으며, 손 디테일보다 신부의 상체와 손전등을 넓게 보여준다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "현우가 신부의 팔을 꽉 붙잡는 접촉과 신체 연결은 명확하지만, 인서트보다 넓은 구도이며 현우의 회색 긴팔 셔츠가 반팔 티셔츠로 바뀌었다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "중앙의 손가락들이 신부의 손목 가까운 검은 소매를 아래로 감싸고 있다. 신부가 쥔 손전등은 화면 오른쪽 벽을 향한다. 눈은 잘려 시선을 확인할 수 없으며, 뒤쪽 커튼 입구로 이동하려던 방향도 분명하지 않다.",
        "built_space": "왼쪽에 피아노 한 대와 의자 한 개, 그 위쪽에 탁상 랜턴 한 개, 오른쪽 위에 벽걸이 랜턴 한 개가 보인다. 뒤에는 묶인 커튼과 어두운 입구 하나가 있어 참조의 지하실 재료와 주요 배치를 대체로 유지한다. 다만 상체와 주변 가구까지 넓게 들어와 손 디테일 인서트의 화면 점유율을 충족하지 못한다.",
        "entities": "신부의 검은 성직자 셔츠, 흰 칼라, 나이 든 손과 황동색 손전등은 참조에 부합한다. 붙잡는 손에는 상처가 있지만 현우의 팔이나 회색 셔츠로 이어지는 부분이 보이지 않아 인물 귀속이 불명확하다. 얼굴 대부분과 하체는 프레임 밖이므로 신원 세부, 다리 부상, 신발 속 카드는 평가할 수 없다. 읽을 수 있는 글씨는 없다.",
        "hard_violations": [
         "신부의 소매 위에 얹힌 붙잡는 손이 손목이나 전완으로 이어지지 않고 독립된 신체 조각처럼 끝나 있어, 현우의 팔로 지지되는 정상적인 해부학적 연결이 성립하지 않는다."
        ],
        "physics": "신부의 손전등은 손가락으로 감싸 쥐어 지지되며 신부의 팔도 상체에 연결된다. 반면 붙잡는 손은 소매와 접촉하기만 할 뿐 손목 뒤로 이어지는 팔이 없고, 힘을 전달할 현우의 신체 연결이 없다. 다른 한 손은 신부의 반대쪽 소매에 이어져 아래로 내려와 있다."
       },
       {
        "label": "A",
        "direction": "오른쪽 현우가 왼쪽으로 팔을 뻗어 신부의 팔꿈치 위 소매를 움켜쥔다. 신부의 턱은 현우 쪽으로 돌아가 있지만 두 사람의 눈은 프레임 밖이다. 비밀 통로는 오른쪽 뒤에 보이며, 신부가 그쪽으로 향하던 순간인지는 몸통과 팔만으로 확실하지 않다.",
        "built_space": "뒤쪽에 피아노 한 대의 일부와 그 위 랜턴 한 개가 보이고, 오른쪽에는 어두운 통로 입구와 커튼 일부가 있다. 낡은 회벽과 가구는 참조 장소와 대체로 일치하며 고정물 중복은 없다. 두 사람의 상체와 얼굴 아래쪽까지 포함해 요구된 손 디테일 인서트보다 넓다.",
        "entities": "검은 셔츠, 흰 성직자 칼라와 목걸이를 착용한 신부, 젊은 남성의 팔과 상처 난 손이 보인다. 보이는 얼굴 일부는 두 인물의 연령대와 대체로 맞지만 정확한 얼굴·머리 일치는 확인할 수 없다. 현우는 참조의 회색 긴팔 단추 셔츠 대신 회색 반팔 티셔츠를 입었다. 손전등과 하체는 프레임 밖이므로 부재로 감점하지 않는다. 추가 인물이나 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "현우의 손은 손목, 전완, 굽힌 팔꿈치, 상완을 거쳐 오른쪽 몸통에 자연스럽게 연결된다. 손가락이 신부의 소매를 강하게 모아 쥐어 천에 당김 주름이 생기며, 신부의 팔도 어깨에서 아래로 이어져 붙잡혀 멈추는 동작이 물리적으로 가능하다. 떠 있는 물체나 분리된 신체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.933
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.683
   },
   "violations": {
    "B": [
     "[gemini-pro] 신부의 몸통에서 이어진 검은 소매 팔에 현우의 멍든 손이 달려 있고 화면 밖에서 들어온 맨팔이 신부의 노인 손과 손전등을 쥐고 있는 등 해부학적으로 불가능한 신체 결합",
     "[gpt-high] 신부의 소매 위에 얹힌 붙잡는 손이 손목이나 전완으로 이어지지 않고 독립된 신체 조각처럼 끝나 있어, 현우의 팔로 지지되는 정상적인 해부학적 연결이 성립하지 않는다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 683
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "인물과 행동 묘사는 정확하나 인서트 클로즈업 프레이밍을 어기고 미디엄 샷으로 넓게 렌더링됨."
   },
   {
    "label": "B",
    "score": 683,
    "verdict_ko": "프레이밍은 지시문에 더 가깝지만, 두 인물의 팔과 복장, 피부가 해부학적으로 불가능하게 결합된 치명적 오류가 있음.  ★위반: [gemini-pro] 신부의 몸통에서 이어진 검은 소매 팔에 현우의 멍든 손이 달려 있고 화면 밖에서 들어온 맨팔이 신부의 노인 손과 손전등을 쥐고 있는 등 해부학적으로 불가능한 신체 결합 / [gpt-high] 신부의 소매 위에 얹힌 붙잡는 손이 손목이나 전완으로 이어지지 않고 독립된 신체 조각처럼 끝나 있어, 현우의 팔로 지지되는 정상적인 해부학적 연결이 성립하지 않는다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S32sh8_sel.png",
    "asset_id": "ac11848d-edd2-4f66-b2ae-c6ea63061703",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-4367-759a-972b-a2c575e63833",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S32sh8"
  }
 },
 "S32sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:29:28.134790+00:00",
  "fingerprint": "d9e351abc8140586b20e1255b0d525d7aaf3e3ef0f1577190dab6e2c837283f1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S32sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S32sh9_sel.png",
  "source_sha256": "51c49ad491cd74f673d4bc5a9a0f56e4d6f2c7f5423bbd91219c6a50593961d3",
  "file": "S32sh9_cine.png",
  "staged_sha256": "55c8c99fb7a4bfb4a80aeaec7ff009a328cd38fb8cc7165e58c0ad0e73bfa7be",
  "latency_ms": 10491
 },
 "S33sh6::signage": {
  "fp": "8ecdc67edb389acf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::27402fa9685326d1": {
  "subjects": [],
  "subject_text": "박철진이 탑승한 전투 헬기 내부\n조종석과 조수석, 통신 장비가 밀집한 비좁은 항공기 실내. 전면과 측면 창을 통해 아래쪽 지형이 내려다보인다.",
  "identity": "canonical",
  "scope_id": "L190",
  "scope_role": "location_interior",
  "scope_sha": "ef30841e56478db8"
 },
 "S33sh6::bgfirst_bg": {
  "input_fingerprint": "914dd9415c65b141",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S33sh6__bgfirst_bg.png",
  "asset_id": "d53a2385-242e-4c71-bb6c-2e769e64e6d2",
  "input_asset_ids": [
   "b50d98cd-e71c-4b90-928e-99153b8d142f",
   "c50aa81f-5e1c-4dd9-bbb0-46f0d8e60222"
  ]
 },
 "S33sh6": {
  "input_fingerprint": "212373fbfbae63ae",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A basement access door has been located in the church floor. The intercut seawall still has severe cracks and seepage, with a vehicle convoy travelling along the embankment road.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A basement access door has been located in the church floor. The intercut seawall still has severe cracks and seepage, with a vehicle convoy travelling along the embankment road.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A basement access door has been located in the church floor. The intercut seawall still has severe cracks and seepage, with a vehicle convoy travelling along the embankment road.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S33sh6__bgfirst_bg.png",
     "asset_id": "d53a2385-242e-4c71-bb6c-2e769e64e6d2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S33sh6.png",
     "asset_id": "b50d98cd-e71c-4b90-928e-99153b8d142f",
     "role": "conti_light"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B02.png",
     "asset_id": "c50aa81f-5e1c-4dd9-bbb0-46f0d8e60222",
     "role": "location_plate"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "남성의 시선과 가리키는 손가락이 우측 하단의 나무 지하실 문을 정확히 향함.",
    "built_space": "레퍼런스와 일치하는 강대상, 나무 의자, 콘크리트 바닥, 지하실 문이 알맞은 위치에 배치됨.",
    "entities": "전술 장비를 착용한 한국인 남성 1명, 나무 지하실 문.",
    "hard_violations": [],
    "physics": "두 발을 바닥에 단단히 딛고 역동적으로 몸을 낮춘 안정적인 자세."
   },
   {
    "label": "B",
    "direction": "남성의 시선은 앞을 향하며, 손가락은 바닥의 지하실 문을 가리킴.",
    "built_space": "강대상과 의자가 보이나 레퍼런스의 정확한 공간 배치와는 차이가 있음.",
    "entities": "주 인물 1명 외에 무장을 한 병사 3명이 배경에 추가로 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장",
     "[gpt-high] 샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
    ],
    "physics": "두 발로 바닥을 딛고 구부정하게 서 있는 자세."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스 공간을 정확히 재현하고, 불필요한 인물 없이 단일 인물의 역동적인 동작과 외침을 프롬프트대로 완벽히 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트에 명시되지 않은 배경 인물들이 다수 등장하여 인물 제한 규칙을 심각하게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남성의 시선과 가리키는 손가락이 우측 하단의 나무 지하실 문을 정확히 향함.",
        "built_space": "레퍼런스와 일치하는 강대상, 나무 의자, 콘크리트 바닥, 지하실 문이 알맞은 위치에 배치됨.",
        "entities": "전술 장비를 착용한 한국인 남성 1명, 나무 지하실 문.",
        "hard_violations": [],
        "physics": "두 발을 바닥에 단단히 딛고 역동적으로 몸을 낮춘 안정적인 자세."
       },
       {
        "label": "B",
        "direction": "남성의 시선은 앞을 향하며, 손가락은 바닥의 지하실 문을 가리킴.",
        "built_space": "강대상과 의자가 보이나 레퍼런스의 정확한 공간 배치와는 차이가 있음.",
        "entities": "주 인물 1명 외에 무장을 한 병사 3명이 배경에 추가로 존재함.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장"
        ],
        "physics": "두 발로 바닥을 딛고 구부정하게 서 있는 자세."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스 공간을 정확히 재현하고, 불필요한 인물 없이 단일 인물의 역동적인 동작과 외침을 프롬프트대로 완벽히 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트에 명시되지 않은 배경 인물들이 다수 등장하여 인물 제한 규칙을 심각하게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "남성의 시선과 가리키는 손가락이 우측 하단의 나무 지하실 문을 정확히 향함.",
        "built_space": "레퍼런스와 일치하는 강대상, 나무 의자, 콘크리트 바닥, 지하실 문이 알맞은 위치에 배치됨.",
        "entities": "전술 장비를 착용한 한국인 남성 1명, 나무 지하실 문.",
        "hard_violations": [],
        "physics": "두 발을 바닥에 단단히 딛고 역동적으로 몸을 낮춘 안정적인 자세."
       },
       {
        "label": "B",
        "direction": "남성의 시선은 앞을 향하며, 손가락은 바닥의 지하실 문을 가리킴.",
        "built_space": "강대상과 의자가 보이나 레퍼런스의 정확한 공간 배치와는 차이가 있음.",
        "entities": "주 인물 1명 외에 무장을 한 병사 3명이 배경에 추가로 존재함.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장"
        ],
        "physics": "두 발로 바닥을 딛고 구부정하게 서 있는 자세."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지하실 문을 가리키며 외치는 동작은 보이지만, 명시되지 않은 인물 최소 3명이 추가되어 실격이며 구도도 요구한 미디엄 숏보다 넓습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "한 명의 민병대원, 오른쪽 아래 나무문, 참고 장소와 외치는 동작을 충실히 구현했지만, 전신과 넓은 바닥을 담아 미디엄 숏 지시에는 실패했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앞쪽 남성의 검지는 오른쪽 아래로 뻗어 있으며 연장 방향이 지하실 나무문 왼쪽 부분에 닿습니다. 얼굴과 시선은 문보다 화면 오른쪽 바깥을 향해 누군가에게 외치는 모습입니다. 왼쪽 배경 인물이 든 총의 총구는 오른쪽 아래 바닥을 향하며, 특정 사람을 겨누지는 않습니다.",
        "built_space": "닫힌 바닥 나무문 1개와 왼쪽 경첩 2개, 뒤쪽 강단 1개와 긴 의자 1개가 보입니다. 주인물은 문 왼쪽 바닥에 웅크려 있습니다. 콘크리트 바닥은 유사하지만 참고 사진과 달리 뒤쪽 창문이 두드러지고 강단 주변 배치도 다릅니다. 나무문은 오른쪽 아래에 있고 인물보다 작지만, 인물의 하퇴까지 담은 구도로 미디엄 숏보다 넓습니다.",
        "entities": "주인물은 한국인 설정에 부합하는 외관의 성인 동아시아계 남성으로, 위장복과 전술 조끼를 착용하고 입을 크게 벌리고 있습니다. 바닥 출입문은 실제 목재 판재와 금속 경첩으로 표현되었습니다. 왼쪽, 중앙 뒤쪽, 오른쪽에 추가 군복 인물이 최소 3명 있어 한 명만 등장해야 하는 지시와 충돌합니다. 머리 조명과 무장도 추가되어 있습니다. 판독 가능한 글자는 보이지 않습니다.",
        "hard_violations": [
         "샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
        ],
        "physics": "주인물은 무릎을 굽히고 몸을 앞으로 기울였으며 뒤쪽 부츠가 바닥에 닿아 몸을 지탱합니다. 앞쪽 발은 화면 밖이므로 접지를 확인할 수 없지만 공중에 뜬 자세는 아닙니다. 가리키는 팔은 어깨에서 자연스럽게 이어지고, 조명은 머리띠에 고정되어 있으며 나무문은 바닥 틀에 받쳐져 있습니다. 배경 인물들도 바닥에 서 있습니다."
       },
       {
        "label": "B",
        "direction": "남성의 검지가 오른쪽 아래를 향하며 그 연장선은 나무문 윗면으로 이어집니다. 고개와 눈도 손끝과 문 쪽으로 내려가 있어 발견한 입구를 지목하는 관계가 명확합니다. 반대쪽 손은 균형을 잡듯 뒤로 벌어져 있습니다.",
        "built_space": "참고 사진과 같은 왼쪽 강단 1개, 그 뒤 작은 책장 1개, 뒤 벽의 긴 의자 1개, 오른쪽 벽 돌출부와 낡은 콘크리트 바닥이 보입니다. 바닥 나무문은 1개이고 왼쪽 경첩 2개가 보이며, 남성은 문 왼쪽의 빈 바닥에 서 있습니다. 문 윗면은 손 아래 오른쪽 하단에 비스듬히 보이고 인물보다 작습니다. 다만 두 발을 포함한 전신과 넓은 바닥을 보여 주어 지정된 미디엄 숏이 아니라 전신 숏에 가깝습니다.",
        "entities": "성인 동아시아계 남성 1명만 등장하며 한국인 민병대원 설정에 부합하는 외관입니다. 어두운 군용 복장과 전술 조끼를 착용하고 입을 크게 벌려 외칩니다. 나무 지하실 문, 금속 경첩, 예배소 바닥이 모두 식별되며 참고 장소의 재료와 색조도 잘 유지됩니다. 추가 인물이나 읽을 수 있는 글자는 없습니다. 프레임 밖의 방조제와 차량 행렬을 끌어들이지 않았습니다.",
        "hard_violations": [],
        "physics": "벌린 두 다리의 부츠가 모두 바닥에 닿고 무릎이 굽혀져 있어 앞으로 기울인 몸을 지탱합니다. 한쪽 팔을 뻗어 가리키고 다른 팔로 균형을 잡는 동작이 물리적으로 자연스럽습니다. 나무문은 바닥과 거의 같은 높이의 틀에 놓여 있으며 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지하실 문을 가리키며 외치는 동작은 보이지만, 명시되지 않은 인물 최소 3명이 추가되어 실격이며 구도도 요구한 미디엄 숏보다 넓습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "한 명의 민병대원, 오른쪽 아래 나무문, 참고 장소와 외치는 동작을 충실히 구현했지만, 전신과 넓은 바닥을 담아 미디엄 숏 지시에는 실패했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앞쪽 남성의 검지는 오른쪽 아래로 뻗어 있으며 연장 방향이 지하실 나무문 왼쪽 부분에 닿습니다. 얼굴과 시선은 문보다 화면 오른쪽 바깥을 향해 누군가에게 외치는 모습입니다. 왼쪽 배경 인물이 든 총의 총구는 오른쪽 아래 바닥을 향하며, 특정 사람을 겨누지는 않습니다.",
        "built_space": "닫힌 바닥 나무문 1개와 왼쪽 경첩 2개, 뒤쪽 강단 1개와 긴 의자 1개가 보입니다. 주인물은 문 왼쪽 바닥에 웅크려 있습니다. 콘크리트 바닥은 유사하지만 참고 사진과 달리 뒤쪽 창문이 두드러지고 강단 주변 배치도 다릅니다. 나무문은 오른쪽 아래에 있고 인물보다 작지만, 인물의 하퇴까지 담은 구도로 미디엄 숏보다 넓습니다.",
        "entities": "주인물은 한국인 설정에 부합하는 외관의 성인 동아시아계 남성으로, 위장복과 전술 조끼를 착용하고 입을 크게 벌리고 있습니다. 바닥 출입문은 실제 목재 판재와 금속 경첩으로 표현되었습니다. 왼쪽, 중앙 뒤쪽, 오른쪽에 추가 군복 인물이 최소 3명 있어 한 명만 등장해야 하는 지시와 충돌합니다. 머리 조명과 무장도 추가되어 있습니다. 판독 가능한 글자는 보이지 않습니다.",
        "hard_violations": [
         "샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
        ],
        "physics": "주인물은 무릎을 굽히고 몸을 앞으로 기울였으며 뒤쪽 부츠가 바닥에 닿아 몸을 지탱합니다. 앞쪽 발은 화면 밖이므로 접지를 확인할 수 없지만 공중에 뜬 자세는 아닙니다. 가리키는 팔은 어깨에서 자연스럽게 이어지고, 조명은 머리띠에 고정되어 있으며 나무문은 바닥 틀에 받쳐져 있습니다. 배경 인물들도 바닥에 서 있습니다."
       },
       {
        "label": "A",
        "direction": "남성의 검지가 오른쪽 아래를 향하며 그 연장선은 나무문 윗면으로 이어집니다. 고개와 눈도 손끝과 문 쪽으로 내려가 있어 발견한 입구를 지목하는 관계가 명확합니다. 반대쪽 손은 균형을 잡듯 뒤로 벌어져 있습니다.",
        "built_space": "참고 사진과 같은 왼쪽 강단 1개, 그 뒤 작은 책장 1개, 뒤 벽의 긴 의자 1개, 오른쪽 벽 돌출부와 낡은 콘크리트 바닥이 보입니다. 바닥 나무문은 1개이고 왼쪽 경첩 2개가 보이며, 남성은 문 왼쪽의 빈 바닥에 서 있습니다. 문 윗면은 손 아래 오른쪽 하단에 비스듬히 보이고 인물보다 작습니다. 다만 두 발을 포함한 전신과 넓은 바닥을 보여 주어 지정된 미디엄 숏이 아니라 전신 숏에 가깝습니다.",
        "entities": "성인 동아시아계 남성 1명만 등장하며 한국인 민병대원 설정에 부합하는 외관입니다. 어두운 군용 복장과 전술 조끼를 착용하고 입을 크게 벌려 외칩니다. 나무 지하실 문, 금속 경첩, 예배소 바닥이 모두 식별되며 참고 장소의 재료와 색조도 잘 유지됩니다. 추가 인물이나 읽을 수 있는 글자는 없습니다. 프레임 밖의 방조제와 차량 행렬을 끌어들이지 않았습니다.",
        "hard_violations": [],
        "physics": "벌린 두 다리의 부츠가 모두 바닥에 닿고 무릎이 굽혀져 있어 앞으로 기울인 몸을 지탱합니다. 한쪽 팔을 뻗어 가리키고 다른 팔로 균형을 잡는 동작이 물리적으로 자연스럽습니다. 나무문은 바닥과 거의 같은 높이의 틀에 놓여 있으며 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.536
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.286
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장",
     "[gpt-high] 샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 286
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스 공간을 정확히 재현하고, 불필요한 인물 없이 단일 인물의 역동적인 동작과 외침을 프롬프트대로 완벽히 구현함."
   },
   {
    "label": "B",
    "score": 286,
    "verdict_ko": "프롬프트에 명시되지 않은 배경 인물들이 다수 등장하여 인물 제한 규칙을 심각하게 위반함.  ★위반: [gemini-pro] 프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장 / [gpt-high] 샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B02.png",
    "asset_id": "c50aa81f-5e1c-4dd9-bbb0-46f0d8e60222",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-450c-7c78-a975-a10ada96450b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S33sh6__bgfirst_bg.png",
   "bg_asset_id": "d53a2385-242e-4c71-bb6c-2e769e64e6d2",
   "bg_record_key": "S33sh6::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S33sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:50:36.778648+00:00",
  "fingerprint": "bed0e94769a6575c1c9ebe7b3230dd702845b5425b97e299b29ce1dca1b7ae1b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S33sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S33sh6_sel.png",
  "source_sha256": "6fc35f7e4a53547014f6a4b73d510b615407200c3fa1a4763d22faec77dbff14",
  "file": "S33sh6_cine.png",
  "staged_sha256": "8059d9853c2b3426f65f35dc932b019c99fa0e02c0aee8b4e1f3fbe856f795fd",
  "latency_ms": 12141
 },
 "S33sh11::signage": {
  "fp": "13e335e6bed67d92",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S33sh11": {
  "input_fingerprint": "38a1c48a1c0f8f63",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 인공제방 위를 달리는 차량을 향해 벌떼처럼 빽빽하게 날아드는 소형 드론들의 실루엣 전경.\n\nLOCATION (lock): In the open air over the seawall road, above a convoy of fleeing vehicles near the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching drone swarm in the lower-left of the frame, foreground, moves toward Convoy on the embankment road; Convoy on the embankment road in the upper-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Approaching small drones (Flying densely toward the convoy) — Rear and side profiles face the camera as the drones converge diagonally into depth; used as Create layered foreground movement, with each drone small enough to preserve the convoy's visibility; Convoy on the embankment road (Travelling in a single line) — Vehicle sides and rear quarters recede along the diagonal road; used as Provide the destination of the drone movement and establish the continuing travel axis; Artificial embankment (Intact before the breach) — The crest and inland face are visible obliquely; used as Separate foreground aerial space from the elevated convoy route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hold a restrained nighttime tonal range, separating the dark drone silhouettes from the more legible road and convoy without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains deeply cracked and leaking, with repair machinery still nearby. Drones are converging on the convoy and surrounding containers before the explosive breach.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 인공제방 위를 달리는 차량을 향해 벌떼처럼 빽빽하게 날아드는 소형 드론들의 실루엣 전경.\n\nLOCATION (lock): In the open air over the seawall road, above a convoy of fleeing vehicles near the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching drone swarm in the lower-left of the frame, foreground, moves toward Convoy on the embankment road; Convoy on the embankment road in the upper-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Approaching small drones (Flying densely toward the convoy) — Rear and side profiles face the camera as the drones converge diagonally into depth; used as Create layered foreground movement, with each drone small enough to preserve the convoy's visibility; Convoy on the embankment road (Travelling in a single line) — Vehicle sides and rear quarters recede along the diagonal road; used as Provide the destination of the drone movement and establish the continuing travel axis; Artificial embankment (Intact before the breach) — The crest and inland face are visible obliquely; used as Separate foreground aerial space from the elevated convoy route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hold a restrained nighttime tonal range, separating the dark drone silhouettes from the more legible road and convoy without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains deeply cracked and leaking, with repair machinery still nearby. Drones are converging on the convoy and surrounding containers before the explosive breach.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 인공제방 위를 달리는 차량을 향해 벌떼처럼 빽빽하게 날아드는 소형 드론들의 실루엣 전경.\n\nLOCATION (lock): In the open air over the seawall road, above a convoy of fleeing vehicles near the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching drone swarm in the lower-left of the frame, foreground, moves toward Convoy on the embankment road; Convoy on the embankment road in the upper-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Approaching small drones (Flying densely toward the convoy) — Rear and side profiles face the camera as the drones converge diagonally into depth; used as Create layered foreground movement, with each drone small enough to preserve the convoy's visibility; Convoy on the embankment road (Travelling in a single line) — Vehicle sides and rear quarters recede along the diagonal road; used as Provide the destination of the drone movement and establish the continuing travel axis; Artificial embankment (Intact before the breach) — The crest and inland face are visible obliquely; used as Separate foreground aerial space from the elevated convoy route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hold a restrained nighttime tonal range, separating the dark drone silhouettes from the more legible road and convoy without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains deeply cracked and leaking, with repair machinery still nearby. Drones are converging on the convoy and surrounding containers before the explosive breach.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "드론들이 좌측에서 우측으로 비행하며 도로 위의 호송대를 향하고 있음.",
    "built_space": "레퍼런스와 동일하게 우측에 바다와 테트라포드가 있는 도로 구도를 띠고 있으나, 요구된 내륙 측면의 단차와 공간 분리가 명확하지 않음.",
    "entities": "드론 떼와 차량 호송대는 존재하나, 프롬프트가 요구한 수리 장비가 보이지 않음.",
    "hard_violations": [
     "[gemini-pro] 지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
    ],
    "physics": "공중에 떠 있어야 할 일부 드론들이 지지대나 그림자 없이 도로 바닥이나 구조물과 융합되어 있음."
   },
   {
    "label": "B",
    "direction": "좌측 하단의 드론 떼가 우측 상단의 고가 도로 위 호송대를 향해 상승하듯 대각선으로 날아가고 있음.",
    "built_space": "카메라가 제방의 내륙 측면을 비스듬히 바라보는 구도이며, 좌측 하단에 균열이 간 지면이 있고 우측 상단에 호송대가 달리는 고가 도로가 배치되어 공간이 분리됨.",
    "entities": "소형 드론 떼, 트럭과 승용차로 이루어진 호송대, 깊게 갈라진 콘크리트 제방, 균열 옆의 주황색 수리 장비가 모두 명확히 확인됨.",
    "hard_violations": [],
    "physics": "드론들은 공중에 안정적으로 떠 있으며, 차량과 수리 장비는 구조물 위에 올바르게 안착되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 프레임 레이아웃(좌측 하단 드론, 우측 상단 호송대)과 내륙 측면 앵글을 정확히 구현했으며, 균열과 수리 장비까지 모두 포함하여 지시사항을 훌륭하게 충족합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "요구된 내륙 측면 구도와 수리 장비가 누락되었으며, 전경의 드론들이 지면을 관통하는 치명적인 물리적 오류가 있어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "드론들이 좌측에서 우측으로 비행하며 도로 위의 호송대를 향하고 있음.",
        "built_space": "레퍼런스와 동일하게 우측에 바다와 테트라포드가 있는 도로 구도를 띠고 있으나, 요구된 내륙 측면의 단차와 공간 분리가 명확하지 않음.",
        "entities": "드론 떼와 차량 호송대는 존재하나, 프롬프트가 요구한 수리 장비가 보이지 않음.",
        "hard_violations": [
         "지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
        ],
        "physics": "공중에 떠 있어야 할 일부 드론들이 지지대나 그림자 없이 도로 바닥이나 구조물과 융합되어 있음."
       },
       {
        "label": "B",
        "direction": "좌측 하단의 드론 떼가 우측 상단의 고가 도로 위 호송대를 향해 상승하듯 대각선으로 날아가고 있음.",
        "built_space": "카메라가 제방의 내륙 측면을 비스듬히 바라보는 구도이며, 좌측 하단에 균열이 간 지면이 있고 우측 상단에 호송대가 달리는 고가 도로가 배치되어 공간이 분리됨.",
        "entities": "소형 드론 떼, 트럭과 승용차로 이루어진 호송대, 깊게 갈라진 콘크리트 제방, 균열 옆의 주황색 수리 장비가 모두 명확히 확인됨.",
        "hard_violations": [],
        "physics": "드론들은 공중에 안정적으로 떠 있으며, 차량과 수리 장비는 구조물 위에 올바르게 안착되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 프레임 레이아웃(좌측 하단 드론, 우측 상단 호송대)과 내륙 측면 앵글을 정확히 구현했으며, 균열과 수리 장비까지 모두 포함하여 지시사항을 훌륭하게 충족합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "요구된 내륙 측면 구도와 수리 장비가 누락되었으며, 전경의 드론들이 지면을 관통하는 치명적인 물리적 오류가 있어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "드론들이 좌측에서 우측으로 비행하며 도로 위의 호송대를 향하고 있음.",
        "built_space": "레퍼런스와 동일하게 우측에 바다와 테트라포드가 있는 도로 구도를 띠고 있으나, 요구된 내륙 측면의 단차와 공간 분리가 명확하지 않음.",
        "entities": "드론 떼와 차량 호송대는 존재하나, 프롬프트가 요구한 수리 장비가 보이지 않음.",
        "hard_violations": [
         "지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
        ],
        "physics": "공중에 떠 있어야 할 일부 드론들이 지지대나 그림자 없이 도로 바닥이나 구조물과 융합되어 있음."
       },
       {
        "label": "B",
        "direction": "좌측 하단의 드론 떼가 우측 상단의 고가 도로 위 호송대를 향해 상승하듯 대각선으로 날아가고 있음.",
        "built_space": "카메라가 제방의 내륙 측면을 비스듬히 바라보는 구도이며, 좌측 하단에 균열이 간 지면이 있고 우측 상단에 호송대가 달리는 고가 도로가 배치되어 공간이 분리됨.",
        "entities": "소형 드론 떼, 트럭과 승용차로 이루어진 호송대, 깊게 갈라진 콘크리트 제방, 균열 옆의 주황색 수리 장비가 모두 명확히 확인됨.",
        "hard_violations": [],
        "physics": "드론들은 공중에 안정적으로 떠 있으며, 차량과 수리 장비는 구조물 위에 올바르게 안착되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "좌하단 드론 전경과 우상단 차량 중경, 차량의 후측면 구도가 더 충실하지만, 여러 줄의 차량과 과도하게 벌어진 제방 손상은 지시와 다릅니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "드론과 차량 사이의 대각선 연결은 보이지만, 차량 정면이 카메라로 다가오고 드론이 도로 전반에 퍼져 지정된 후측면 행렬과 전경 구도에서 벗어납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "드론 무리는 좌하단에서 중앙의 제방 아래쪽으로 좁아지며, 그 위 우상단 도로의 차량들이 접근 대상으로 읽힙니다. 다만 개별 기체의 전후 방향은 실루엣만으로 확정하기 어렵고, 무리의 수렴점도 차량보다는 제방 벽 아래에 가깝습니다. 차량은 후미와 측면을 보이며 도로를 따라 좌상단 원경으로 이어집니다.",
        "built_space": "높은 콘크리트 제방 하나와 상부 도로 하나, 도로 양쪽의 연속 난간이 보입니다. 경사 벽과 소파블록은 참고 장소의 주요 구조를 유지합니다. 하부에는 넓은 콘크리트 작업면과 장비 두 대가 보이며, 벽 중앙부터 작업면까지 큰 균열과 떨어져 나온 덩어리가 이어집니다. 상부 도로는 연결되어 있지만 하부 손상은 단순한 균열보다 붕괴에 가깝습니다. 차량은 한 줄이 아니라 여러 줄로 배치되어 있습니다.",
        "entities": "수십 대의 소형 다중회전익 드론, 화물차와 승용차 행렬, 콘크리트 제방, 난간, 소파블록, 물과 보수용으로 보이는 장비가 있습니다. 사람과 얼굴은 보이지 않습니다. 별도의 정착지 컨테이너는 식별되지 않으며 트럭 적재함만 명확합니다. 야간의 젖은 재질과 차량 불빛은 부합합니다. 읽을 수 있는 문구는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "드론에는 회전익과 회전 흐림이 보여 비행을 지탱하는 추진 장치가 있습니다. 차량 바퀴는 상부 도로에, 장비는 하부 작업면에 놓여 있습니다. 부서진 콘크리트도 하부 지면에 쌓여 있어 무지지 부유물은 보이지 않습니다. 물 고임과 젖은 흔적은 있지만 균열에서 실제로 새어 나오는 물줄기는 분명하지 않습니다."
       },
       {
        "label": "B",
        "direction": "드론은 좌하단과 도로 전경에서 우상단 차량 행렬 쪽으로 펼쳐져 있습니다. 일부 날개형 기체는 그 방향으로 향하지만 다중회전익 기체들의 전후 방향은 일정하게 판독되지 않습니다. 가장 눈에 띄는 화물차와 여러 승용차는 전조등과 정면을 카메라 쪽으로 향하고 있어, 후측면을 보이며 깊이 들어가는 지정 차량 행렬과 다릅니다.",
        "built_space": "연속된 콘크리트 제방 하나, 상부 도로 하나, 양쪽 난간과 바깥 경사면 아래의 소파블록이 보입니다. 참고 장소의 벽 재질과 난간 구조는 잘 유지됩니다. 도로와 벽에 균열이 있으나 큰 개구부는 없습니다. 차량은 여러 줄이며 일부는 서로 반대 방향을 향합니다. 보수 장비는 식별되지 않고, 내륙측 벽면보다는 소파블록이 있는 바다측 경사면이 주로 드러납니다.",
        "entities": "다중회전익 드론과 날개형 소형 무인기, 화물차와 승용차, 난간, 콘크리트 제방, 소파블록과 수면이 보입니다. 사람이나 얼굴은 없습니다. 정착지 컨테이너와 보수 장비는 확인되지 않습니다. 야간 분위기와 젖은 도로는 맞으며 읽을 수 있는 글자는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "다중회전익 드론은 로터로 지지되는 비행 상태이며, 날개형 기체도 날개를 가진 비행체로 보입니다. 정지 화면이라 추진 세부는 불명확하지만 무지지 물체로 단정할 근거는 없습니다. 차량은 바퀴로 도로에 지지되고, 난간과 소파블록도 구조물과 지면에 연결되어 있습니다. 균열은 보이지만 누수는 명확하지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "좌하단 드론 전경과 우상단 차량 중경, 차량의 후측면 구도가 더 충실하지만, 여러 줄의 차량과 과도하게 벌어진 제방 손상은 지시와 다릅니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "드론과 차량 사이의 대각선 연결은 보이지만, 차량 정면이 카메라로 다가오고 드론이 도로 전반에 퍼져 지정된 후측면 행렬과 전경 구도에서 벗어납니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "드론 무리는 좌하단에서 중앙의 제방 아래쪽으로 좁아지며, 그 위 우상단 도로의 차량들이 접근 대상으로 읽힙니다. 다만 개별 기체의 전후 방향은 실루엣만으로 확정하기 어렵고, 무리의 수렴점도 차량보다는 제방 벽 아래에 가깝습니다. 차량은 후미와 측면을 보이며 도로를 따라 좌상단 원경으로 이어집니다.",
        "built_space": "높은 콘크리트 제방 하나와 상부 도로 하나, 도로 양쪽의 연속 난간이 보입니다. 경사 벽과 소파블록은 참고 장소의 주요 구조를 유지합니다. 하부에는 넓은 콘크리트 작업면과 장비 두 대가 보이며, 벽 중앙부터 작업면까지 큰 균열과 떨어져 나온 덩어리가 이어집니다. 상부 도로는 연결되어 있지만 하부 손상은 단순한 균열보다 붕괴에 가깝습니다. 차량은 한 줄이 아니라 여러 줄로 배치되어 있습니다.",
        "entities": "수십 대의 소형 다중회전익 드론, 화물차와 승용차 행렬, 콘크리트 제방, 난간, 소파블록, 물과 보수용으로 보이는 장비가 있습니다. 사람과 얼굴은 보이지 않습니다. 별도의 정착지 컨테이너는 식별되지 않으며 트럭 적재함만 명확합니다. 야간의 젖은 재질과 차량 불빛은 부합합니다. 읽을 수 있는 문구는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "드론에는 회전익과 회전 흐림이 보여 비행을 지탱하는 추진 장치가 있습니다. 차량 바퀴는 상부 도로에, 장비는 하부 작업면에 놓여 있습니다. 부서진 콘크리트도 하부 지면에 쌓여 있어 무지지 부유물은 보이지 않습니다. 물 고임과 젖은 흔적은 있지만 균열에서 실제로 새어 나오는 물줄기는 분명하지 않습니다."
       },
       {
        "label": "A",
        "direction": "드론은 좌하단과 도로 전경에서 우상단 차량 행렬 쪽으로 펼쳐져 있습니다. 일부 날개형 기체는 그 방향으로 향하지만 다중회전익 기체들의 전후 방향은 일정하게 판독되지 않습니다. 가장 눈에 띄는 화물차와 여러 승용차는 전조등과 정면을 카메라 쪽으로 향하고 있어, 후측면을 보이며 깊이 들어가는 지정 차량 행렬과 다릅니다.",
        "built_space": "연속된 콘크리트 제방 하나, 상부 도로 하나, 양쪽 난간과 바깥 경사면 아래의 소파블록이 보입니다. 참고 장소의 벽 재질과 난간 구조는 잘 유지됩니다. 도로와 벽에 균열이 있으나 큰 개구부는 없습니다. 차량은 여러 줄이며 일부는 서로 반대 방향을 향합니다. 보수 장비는 식별되지 않고, 내륙측 벽면보다는 소파블록이 있는 바다측 경사면이 주로 드러납니다.",
        "entities": "다중회전익 드론과 날개형 소형 무인기, 화물차와 승용차, 난간, 콘크리트 제방, 소파블록과 수면이 보입니다. 사람이나 얼굴은 없습니다. 정착지 컨테이너와 보수 장비는 확인되지 않습니다. 야간 분위기와 젖은 도로는 맞으며 읽을 수 있는 글자는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "다중회전익 드론은 로터로 지지되는 비행 상태이며, 날개형 기체도 날개를 가진 비행체로 보입니다. 정지 화면이라 추진 세부는 불명확하지만 무지지 물체로 단정할 근거는 없습니다. 차량은 바퀴로 도로에 지지되고, 난간과 소파블록도 구조물과 지면에 연결되어 있습니다. 균열은 보이지만 누수는 명확하지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.214,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.964,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 964
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 프레임 레이아웃(좌측 하단 드론, 우측 상단 호송대)과 내륙 측면 앵글을 정확히 구현했으며, 균열과 수리 장비까지 모두 포함하여 지시사항을 훌륭하게 충족합니다."
   },
   {
    "label": "A",
    "score": 964,
    "verdict_ko": "요구된 내륙 측면 구도와 수리 장비가 누락되었으며, 전경의 드론들이 지면을 관통하는 치명적인 물리적 오류가 있어 감점되었습니다.  ★위반: [gemini-pro] 지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B01.png",
    "asset_id": "9dbf00e6-e0e7-42db-8305-7544a83a95ac",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-4845-7921-b9d6-16112ec905e7",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S33sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:51:52.255857+00:00",
  "fingerprint": "cbeb9063194db515dd6fc26bcf870403d74a722d8bdcc7b1ec708dd95c090be6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S33sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S33sh11_sel.png",
  "source_sha256": "2637fa70cf1e02aefdb2cce477e1124332c670ed752636c5cdd3356913474ec1",
  "file": "S33sh11_cine.png",
  "staged_sha256": "fb7dc11330436265b14e320ca3185ac47889631581e726d4cbd1660b8f9a8b39",
  "latency_ms": 13828
 },
 "S33sh14::signage": {
  "fp": "fb42e2bd9ca127c6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S33sh14": {
  "input_fingerprint": "245ab460d0a4f27e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 폭발로 산산조각이 나며 무너져 내린 제방 틈새로 거대한 검은 바닷물이 폭포수처럼 쏟아져 들어오는 압도적인 광경.\n\nLOCATION (lock): At the breached upper section of the coastal seawall, where seawater pours into the refugee settlement at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Breached upper embankment (Broken by explosions and collapsing) — The inland face and broken edges of the upper opening are seen obliquely from below; used as Bracket the falling water and retain the established wall orientation during withdrawal; Seawater pouring through the breach (Rushing inward in a massive descending flow); used as Carry the principal movement from the elevated opening toward the bottom of the frame; Repair machinery and construction equipment (Being swept away by seawater) — Only partial forms remain visible below the breach near the lower frame edge; used as Supply scale without redirecting attention away from the breach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime exposure with controlled tonal separation between black seawater, broken embankment edges, and the remaining wall.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall's upper section has been blown apart, opening a breach through which seawater pours into the settlement. Nearby repair machinery and construction equipment are being swept away.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 폭발로 산산조각이 나며 무너져 내린 제방 틈새로 거대한 검은 바닷물이 폭포수처럼 쏟아져 들어오는 압도적인 광경.\n\nLOCATION (lock): At the breached upper section of the coastal seawall, where seawater pours into the refugee settlement at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Breached upper embankment (Broken by explosions and collapsing) — The inland face and broken edges of the upper opening are seen obliquely from below; used as Bracket the falling water and retain the established wall orientation during withdrawal; Seawater pouring through the breach (Rushing inward in a massive descending flow); used as Carry the principal movement from the elevated opening toward the bottom of the frame; Repair machinery and construction equipment (Being swept away by seawater) — Only partial forms remain visible below the breach near the lower frame edge; used as Supply scale without redirecting attention away from the breach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime exposure with controlled tonal separation between black seawater, broken embankment edges, and the remaining wall.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall's upper section has been blown apart, opening a breach through which seawater pours into the settlement. Nearby repair machinery and construction equipment are being swept away.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 폭발로 산산조각이 나며 무너져 내린 제방 틈새로 거대한 검은 바닷물이 폭포수처럼 쏟아져 들어오는 압도적인 광경.\n\nLOCATION (lock): At the breached upper section of the coastal seawall, where seawater pours into the refugee settlement at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Breached upper embankment (Broken by explosions and collapsing) — The inland face and broken edges of the upper opening are seen obliquely from below; used as Bracket the falling water and retain the established wall orientation during withdrawal; Seawater pouring through the breach (Rushing inward in a massive descending flow); used as Carry the principal movement from the elevated opening toward the bottom of the frame; Repair machinery and construction equipment (Being swept away by seawater) — Only partial forms remain visible below the breach near the lower frame edge; used as Supply scale without redirecting attention away from the breach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime exposure with controlled tonal separation between black seawater, broken embankment edges, and the remaining wall.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall's upper section has been blown apart, opening a breach through which seawater pours into the settlement. Nearby repair machinery and construction equipment are being swept away.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "거대한 바닷물이 무너진 제방 틈새에서 화면 하단을 향해 폭포수처럼 쏟아져 내리고 있음.",
    "built_space": "도로와 난간이 있는 거대한 콘크리트 방조제 중앙이 크게 붕괴되어 있으며, 화면 아래쪽 끄트머리에 건설 장비들이 위치함.",
    "entities": "검은 바닷물(일치), 무너진 상단 제방(일치), 화면 하단의 건설 장비 일부(일치).",
    "hard_violations": [
     "[gpt-high] 파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
    ],
    "physics": "쏟아지는 물결은 중력에 의해 자연스럽게 낙하하며, 차량과 중장비들은 바닥이나 남은 구조물에 올바르게 지탱되어 있음."
   },
   {
    "label": "B",
    "direction": "방조제의 무너진 틈에서 바닷물이 화면 중앙에서 아래쪽을 향해 거세게 쏟아지고 있음.",
    "built_space": "무너진 방조제 왼편으로 판자촌 건물이 일부 보이며, 아래쪽에는 굴삭기, 트럭, 대형 철골 구조물들이 어지럽게 놓여 있음.",
    "entities": "바닷물(흰 거품이 두드러져 검은 바닷물 지시 불일치), 무너진 제방(일치), 건설 장비(프레임 하단에 부분적으로 보이라는 지시를 어기고 온전하게 전경을 차지함).",
    "hard_violations": [],
    "physics": "물이 중력에 따라 쏟아져 내리며, 굴삭기와 트럭 등은 지면에 제대로 지탱되어 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '검은 바닷물'의 톤을 잘 살렸으며, 화면 하단 가장자리에 건설 장비가 부분적으로만 보이도록 한 구도 지시를 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "바닷물에 흰 거품이 많아 '검은 바닷물' 지시와 거리가 멀고, 전경에 배치된 건설 장비와 철골이 너무 커서 시선을 분산시킵니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "거대한 바닷물이 무너진 제방 틈새에서 화면 하단을 향해 폭포수처럼 쏟아져 내리고 있음.",
        "built_space": "도로와 난간이 있는 거대한 콘크리트 방조제 중앙이 크게 붕괴되어 있으며, 화면 아래쪽 끄트머리에 건설 장비들이 위치함.",
        "entities": "검은 바닷물(일치), 무너진 상단 제방(일치), 화면 하단의 건설 장비 일부(일치).",
        "hard_violations": [],
        "physics": "쏟아지는 물결은 중력에 의해 자연스럽게 낙하하며, 차량과 중장비들은 바닥이나 남은 구조물에 올바르게 지탱되어 있음."
       },
       {
        "label": "B",
        "direction": "방조제의 무너진 틈에서 바닷물이 화면 중앙에서 아래쪽을 향해 거세게 쏟아지고 있음.",
        "built_space": "무너진 방조제 왼편으로 판자촌 건물이 일부 보이며, 아래쪽에는 굴삭기, 트럭, 대형 철골 구조물들이 어지럽게 놓여 있음.",
        "entities": "바닷물(흰 거품이 두드러져 검은 바닷물 지시 불일치), 무너진 제방(일치), 건설 장비(프레임 하단에 부분적으로 보이라는 지시를 어기고 온전하게 전경을 차지함).",
        "hard_violations": [],
        "physics": "물이 중력에 따라 쏟아져 내리며, 굴삭기와 트럭 등은 지면에 제대로 지탱되어 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '검은 바닷물'의 톤을 잘 살렸으며, 화면 하단 가장자리에 건설 장비가 부분적으로만 보이도록 한 구도 지시를 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "바닷물에 흰 거품이 많아 '검은 바닷물' 지시와 거리가 멀고, 전경에 배치된 건설 장비와 철골이 너무 커서 시선을 분산시킵니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "거대한 바닷물이 무너진 제방 틈새에서 화면 하단을 향해 폭포수처럼 쏟아져 내리고 있음.",
        "built_space": "도로와 난간이 있는 거대한 콘크리트 방조제 중앙이 크게 붕괴되어 있으며, 화면 아래쪽 끄트머리에 건설 장비들이 위치함.",
        "entities": "검은 바닷물(일치), 무너진 상단 제방(일치), 화면 하단의 건설 장비 일부(일치).",
        "hard_violations": [],
        "physics": "쏟아지는 물결은 중력에 의해 자연스럽게 낙하하며, 차량과 중장비들은 바닥이나 남은 구조물에 올바르게 지탱되어 있음."
       },
       {
        "label": "B",
        "direction": "방조제의 무너진 틈에서 바닷물이 화면 중앙에서 아래쪽을 향해 거세게 쏟아지고 있음.",
        "built_space": "무너진 방조제 왼편으로 판자촌 건물이 일부 보이며, 아래쪽에는 굴삭기, 트럭, 대형 철골 구조물들이 어지럽게 놓여 있음.",
        "entities": "바닷물(흰 거품이 두드러져 검은 바닷물 지시 불일치), 무너진 제방(일치), 건설 장비(프레임 하단에 부분적으로 보이라는 지시를 어기고 온전하게 전경을 차지함).",
        "hard_violations": [],
        "physics": "물이 중력에 따라 쏟아져 내리며, 굴삭기와 트럭 등은 지면에 제대로 지탱되어 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "파괴된 제방 사이에서 정착지로 낙하하는 검은 해수와 장소의 재질을 잘 구현했지만, 장비를 하단의 일부 형태로만 보여야 한다는 구도보다 노출이 많습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "거대한 낙수와 하단에 잘린 장비는 적절하지만, 정상 난간보다 낮은 틈새 뒤로 온전한 차량 도로를 추가해 기준 제방의 도로 배치를 어긋나게 했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "물은 중앙의 파괴된 개구부에서 카메라 쪽 내륙 공간으로 넘어와 화면 아래로 떨어집니다. 하단에서는 포말이 굴착기와 덤프트럭 주변으로 퍼집니다. 사람의 시선이나 무기는 없으며, 굴착기 버킷은 아래의 침수 구역을 향합니다.",
        "built_space": "하나의 콘크리트 제방이 중앙의 큰 파열부 양쪽으로 이어지고, 상단 난간도 그 양쪽에 남아 있습니다. 파열부 왼쪽에는 철근이 드러난 도로 슬래브 잔해가 보입니다. 왼쪽 제방 아래에는 낮은 정착지 건물들이, 하단에는 굴착기와 트럭 및 강재가 있습니다. 기준 사진의 경사진 콘크리트 벽체와 금속 난간은 유지되지만, 장비와 바닥을 넓게 내려다보여 개구부 아래에서 올려다보는 시점은 약합니다.",
        "entities": "검은 해수, 흰 포말, 부서진 콘크리트와 철근, 남은 제방, 수리·건설 장비가 보입니다. 왼쪽 굴착기와 중앙 덤프트럭은 대부분의 형태가 드러나며, 오른쪽 아래에는 다른 장비 일부가 잘려 있습니다. 인물이나 얼굴, 읽을 수 있는 문구는 보이지 않습니다. 야간의 낮은 노출과 젖은 콘크리트·금속의 재질은 요청에 부합합니다.",
        "hard_violations": [],
        "physics": "물은 높은 개구부에서 중력 방향으로 낙하하고 아래 수면에서 포말을 만듭니다. 차량 하부와 굴착기 궤도는 침수 구역에 잠겨 있으며, 굴착기 팔과 버킷은 기계 관절에 연결되어 있습니다. 강재는 잔해와 바닥에 걸쳐 있습니다. 지지 없이 공중에 뜬 물체는 보이지 않지만, 장비가 실제로 떠밀리는 순간보다는 물에 잠긴 상태가 더 명확합니다."
       },
       {
        "label": "B",
        "direction": "검은 물은 중앙 상부의 틈에서 전경 아래로 폭포처럼 떨어져 오른쪽과 하단 장비를 덮습니다. 굴착기 팔은 오른쪽 아래로 뻗어 있습니다. 사람의 시선이나 겨냥하는 무기는 없습니다. 틈새 뒤 차량들은 물의 낙하 방향을 가로지르는 도로에 놓여 있습니다.",
        "built_space": "좌우에 파손된 콘크리트 제방과 정상 난간이 있고, 오른쪽 벽 아래에는 정착지 건물들이 이어집니다. 그런데 두 정상 난간보다 낮은 위치에서 차량 여러 대가 선 온전한 도로가 파열부 뒤를 가로지릅니다. 이는 기준 사진의 제방 정상 도로와 별개의 낮은 도로층처럼 보여 장소의 고정 구조와 맞지 않습니다. 하단 장비는 부분적으로 잘렸지만, 전경 운전실과 오른쪽 굴착기 팔이 상당한 면적을 차지합니다.",
        "entities": "대규모 검은 낙수, 부서진 콘크리트 가장자리, 노출 철근, 제방 난간과 노란 건설 장비가 있습니다. 배경에는 승용차 네 대가 식별되며, 오른쪽에는 작은 건물들과 노란 상자형 설비가 보입니다. 인물이나 얼굴, 읽을 수 있는 글자는 없습니다. 야간 분위기와 물·콘크리트·금속의 물성은 대체로 적절합니다.",
        "hard_violations": [
         "파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
        ],
        "physics": "낙수는 개구부에서 아래로 이어지고 장비 주변의 충돌 지점에 포말이 생깁니다. 전경 운전실은 일부 잠긴 차체에 연결되어 있고 오른쪽 굴착기 팔도 차체와 연결되어 있습니다. 기울어진 장비는 물에 휩쓸리는 상태로 읽힙니다. 차량은 배경 도로에 지지되어 있으며, 지지 없이 공중에 떠 있는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "파괴된 제방 사이에서 정착지로 낙하하는 검은 해수와 장소의 재질을 잘 구현했지만, 장비를 하단의 일부 형태로만 보여야 한다는 구도보다 노출이 많습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "거대한 낙수와 하단에 잘린 장비는 적절하지만, 정상 난간보다 낮은 틈새 뒤로 온전한 차량 도로를 추가해 기준 제방의 도로 배치를 어긋나게 했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "물은 중앙의 파괴된 개구부에서 카메라 쪽 내륙 공간으로 넘어와 화면 아래로 떨어집니다. 하단에서는 포말이 굴착기와 덤프트럭 주변으로 퍼집니다. 사람의 시선이나 무기는 없으며, 굴착기 버킷은 아래의 침수 구역을 향합니다.",
        "built_space": "하나의 콘크리트 제방이 중앙의 큰 파열부 양쪽으로 이어지고, 상단 난간도 그 양쪽에 남아 있습니다. 파열부 왼쪽에는 철근이 드러난 도로 슬래브 잔해가 보입니다. 왼쪽 제방 아래에는 낮은 정착지 건물들이, 하단에는 굴착기와 트럭 및 강재가 있습니다. 기준 사진의 경사진 콘크리트 벽체와 금속 난간은 유지되지만, 장비와 바닥을 넓게 내려다보여 개구부 아래에서 올려다보는 시점은 약합니다.",
        "entities": "검은 해수, 흰 포말, 부서진 콘크리트와 철근, 남은 제방, 수리·건설 장비가 보입니다. 왼쪽 굴착기와 중앙 덤프트럭은 대부분의 형태가 드러나며, 오른쪽 아래에는 다른 장비 일부가 잘려 있습니다. 인물이나 얼굴, 읽을 수 있는 문구는 보이지 않습니다. 야간의 낮은 노출과 젖은 콘크리트·금속의 재질은 요청에 부합합니다.",
        "hard_violations": [],
        "physics": "물은 높은 개구부에서 중력 방향으로 낙하하고 아래 수면에서 포말을 만듭니다. 차량 하부와 굴착기 궤도는 침수 구역에 잠겨 있으며, 굴착기 팔과 버킷은 기계 관절에 연결되어 있습니다. 강재는 잔해와 바닥에 걸쳐 있습니다. 지지 없이 공중에 뜬 물체는 보이지 않지만, 장비가 실제로 떠밀리는 순간보다는 물에 잠긴 상태가 더 명확합니다."
       },
       {
        "label": "A",
        "direction": "검은 물은 중앙 상부의 틈에서 전경 아래로 폭포처럼 떨어져 오른쪽과 하단 장비를 덮습니다. 굴착기 팔은 오른쪽 아래로 뻗어 있습니다. 사람의 시선이나 겨냥하는 무기는 없습니다. 틈새 뒤 차량들은 물의 낙하 방향을 가로지르는 도로에 놓여 있습니다.",
        "built_space": "좌우에 파손된 콘크리트 제방과 정상 난간이 있고, 오른쪽 벽 아래에는 정착지 건물들이 이어집니다. 그런데 두 정상 난간보다 낮은 위치에서 차량 여러 대가 선 온전한 도로가 파열부 뒤를 가로지릅니다. 이는 기준 사진의 제방 정상 도로와 별개의 낮은 도로층처럼 보여 장소의 고정 구조와 맞지 않습니다. 하단 장비는 부분적으로 잘렸지만, 전경 운전실과 오른쪽 굴착기 팔이 상당한 면적을 차지합니다.",
        "entities": "대규모 검은 낙수, 부서진 콘크리트 가장자리, 노출 철근, 제방 난간과 노란 건설 장비가 있습니다. 배경에는 승용차 네 대가 식별되며, 오른쪽에는 작은 건물들과 노란 상자형 설비가 보입니다. 인물이나 얼굴, 읽을 수 있는 글자는 없습니다. 야간 분위기와 물·콘크리트·금속의 물성은 대체로 적절합니다.",
        "hard_violations": [
         "파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
        ],
        "physics": "낙수는 개구부에서 아래로 이어지고 장비 주변의 충돌 지점에 포말이 생깁니다. 전경 운전실은 일부 잠긴 차체에 연결되어 있고 오른쪽 굴착기 팔도 차체와 연결되어 있습니다. 기울어진 장비는 물에 휩쓸리는 상태로 읽힙니다. 차량은 배경 도로에 지지되어 있으며, 지지 없이 공중에 떠 있는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.571
   },
   "violations": {
    "A": [
     "[gpt-high] 파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1321,
   "B": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "프롬프트가 요구한 '검은 바닷물'의 톤을 잘 살렸으며, 화면 하단 가장자리에 건설 장비가 부분적으로만 보이도록 한 구도 지시를 정확히 따랐습니다.  ★위반: [gpt-high] 파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "바닷물에 흰 거품이 많아 '검은 바닷물' 지시와 거리가 멀고, 전경에 배치된 건설 장비와 철골이 너무 커서 시선을 분산시킵니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B01.png",
    "asset_id": "9dbf00e6-e0e7-42db-8305-7544a83a95ac",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-49f0-7fe0-9bb1-b0bc848c9ae2",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S33sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:53:01.154664+00:00",
  "fingerprint": "18173c8537eba4d38bf859ec51d558d057bb7222ee07547151af661f2f374316",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S33sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S33sh14_sel.png",
  "source_sha256": "d34fb74b7fd0b1b0c5756ad3781b444ccd9886d97573b32ef05d1813e2cec3b1",
  "file": "S33sh14_cine.png",
  "staged_sha256": "bfcc2d65ce57981b2ae5f0c76e9fbe841b684acddfedf79a5aedbf15dfb725f1",
  "latency_ms": 12394
 },
 "S34sh3::signage": {
  "fp": "a8efc31dab6ae9e0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::698a95a010babfcd": {
  "subjects": [],
  "subject_text": "인천 난민촌 관리사무소 3층 복도\n여러 방의 출입문과 외부를 내다보는 창이 이어지는 복도. 구석에 소화기가 비치돼 있고 위층으로 향하는 계단과 연결된다.",
  "identity": "canonical",
  "scope_id": "L188",
  "scope_role": "location_interior",
  "scope_sha": "37994dc144aebf36"
 },
 "S34sh3::bgfirst_bg": {
  "input_fingerprint": "5194da41211331df",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh3__bgfirst_bg.png",
  "asset_id": "cdae0e3e-8869-4d01-8715-a1ca5773fcb2",
  "input_asset_ids": [
   "3c73b6f6-486c-4b9f-93e3-d1822cce367e",
   "ed668b3a-a7db-491e-8f0e-bd093bd7b2fd"
  ]
 },
 "S34sh3": {
  "input_fingerprint": "c24ffaf99dc14341",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door is still shut and resisting forced entry; its handle has not yet broken free. The management office remains lit. 현우: He holds the fire extinguisher for battering the door, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door is still shut and resisting forced entry; its handle has not yet broken free. The management office remains lit. 현우: He holds the fire extinguisher for battering the door, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door is still shut and resisting forced entry; its handle has not yet broken free. The management office remains lit. 현우: He holds the fire extinguisher for battering the door, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh3__bgfirst_bg.png",
     "asset_id": "cdae0e3e-8869-4d01-8715-a1ca5773fcb2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S34sh3.png",
     "asset_id": "3c73b6f6-486c-4b9f-93e3-d1822cce367e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B02.png",
     "asset_id": "ed668b3a-a7db-491e-8f0e-bd093bd7b2fd",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "소화기의 끝부분이 문의 잠금장치 위치를 향해 강하게 타격을 가하고 있으며, 현우의 시선 역시 충돌 지점을 향하고 있음.",
    "built_space": "3층 복도의 위치로, 왼쪽에 창문이 있고 오른쪽에 문이 있음. 레퍼런스 사진에 있는 문 오른쪽 벽면의 붉은색 소화기 함이 정확한 위치에 구현됨.",
    "entities": "현우(얼굴과 헤어스타일은 일치하나, 겉옷 없음 지시를 상의 완전히 탈의한 것으로 묘사함), 붉은색 소화기. 프롬프트가 요구한 돌출된 문고리는 소화기 밸브와 스파크 이펙트에 가려져 명확히 보이지 않음.",
    "hard_violations": [
     "[gemini-pro] 소화기 밸브 쪽에 손가락과 손톱이 뚜렷한 정체불명의 세 번째 손이 묘사됨 (Extra body parts)"
    ],
    "physics": "두 다리로 바닥을 딛고 체중을 실어 소화기를 휘두르는 역동적인 자세임. 본인의 왼손은 소화기 몸통을, 오른손은 바닥면을 쥐고 지탱함. 그러나 충돌 지점의 밸브에 소유자를 알 수 없는 손이 불가능하게 나타나 있음."
   },
   {
    "label": "B",
    "direction": "소화기가 문고리를 향해 닿아 있으며, 현우의 시선은 문고리 쪽을 향함.",
    "built_space": "3층 복도의 공간. 왼쪽에 창문과 오른쪽에 문이 있으나, 레퍼런스 사진상 문 오른쪽에 반드시 있어야 할 붉은색 소화기 함이 완전히 누락되어 빈 벽으로 묘사됨.",
    "entities": "현우(얼굴과 회색 셔츠 복장이 레퍼런스와 일치함), 붉은색 소화기, 돌출된 은색 문고리.",
    "hard_violations": [
     "[gemini-pro] 오른손에 여러 개의 손가락이 겹쳐져 비정상적으로 많이 묘사됨 (Physically impossible anatomy)",
     "[gemini-pro] 소화기의 검은색 고무 호스가 은색 문고리 끝부분과 물리적으로 연결되어 하나의 물체처럼 융합됨 (Physically impossible object)"
    ],
    "physics": "왼손은 소화기의 목 부분을 쥐고, 오른손은 몸통 아래를 받치고 서 있으나 역동적인 타격 자세라기보다는 멈춰 있는 것에 가까움. 손가락의 기형과 호스의 융합이 물리 법칙에 완전히 위배됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "역동적인 타격 자세와 배경의 소화기 함 등은 잘 구현되었으나, 겉옷 없음을 상의 탈의로 묘사하였고 소화기 밸브에 세 번째 손이 나타나는 치명적인 오류가 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "캐릭터의 셔츠 복장은 레퍼런스와 일치하나, 소화기 호스가 문고리와 융합되어 있고 오른손의 형태가 기형적이며 배경의 붉은 소화기 함이 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "소화기의 끝부분이 문의 잠금장치 위치를 향해 강하게 타격을 가하고 있으며, 현우의 시선 역시 충돌 지점을 향하고 있음.",
        "built_space": "3층 복도의 위치로, 왼쪽에 창문이 있고 오른쪽에 문이 있음. 레퍼런스 사진에 있는 문 오른쪽 벽면의 붉은색 소화기 함이 정확한 위치에 구현됨.",
        "entities": "현우(얼굴과 헤어스타일은 일치하나, 겉옷 없음 지시를 상의 완전히 탈의한 것으로 묘사함), 붉은색 소화기. 프롬프트가 요구한 돌출된 문고리는 소화기 밸브와 스파크 이펙트에 가려져 명확히 보이지 않음.",
        "hard_violations": [
         "소화기 밸브 쪽에 손가락과 손톱이 뚜렷한 정체불명의 세 번째 손이 묘사됨 (Extra body parts)"
        ],
        "physics": "두 다리로 바닥을 딛고 체중을 실어 소화기를 휘두르는 역동적인 자세임. 본인의 왼손은 소화기 몸통을, 오른손은 바닥면을 쥐고 지탱함. 그러나 충돌 지점의 밸브에 소유자를 알 수 없는 손이 불가능하게 나타나 있음."
       },
       {
        "label": "B",
        "direction": "소화기가 문고리를 향해 닿아 있으며, 현우의 시선은 문고리 쪽을 향함.",
        "built_space": "3층 복도의 공간. 왼쪽에 창문과 오른쪽에 문이 있으나, 레퍼런스 사진상 문 오른쪽에 반드시 있어야 할 붉은색 소화기 함이 완전히 누락되어 빈 벽으로 묘사됨.",
        "entities": "현우(얼굴과 회색 셔츠 복장이 레퍼런스와 일치함), 붉은색 소화기, 돌출된 은색 문고리.",
        "hard_violations": [
         "오른손에 여러 개의 손가락이 겹쳐져 비정상적으로 많이 묘사됨 (Physically impossible anatomy)",
         "소화기의 검은색 고무 호스가 은색 문고리 끝부분과 물리적으로 연결되어 하나의 물체처럼 융합됨 (Physically impossible object)"
        ],
        "physics": "왼손은 소화기의 목 부분을 쥐고, 오른손은 몸통 아래를 받치고 서 있으나 역동적인 타격 자세라기보다는 멈춰 있는 것에 가까움. 손가락의 기형과 호스의 융합이 물리 법칙에 완전히 위배됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "역동적인 타격 자세와 배경의 소화기 함 등은 잘 구현되었으나, 겉옷 없음을 상의 탈의로 묘사하였고 소화기 밸브에 세 번째 손이 나타나는 치명적인 오류가 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "캐릭터의 셔츠 복장은 레퍼런스와 일치하나, 소화기 호스가 문고리와 융합되어 있고 오른손의 형태가 기형적이며 배경의 붉은 소화기 함이 누락되었습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "소화기의 끝부분이 문의 잠금장치 위치를 향해 강하게 타격을 가하고 있으며, 현우의 시선 역시 충돌 지점을 향하고 있음.",
        "built_space": "3층 복도의 위치로, 왼쪽에 창문이 있고 오른쪽에 문이 있음. 레퍼런스 사진에 있는 문 오른쪽 벽면의 붉은색 소화기 함이 정확한 위치에 구현됨.",
        "entities": "현우(얼굴과 헤어스타일은 일치하나, 겉옷 없음 지시를 상의 완전히 탈의한 것으로 묘사함), 붉은색 소화기. 프롬프트가 요구한 돌출된 문고리는 소화기 밸브와 스파크 이펙트에 가려져 명확히 보이지 않음.",
        "hard_violations": [
         "소화기 밸브 쪽에 손가락과 손톱이 뚜렷한 정체불명의 세 번째 손이 묘사됨 (Extra body parts)"
        ],
        "physics": "두 다리로 바닥을 딛고 체중을 실어 소화기를 휘두르는 역동적인 자세임. 본인의 왼손은 소화기 몸통을, 오른손은 바닥면을 쥐고 지탱함. 그러나 충돌 지점의 밸브에 소유자를 알 수 없는 손이 불가능하게 나타나 있음."
       },
       {
        "label": "B",
        "direction": "소화기가 문고리를 향해 닿아 있으며, 현우의 시선은 문고리 쪽을 향함.",
        "built_space": "3층 복도의 공간. 왼쪽에 창문과 오른쪽에 문이 있으나, 레퍼런스 사진상 문 오른쪽에 반드시 있어야 할 붉은색 소화기 함이 완전히 누락되어 빈 벽으로 묘사됨.",
        "entities": "현우(얼굴과 회색 셔츠 복장이 레퍼런스와 일치함), 붉은색 소화기, 돌출된 은색 문고리.",
        "hard_violations": [
         "오른손에 여러 개의 손가락이 겹쳐져 비정상적으로 많이 묘사됨 (Physically impossible anatomy)",
         "소화기의 검은색 고무 호스가 은색 문고리 끝부분과 물리적으로 연결되어 하나의 물체처럼 융합됨 (Physically impossible object)"
        ],
        "physics": "왼손은 소화기의 목 부분을 쥐고, 오른손은 몸통 아래를 받치고 서 있으나 역동적인 타격 자세라기보다는 멈춰 있는 것에 가까움. 손가락의 기형과 호스의 융합이 물리 법칙에 완전히 위배됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "양손의 힘이 대각선 소화기를 통해 아직 붙어 있는 문고리에 전달되는 순간이 명확하고, 회색 셔츠와 얼굴 타박상도 참조에 충실하나 미디엄 숏치고 하체가 다소 많이 보인다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "장소와 두 손으로 가격하는 동작은 잘 맞지만, 소화기와 과한 불꽃이 문고리를 가려 정확한 충돌 관계가 약해지고 참조의 회색 셔츠까지 벗긴 점도 불일치한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 오른쪽 문고리의 충돌 지점을 향한다. 소화기는 왼쪽 아래에서 오른쪽 위로 뻗으며 밸브 쪽 끝이 수평으로 돌출된 문고리 끝에 닿아 있다. 목표가 문판이 아니라 문고리라는 것을 읽을 수 있다.",
        "built_space": "오른쪽에 닫힌 회색 금속문 하나와 부착된 레버 손잡이 하나가 보인다. 문 바깥 면과 손잡이를 비스듬히 보고 있으며 현우는 복도 쪽에 서 있다. 투톤 벽, 높은 창, 천장 배관, 복도 끝 유리문과 낡은 바닥은 장소 참조와 대체로 맞는다. 참조의 붉은 벽함은 오른쪽 화면 밖이다. 충돌점은 중앙 오른쪽보다 우측 가장자리에 조금 더 치우쳐 있다.",
        "entities": "인물은 한 명이며 앳된 동아시아계 남성의 외형, 검은 머리, 마른 체격과 뺨의 타박상이 보인다. 국적은 외형만으로 확인할 수 없다. 회색 셔츠, 갈색 벨트와 올리브색 카고 바지는 인물 참조와 부합하고 별도 겉옷은 없다. 붉은 소화기 하나와 금속 문고리가 식별된다. 소화기 라벨은 흐려 확실히 읽히는 문구가 없다. 다리 부상과 신발 속 카드는 이 구도에서 확인할 수 없다.",
        "hard_violations": [],
        "physics": "한 손은 소화기의 호스 쪽을 잡고 다른 손은 통 아래를 받친다. 호스를 잡는 방식은 타격용으로 이상적이지 않지만 통을 받치는 손이 있어 소화기가 떠 있지는 않다. 벌린 다리와 앞으로 기울인 몸통이 밀어치는 힘을 전달하며 발의 접지는 화면 밖이다. 문고리는 문에 붙어 있고 문도 닫혀 있어 아직 진입을 막는 상태가 성립한다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 충돌 부위를 내려다보고, 소화기는 왼쪽 아래에서 오른쪽 위의 손잡이 부근을 향한다. 다만 소화기 끝과 밝은 불꽃이 돌출 손잡이를 상당 부분 가려, 문고리 자체를 때리는지 그 부착부를 때리는지 A보다 불명확하다.",
        "built_space": "닫힌 회색 금속문 하나, 충돌 부위의 손잡이 장치 하나, 오른쪽의 빈 붉은 매립 벽함 하나가 보인다. 높은 창, 투톤 벽, 천장 배관과 복도 끝 유리문도 참조의 공간 구성을 잘 따른다. 현우는 문 바깥 복도에 있으며 문의 외측 면을 비스듬히 보는 구도다. 충돌점은 중앙 오른쪽에 있으나 요구된 돌출 손잡이의 형태가 충분히 드러나지 않는다.",
        "entities": "검은 머리와 뺨의 타박상이 있는 젊은 동아시아계 남성 한 명이 보인다. 올리브색 카고 바지와 갈색 벨트는 참조와 맞지만 상체는 맨몸으로, 참조의 회색 셔츠가 없다. 붉은 소화기 하나를 양손으로 쥐고 있다. 읽을 수 있는 글자는 보이지 않는다. 다리 부상과 신발 속 카드는 화면에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 손이 소화기 통의 앞뒤를 직접 잡아 무게를 지지한다. 굽힌 팔, 비튼 몸통과 벌린 다리가 가격 동작으로 연결되고, 발은 화면 밖이므로 접지 자체는 확인되지 않는다. 지지 없이 떠 있는 물체는 없다. 금속 충돌에서 작은 불꽃은 가능하지만 여기서는 불꽃이 크게 강조되어 손잡이의 접촉 상태와 파손 여부를 가린다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "양손의 힘이 대각선 소화기를 통해 아직 붙어 있는 문고리에 전달되는 순간이 명확하고, 회색 셔츠와 얼굴 타박상도 참조에 충실하나 미디엄 숏치고 하체가 다소 많이 보인다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "장소와 두 손으로 가격하는 동작은 잘 맞지만, 소화기와 과한 불꽃이 문고리를 가려 정확한 충돌 관계가 약해지고 참조의 회색 셔츠까지 벗긴 점도 불일치한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 오른쪽 문고리의 충돌 지점을 향한다. 소화기는 왼쪽 아래에서 오른쪽 위로 뻗으며 밸브 쪽 끝이 수평으로 돌출된 문고리 끝에 닿아 있다. 목표가 문판이 아니라 문고리라는 것을 읽을 수 있다.",
        "built_space": "오른쪽에 닫힌 회색 금속문 하나와 부착된 레버 손잡이 하나가 보인다. 문 바깥 면과 손잡이를 비스듬히 보고 있으며 현우는 복도 쪽에 서 있다. 투톤 벽, 높은 창, 천장 배관, 복도 끝 유리문과 낡은 바닥은 장소 참조와 대체로 맞는다. 참조의 붉은 벽함은 오른쪽 화면 밖이다. 충돌점은 중앙 오른쪽보다 우측 가장자리에 조금 더 치우쳐 있다.",
        "entities": "인물은 한 명이며 앳된 동아시아계 남성의 외형, 검은 머리, 마른 체격과 뺨의 타박상이 보인다. 국적은 외형만으로 확인할 수 없다. 회색 셔츠, 갈색 벨트와 올리브색 카고 바지는 인물 참조와 부합하고 별도 겉옷은 없다. 붉은 소화기 하나와 금속 문고리가 식별된다. 소화기 라벨은 흐려 확실히 읽히는 문구가 없다. 다리 부상과 신발 속 카드는 이 구도에서 확인할 수 없다.",
        "hard_violations": [],
        "physics": "한 손은 소화기의 호스 쪽을 잡고 다른 손은 통 아래를 받친다. 호스를 잡는 방식은 타격용으로 이상적이지 않지만 통을 받치는 손이 있어 소화기가 떠 있지는 않다. 벌린 다리와 앞으로 기울인 몸통이 밀어치는 힘을 전달하며 발의 접지는 화면 밖이다. 문고리는 문에 붙어 있고 문도 닫혀 있어 아직 진입을 막는 상태가 성립한다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 충돌 부위를 내려다보고, 소화기는 왼쪽 아래에서 오른쪽 위의 손잡이 부근을 향한다. 다만 소화기 끝과 밝은 불꽃이 돌출 손잡이를 상당 부분 가려, 문고리 자체를 때리는지 그 부착부를 때리는지 A보다 불명확하다.",
        "built_space": "닫힌 회색 금속문 하나, 충돌 부위의 손잡이 장치 하나, 오른쪽의 빈 붉은 매립 벽함 하나가 보인다. 높은 창, 투톤 벽, 천장 배관과 복도 끝 유리문도 참조의 공간 구성을 잘 따른다. 현우는 문 바깥 복도에 있으며 문의 외측 면을 비스듬히 보는 구도다. 충돌점은 중앙 오른쪽에 있으나 요구된 돌출 손잡이의 형태가 충분히 드러나지 않는다.",
        "entities": "검은 머리와 뺨의 타박상이 있는 젊은 동아시아계 남성 한 명이 보인다. 올리브색 카고 바지와 갈색 벨트는 참조와 맞지만 상체는 맨몸으로, 참조의 회색 셔츠가 없다. 붉은 소화기 하나를 양손으로 쥐고 있다. 읽을 수 있는 글자는 보이지 않는다. 다리 부상과 신발 속 카드는 화면에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 손이 소화기 통의 앞뒤를 직접 잡아 무게를 지지한다. 굽힌 팔, 비튼 몸통과 벌린 다리가 가격 동작으로 연결되고, 발은 화면 밖이므로 접지 자체는 확인되지 않는다. 지지 없이 떠 있는 물체는 없다. 금속 충돌에서 작은 불꽃은 가능하지만 여기서는 불꽃이 크게 강조되어 손잡이의 접촉 상태와 파손 여부를 가린다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 소화기 밸브 쪽에 손가락과 손톱이 뚜렷한 정체불명의 세 번째 손이 묘사됨 (Extra body parts)"
    ],
    "B": [
     "[gemini-pro] 오른손에 여러 개의 손가락이 겹쳐져 비정상적으로 많이 묘사됨 (Physically impossible anatomy)",
     "[gemini-pro] 소화기의 검은색 고무 호스가 은색 문고리 끝부분과 물리적으로 연결되어 하나의 물체처럼 융합됨 (Physically impossible object)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1500,
   "B": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "역동적인 타격 자세와 배경의 소화기 함 등은 잘 구현되었으나, 겉옷 없음을 상의 탈의로 묘사하였고 소화기 밸브에 세 번째 손이 나타나는 치명적인 오류가 있습니다.  ★위반: [gemini-pro] 소화기 밸브 쪽에 손가락과 손톱이 뚜렷한 정체불명의 세 번째 손이 묘사됨 (Extra body parts)"
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "캐릭터의 셔츠 복장은 레퍼런스와 일치하나, 소화기 호스가 문고리와 융합되어 있고 오른손의 형태가 기형적이며 배경의 붉은 소화기 함이 누락되었습니다.  ★위반: [gemini-pro] 오른손에 여러 개의 손가락이 겹쳐져 비정상적으로 많이 묘사됨 (Physically impossible anatomy) / [gemini-pro] 소화기의 검은색 고무 호스가 은색 문고리 끝부분과 물리적으로 연결되어 하나의 물체처럼 융합됨 (Physically impossible object)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B02.png",
    "asset_id": "ed668b3a-a7db-491e-8f0e-bd093bd7b2fd",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-4b90-7492-89d7-93ec96a6629b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh3__bgfirst_bg.png",
   "bg_asset_id": "cdae0e3e-8869-4d01-8715-a1ca5773fcb2",
   "bg_record_key": "S34sh3::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S34sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:31:26.177281+00:00",
  "fingerprint": "86d2cdf24c85e652f5510dc92fb98a640125414a8966d737b69d0f8d636ee1ac",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S34sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S34sh3_sel.png",
  "source_sha256": "7d54ac966f0f29384a157f89325f1c91c8f6998633982bf247677606a9973ab1",
  "file": "S34sh3_cine.png",
  "staged_sha256": "c721450c123617d7e7d073bdb1a3e95cd4f2766f7090f4d6fa53d32d0812c8f6",
  "latency_ms": 9968
 },
 "S34sh7::signage": {
  "fp": "42b46010b0fc40a3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S34sh7": {
  "input_fingerprint": "9c38bd54d724797b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 미연과 현우가 바닥에 무릎을 꿇은 채 서로를 빈틈없이 꽉 끌어안은 굳은 자세.\n\nLOCATION (lock): Just inside the opened detention room on the administration building's third floor, under the building's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office floor (Supporting the kneeling pair) — A narrow area is visible beneath their knees; used as Confirm their lowered posture while leaving the embrace as the primary subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the interior illumination restrained and gently modeled, allowing the contact between their faces, shoulders, and arms to remain readable without introducing a new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door handle has broken off and the third-floor room door is now open. The management office remains lit. 현우: He is inside the opened room in an embrace posture, with persistent facial bruises and an injured leg, without his outer garment. The contact card remains in his shoe. 미연: She is no longer shut behind the door and remains visibly injured, with a badly swollen face.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 미연과 현우가 바닥에 무릎을 꿇은 채 서로를 빈틈없이 꽉 끌어안은 굳은 자세.\n\nLOCATION (lock): Just inside the opened detention room on the administration building's third floor, under the building's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office floor (Supporting the kneeling pair) — A narrow area is visible beneath their knees; used as Confirm their lowered posture while leaving the embrace as the primary subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the interior illumination restrained and gently modeled, allowing the contact between their faces, shoulders, and arms to remain readable without introducing a new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door handle has broken off and the third-floor room door is now open. The management office remains lit. 현우: He is inside the opened room in an embrace posture, with persistent facial bruises and an injured leg, without his outer garment. The contact card remains in his shoe. 미연: She is no longer shut behind the door and remains visibly injured, with a badly swollen face.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 미연과 현우가 바닥에 무릎을 꿇은 채 서로를 빈틈없이 꽉 끌어안은 굳은 자세.\n\nLOCATION (lock): Just inside the opened detention room on the administration building's third floor, under the building's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office floor (Supporting the kneeling pair) — A narrow area is visible beneath their knees; used as Confirm their lowered posture while leaving the embrace as the primary subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the interior illumination restrained and gently modeled, allowing the contact between their faces, shoulders, and arms to remain readable without introducing a new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door handle has broken off and the third-floor room door is now open. The management office remains lit. 현우: He is inside the opened room in an embrace posture, with persistent facial bruises and an injured leg, without his outer garment. The contact card remains in his shoe. 미연: She is no longer shut behind the door and remains visibly injured, with a badly swollen face.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 무릎을 꿇고 서로 꽉 끌어안은 채 시선은 허공을 향해 엇갈려 있음.",
    "built_space": "이전 샷과 동일한 분할 도색된 회색/흰색 벽면의 방 내부이며, 화면 양옆 문틀 사이로 바닥이 보임.",
    "entities": "미연(체크 셔츠)과 현우(회색 셔츠)의 얼굴과 신체가 바르게 일치하나, 현우가 프롬프트의 지시와 달리 겉옷을 입고 있음.",
    "hard_violations": [
     "[gpt-high] 현우의 가슴에 참조에 없는 이름표를 추가하고 읽을 수 있는 한글 이름을 노출하여, 이미지 어디에도 판독 가능한 글자를 두지 말라는 지시를 위반했다."
    ],
    "physics": "두 사람 모두 바닥에 무릎을 대고 물리적으로 서로의 체중을 지탱하며 자연스럽게 안고 있음."
   },
   {
    "label": "B",
    "direction": "바닥에 무릎을 꿇고 서로를 단단히 끌어안고 있으며 시선이 엇갈림.",
    "built_space": "프롬프트에 명시된 부서진 문고리가 달린 열린 문이 보이며, 그 너머로 복도 공간이 구현됨.",
    "entities": "미연과 현우의 얼굴과 몸(복장 포함)이 완전히 뒤바뀌어, 미연의 얼굴이 카고 바지를 입은 근육질 몸에 붙어 있음.",
    "hard_violations": [
     "[gemini-pro] 두 인물의 머리와 신체(복장)가 서로 뒤바뀐 해부학적 및 정체성 완전 오류"
    ],
    "physics": "바닥에 무릎을 꿇고 몸을 지탱하고 있으나, 신체가 뒤섞여 물리적 설득력을 잃음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 인물의 정체성과 신체가 올바르게 짝지어졌고 이전 샷의 배경을 잘 유지했으나, 현우가 겉옷을 벗지 않고 바지 색상이 다른 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "부서진 문고리와 배경 연출은 훌륭하지만, 두 인물의 머리와 몸(복장 포함)이 완전히 뒤바뀌는 치명적인 오류가 발생하여 사용할 수 없습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 무릎을 꿇고 서로 꽉 끌어안은 채 시선은 허공을 향해 엇갈려 있음.",
        "built_space": "이전 샷과 동일한 분할 도색된 회색/흰색 벽면의 방 내부이며, 화면 양옆 문틀 사이로 바닥이 보임.",
        "entities": "미연(체크 셔츠)과 현우(회색 셔츠)의 얼굴과 신체가 바르게 일치하나, 현우가 프롬프트의 지시와 달리 겉옷을 입고 있음.",
        "hard_violations": [],
        "physics": "두 사람 모두 바닥에 무릎을 대고 물리적으로 서로의 체중을 지탱하며 자연스럽게 안고 있음."
       },
       {
        "label": "B",
        "direction": "바닥에 무릎을 꿇고 서로를 단단히 끌어안고 있으며 시선이 엇갈림.",
        "built_space": "프롬프트에 명시된 부서진 문고리가 달린 열린 문이 보이며, 그 너머로 복도 공간이 구현됨.",
        "entities": "미연과 현우의 얼굴과 몸(복장 포함)이 완전히 뒤바뀌어, 미연의 얼굴이 카고 바지를 입은 근육질 몸에 붙어 있음.",
        "hard_violations": [
         "두 인물의 머리와 신체(복장)가 서로 뒤바뀐 해부학적 및 정체성 완전 오류"
        ],
        "physics": "바닥에 무릎을 꿇고 몸을 지탱하고 있으나, 신체가 뒤섞여 물리적 설득력을 잃음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 인물의 정체성과 신체가 올바르게 짝지어졌고 이전 샷의 배경을 잘 유지했으나, 현우가 겉옷을 벗지 않고 바지 색상이 다른 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "부서진 문고리와 배경 연출은 훌륭하지만, 두 인물의 머리와 몸(복장 포함)이 완전히 뒤바뀌는 치명적인 오류가 발생하여 사용할 수 없습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 무릎을 꿇고 서로 꽉 끌어안은 채 시선은 허공을 향해 엇갈려 있음.",
        "built_space": "이전 샷과 동일한 분할 도색된 회색/흰색 벽면의 방 내부이며, 화면 양옆 문틀 사이로 바닥이 보임.",
        "entities": "미연(체크 셔츠)과 현우(회색 셔츠)의 얼굴과 신체가 바르게 일치하나, 현우가 프롬프트의 지시와 달리 겉옷을 입고 있음.",
        "hard_violations": [],
        "physics": "두 사람 모두 바닥에 무릎을 대고 물리적으로 서로의 체중을 지탱하며 자연스럽게 안고 있음."
       },
       {
        "label": "B",
        "direction": "바닥에 무릎을 꿇고 서로를 단단히 끌어안고 있으며 시선이 엇갈림.",
        "built_space": "프롬프트에 명시된 부서진 문고리가 달린 열린 문이 보이며, 그 너머로 복도 공간이 구현됨.",
        "entities": "미연과 현우의 얼굴과 몸(복장 포함)이 완전히 뒤바뀌어, 미연의 얼굴이 카고 바지를 입은 근육질 몸에 붙어 있음.",
        "hard_violations": [
         "두 인물의 머리와 신체(복장)가 서로 뒤바뀐 해부학적 및 정체성 완전 오류"
        ],
        "physics": "바닥에 무릎을 꿇고 몸을 지탱하고 있으나, 신체가 뒤섞여 물리적 설득력을 잃음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "무릎을 지탱한 밀착 포옹, 겉옷을 벗은 현우와 부상 표현은 충실하지만, 전신에 가까운 구도와 복도 중심 배경 때문에 지정된 미디엄 숏과 구금실 내부성이 약하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "열린 방 안의 무릎 꿇은 포옹은 명확하지만, 현우의 가슴에 읽히는 이름표를 추가해 문자 금지를 위반하고 바지·신발도 참조와 달라졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 미연의 어깨 너머 아래쪽을 보고, 미연은 눈을 거의 감은 채 현우의 어깨에 얼굴을 붙인다. 두 사람의 팔과 손은 서로의 어깨와 등을 향해 감겨 있으며 카메라를 응시하지 않는다. 겨누는 물체는 없다.",
        "built_space": "왼쪽에 열린 문짝 하나와 사무실 출입구가 보이고, 그 안에는 책상·서랍장·모니터가 있다. 오른쪽 뒤로 긴 복도와 천장 조명들이 이어지며 오른쪽 벽에는 스위치 하나가 있다. 회백색 이중 도장 벽과 닳은 회색 바닥은 참조와 유사하다. 두 사람은 열린 문 옆의 전경 바닥에 있지만 구금실 안쪽인지 복도 접점인지 명료하지 않다. 문에는 금속 손잡이 부품이 남아 있어 파손 상태가 확실히 판독되지는 않는다. 무릎뿐 아니라 신발까지 보여 지정된 미디엄 숏보다 넓다.",
        "entities": "인물은 중년 한국인 여성으로 보이는 미연과 앳된 한국계 남성으로 보이는 현우, 두 명뿐이다. 미연의 짧은 검은 머리와 갈색 체크 셔츠는 이전 장면을 따른다. 현우의 헝클어진 검은 머리, 얼굴 멍, 올리브색 카고 바지와 부츠는 참조 및 상태 지시와 대체로 맞으며, 상의는 민소매 속옷으로 겉옷을 벗은 상태를 표현한다. 미연의 얼굴에는 멍이 있으나 심한 부기는 약하다. 현우의 무릎에는 피 묻은 손상이 보인다. 신발 속 연락 카드는 보이지 않으며 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 접힌 무릎과 정강이를 바닥에 대고 있으며 뒤로 접힌 신발도 바닥에 닿는다. 서로의 몸을 감싼 손과 팔이 상체를 밀착시키고, 기울어진 몸통은 무릎의 지지 범위 안에 있다. 떠 있는 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 감고 현우 쪽으로 머리를 기울이며, 현우는 미연 가까이에서 아래쪽을 바라본다. 두 사람의 얼굴과 어깨가 닿고 손은 서로의 어깨와 옆구리를 감싼다. 시선이나 손이 포옹과 무관한 대상을 향하지 않는다.",
        "built_space": "카메라는 하나의 출입구를 통해 방 안을 본다. 양쪽 문틀과 왼쪽의 열린 문 가장자리·경첩이 보이며, 뒤에는 참조와 비슷한 회백색 이중 도장 벽과 걸레받이, 회색 바닥이 있다. 두 사람의 무릎은 문틀보다 안쪽 바닥에 놓여 구금실 내부 배치는 명확하다. 왼쪽 문 가장자리의 둥근 금속 부품만으로 손잡이 파손 여부는 확인하기 어렵다. 발과 신발까지 포함해 포옹 중심의 미디엄 숏보다 넓게 잡았다.",
        "entities": "두 인물의 얼굴과 검은 머리는 미연과 현우의 참조에 대체로 부합한다. 미연은 이전 장면의 체크 셔츠를 입었고 얼굴에 멍이 있지만 심한 부기는 뚜렷하지 않다. 현우는 얼굴과 무릎에 부상이 있으나 회색 셔츠에 이름표가 추가되었고, 참조의 올리브색 카고 바지·부츠가 남색 바지·검은 운동화로 바뀌었다. 가슴의 밝은 이름표에는 '현우'로 읽히는 글자가 보인다. 신발 안 연락 카드는 확인되지 않는다.",
        "hard_violations": [
         "현우의 가슴에 참조에 없는 이름표를 추가하고 읽을 수 있는 한글 이름을 노출하여, 이미지 어디에도 판독 가능한 글자를 두지 말라는 지시를 위반했다."
        ],
        "physics": "두 사람의 무릎과 접힌 정강이가 바닥에 닿고 신발 끝도 바닥의 지지를 받는다. 현우의 손은 미연의 어깨와 옆구리에, 미연의 보이는 손은 현우의 어깨에 놓여 실제로 가능한 포옹을 이룬다. 공중에 뜬 인물이나 지지 없는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "무릎을 지탱한 밀착 포옹, 겉옷을 벗은 현우와 부상 표현은 충실하지만, 전신에 가까운 구도와 복도 중심 배경 때문에 지정된 미디엄 숏과 구금실 내부성이 약하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "열린 방 안의 무릎 꿇은 포옹은 명확하지만, 현우의 가슴에 읽히는 이름표를 추가해 문자 금지를 위반하고 바지·신발도 참조와 달라졌다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 미연의 어깨 너머 아래쪽을 보고, 미연은 눈을 거의 감은 채 현우의 어깨에 얼굴을 붙인다. 두 사람의 팔과 손은 서로의 어깨와 등을 향해 감겨 있으며 카메라를 응시하지 않는다. 겨누는 물체는 없다.",
        "built_space": "왼쪽에 열린 문짝 하나와 사무실 출입구가 보이고, 그 안에는 책상·서랍장·모니터가 있다. 오른쪽 뒤로 긴 복도와 천장 조명들이 이어지며 오른쪽 벽에는 스위치 하나가 있다. 회백색 이중 도장 벽과 닳은 회색 바닥은 참조와 유사하다. 두 사람은 열린 문 옆의 전경 바닥에 있지만 구금실 안쪽인지 복도 접점인지 명료하지 않다. 문에는 금속 손잡이 부품이 남아 있어 파손 상태가 확실히 판독되지는 않는다. 무릎뿐 아니라 신발까지 보여 지정된 미디엄 숏보다 넓다.",
        "entities": "인물은 중년 한국인 여성으로 보이는 미연과 앳된 한국계 남성으로 보이는 현우, 두 명뿐이다. 미연의 짧은 검은 머리와 갈색 체크 셔츠는 이전 장면을 따른다. 현우의 헝클어진 검은 머리, 얼굴 멍, 올리브색 카고 바지와 부츠는 참조 및 상태 지시와 대체로 맞으며, 상의는 민소매 속옷으로 겉옷을 벗은 상태를 표현한다. 미연의 얼굴에는 멍이 있으나 심한 부기는 약하다. 현우의 무릎에는 피 묻은 손상이 보인다. 신발 속 연락 카드는 보이지 않으며 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 접힌 무릎과 정강이를 바닥에 대고 있으며 뒤로 접힌 신발도 바닥에 닿는다. 서로의 몸을 감싼 손과 팔이 상체를 밀착시키고, 기울어진 몸통은 무릎의 지지 범위 안에 있다. 떠 있는 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "미연은 눈을 감고 현우 쪽으로 머리를 기울이며, 현우는 미연 가까이에서 아래쪽을 바라본다. 두 사람의 얼굴과 어깨가 닿고 손은 서로의 어깨와 옆구리를 감싼다. 시선이나 손이 포옹과 무관한 대상을 향하지 않는다.",
        "built_space": "카메라는 하나의 출입구를 통해 방 안을 본다. 양쪽 문틀과 왼쪽의 열린 문 가장자리·경첩이 보이며, 뒤에는 참조와 비슷한 회백색 이중 도장 벽과 걸레받이, 회색 바닥이 있다. 두 사람의 무릎은 문틀보다 안쪽 바닥에 놓여 구금실 내부 배치는 명확하다. 왼쪽 문 가장자리의 둥근 금속 부품만으로 손잡이 파손 여부는 확인하기 어렵다. 발과 신발까지 포함해 포옹 중심의 미디엄 숏보다 넓게 잡았다.",
        "entities": "두 인물의 얼굴과 검은 머리는 미연과 현우의 참조에 대체로 부합한다. 미연은 이전 장면의 체크 셔츠를 입었고 얼굴에 멍이 있지만 심한 부기는 뚜렷하지 않다. 현우는 얼굴과 무릎에 부상이 있으나 회색 셔츠에 이름표가 추가되었고, 참조의 올리브색 카고 바지·부츠가 남색 바지·검은 운동화로 바뀌었다. 가슴의 밝은 이름표에는 '현우'로 읽히는 글자가 보인다. 신발 안 연락 카드는 확인되지 않는다.",
        "hard_violations": [
         "현우의 가슴에 참조에 없는 이름표를 추가하고 읽을 수 있는 한글 이름을 노출하여, 이미지 어디에도 판독 가능한 글자를 두지 말라는 지시를 위반했다."
        ],
        "physics": "두 사람의 무릎과 접힌 정강이가 바닥에 닿고 신발 끝도 바닥의 지지를 받는다. 현우의 손은 미연의 어깨와 옆구리에, 미연의 보이는 손은 현우의 어깨에 놓여 실제로 가능한 포옹을 이룬다. 공중에 뜬 인물이나 지지 없는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 두 인물의 머리와 신체(복장)가 서로 뒤바뀐 해부학적 및 정체성 완전 오류"
    ],
    "A": [
     "[gpt-high] 현우의 가슴에 참조에 없는 이름표를 추가하고 읽을 수 있는 한글 이름을 노출하여, 이미지 어디에도 판독 가능한 글자를 두지 말라는 지시를 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "두 인물의 정체성과 신체가 올바르게 짝지어졌고 이전 샷의 배경을 잘 유지했으나, 현우가 겉옷을 벗지 않고 바지 색상이 다른 점이 감점 요인입니다.  ★위반: [gpt-high] 현우의 가슴에 참조에 없는 이름표를 추가하고 읽을 수 있는 한글 이름을 노출하여, 이미지 어디에도 판독 가능한 글자를 두지 말라는 지시를 위반했다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "부서진 문고리와 배경 연출은 훌륭하지만, 두 인물의 머리와 몸(복장 포함)이 완전히 뒤바뀌는 치명적인 오류가 발생하여 사용할 수 없습니다.  ★위반: [gemini-pro] 두 인물의 머리와 신체(복장)가 서로 뒤바뀐 해부학적 및 정체성 완전 오류"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh9_sel.png",
    "asset_id": "db362481-4e6e-4ffa-b9cf-823c47b364e3",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-4eda-731a-ae97-535536e21bcc",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S31sh9"
  }
 },
 "S34sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:32:41.508975+00:00",
  "fingerprint": "4f7e2a57409fab131c83cf24eb9c0239cc38bb9e8f73a83ac3a087dbe5a10785",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S34sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S34sh7_sel.png",
  "source_sha256": "278d15f7cce40a14f109b231de3e2a56b54fb8471d4887989037507e394e2479",
  "file": "S34sh7_cine.png",
  "staged_sha256": "b221eaa242e48768976f816fdccde526544a1342a6b82904bcca34dfd6808527",
  "latency_ms": 10198
 },
 "S34sh10::signage": {
  "fp": "516c5a4e894f9913",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S34sh10": {
  "input_fingerprint": "4a0eccb736e03ade",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 강렬한 붉은 사이렌 불빛이 방 안을 비추는 찰나, 천장을 올려다보며 얼어붙은 두 사람의 굳은 상체.\n\nLOCATION (lock): Inside the third-floor detention room of the refugee administration building, washed by the red emergency light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office ceiling (Above the pair as they react to the alarm) — Its underside is partially visible above their raised faces; used as Provide a real spatial destination for the upward reaction without requiring a visible alarm fixture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Capture the specified intense red alarm-light pulse across their faces and the room, preserving enough shadow detail to distinguish both reactions without showing an invented fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the room surfaces and the opened doorway with its broken handle. Exclude the earlier unalarmed lighting state; allow the emergency red light to affect the room.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door remains open with its handle broken off. The management office remains lit; no additional alarm-light color is established. 현우: He remains in the opened room with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 미연: She remains in the opened room with her badly swollen face and beating injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 강렬한 붉은 사이렌 불빛이 방 안을 비추는 찰나, 천장을 올려다보며 얼어붙은 두 사람의 굳은 상체.\n\nLOCATION (lock): Inside the third-floor detention room of the refugee administration building, washed by the red emergency light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office ceiling (Above the pair as they react to the alarm) — Its underside is partially visible above their raised faces; used as Provide a real spatial destination for the upward reaction without requiring a visible alarm fixture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Capture the specified intense red alarm-light pulse across their faces and the room, preserving enough shadow detail to distinguish both reactions without showing an invented fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the room surfaces and the opened doorway with its broken handle. Exclude the earlier unalarmed lighting state; allow the emergency red light to affect the room.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door remains open with its handle broken off. The management office remains lit; no additional alarm-light color is established. 현우: He remains in the opened room with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 미연: She remains in the opened room with her badly swollen face and beating injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 강렬한 붉은 사이렌 불빛이 방 안을 비추는 찰나, 천장을 올려다보며 얼어붙은 두 사람의 굳은 상체.\n\nLOCATION (lock): Inside the third-floor detention room of the refugee administration building, washed by the red emergency light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office ceiling (Above the pair as they react to the alarm) — Its underside is partially visible above their raised faces; used as Provide a real spatial destination for the upward reaction without requiring a visible alarm fixture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Capture the specified intense red alarm-light pulse across their faces and the room, preserving enough shadow detail to distinguish both reactions without showing an invented fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the room surfaces and the opened doorway with its broken handle. Exclude the earlier unalarmed lighting state; allow the emergency red light to affect the room.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door remains open with its handle broken off. The management office remains lit; no additional alarm-light color is established. 현우: He remains in the opened room with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 미연: She remains in the opened room with her badly swollen face and beating injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물 모두 시선을 위로 향한 채 프레임 밖의 천장 공간을 올려다보고 있음.",
    "built_space": "투톤 벽면과 노출 파이프가 있는 콘크리트 천장 구조의 방 내부. 배경에 닫힌 철문이 보이며 전체 공간이 붉은 조명으로 물들어 있음.",
    "entities": "미연은 지시된 체크무늬 셔츠와 부은 얼굴을, 현우는 레퍼런스와 동일한 회색 버튼업 셔츠와 상처 난 얼굴을 정확히 반영함.",
    "hard_violations": [
     "[gpt-high] 계속 열려 있어야 하는 출입문을 두 사람 뒤의 닫힌 문으로 표현하고 오른쪽에 별도 개구부까지 배치하여, 고정된 출입구 상태와 공간 구성을 바꿨습니다."
    ],
    "physics": "두 사람 모두 하체를 지면에 고정한 채 웅크리거나 무릎을 꿇은 안정적인 자세로 상체를 굳힌 상태임."
   },
   {
    "label": "B",
    "direction": "두 사람 모두 고개를 들어 위쪽 천장을 주시하고 있음.",
    "built_space": "문틀 너머로 보이는 방 안. 투톤 벽과 환풍기가 있는 천장이 보이며 붉은 비상등이 공간을 채우고 있음.",
    "entities": "미연은 체크 셔츠를 입고 있으나, 현우는 이전 샷에서 고정된 긴팔 셔츠 대신 민소매 티셔츠를 입고 있어 설정에 어긋남.",
    "hard_violations": [
     "[gemini-pro] 이전 샷에서 고정(locked)된 현우의 의상(회색 셔츠)을 민소매로 임의 변경함."
    ],
    "physics": "무릎을 꿇거나 낮게 앉아 바닥에 체중을 지탱한 채 얼어붙은 포즈를 유지함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 의상(현우의 회색 셔츠)과 인물 외형을 정확히 유지했으며, 지시된 미디엄 샷과 붉은 조명 아래 천장을 올려다보는 앵글을 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷과 완벽히 일치해야 하는 의상 조건(현우의 긴팔 셔츠)을 무시하고 민소매로 임의 변경하여 심각한 위반을 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 시선을 위로 향한 채 프레임 밖의 천장 공간을 올려다보고 있음.",
        "built_space": "투톤 벽면과 노출 파이프가 있는 콘크리트 천장 구조의 방 내부. 배경에 닫힌 철문이 보이며 전체 공간이 붉은 조명으로 물들어 있음.",
        "entities": "미연은 지시된 체크무늬 셔츠와 부은 얼굴을, 현우는 레퍼런스와 동일한 회색 버튼업 셔츠와 상처 난 얼굴을 정확히 반영함.",
        "hard_violations": [],
        "physics": "두 사람 모두 하체를 지면에 고정한 채 웅크리거나 무릎을 꿇은 안정적인 자세로 상체를 굳힌 상태임."
       },
       {
        "label": "B",
        "direction": "두 사람 모두 고개를 들어 위쪽 천장을 주시하고 있음.",
        "built_space": "문틀 너머로 보이는 방 안. 투톤 벽과 환풍기가 있는 천장이 보이며 붉은 비상등이 공간을 채우고 있음.",
        "entities": "미연은 체크 셔츠를 입고 있으나, 현우는 이전 샷에서 고정된 긴팔 셔츠 대신 민소매 티셔츠를 입고 있어 설정에 어긋남.",
        "hard_violations": [
         "이전 샷에서 고정(locked)된 현우의 의상(회색 셔츠)을 민소매로 임의 변경함."
        ],
        "physics": "무릎을 꿇거나 낮게 앉아 바닥에 체중을 지탱한 채 얼어붙은 포즈를 유지함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 의상(현우의 회색 셔츠)과 인물 외형을 정확히 유지했으며, 지시된 미디엄 샷과 붉은 조명 아래 천장을 올려다보는 앵글을 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷과 완벽히 일치해야 하는 의상 조건(현우의 긴팔 셔츠)을 무시하고 민소매로 임의 변경하여 심각한 위반을 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 시선을 위로 향한 채 프레임 밖의 천장 공간을 올려다보고 있음.",
        "built_space": "투톤 벽면과 노출 파이프가 있는 콘크리트 천장 구조의 방 내부. 배경에 닫힌 철문이 보이며 전체 공간이 붉은 조명으로 물들어 있음.",
        "entities": "미연은 지시된 체크무늬 셔츠와 부은 얼굴을, 현우는 레퍼런스와 동일한 회색 버튼업 셔츠와 상처 난 얼굴을 정확히 반영함.",
        "hard_violations": [],
        "physics": "두 사람 모두 하체를 지면에 고정한 채 웅크리거나 무릎을 꿇은 안정적인 자세로 상체를 굳힌 상태임."
       },
       {
        "label": "B",
        "direction": "두 사람 모두 고개를 들어 위쪽 천장을 주시하고 있음.",
        "built_space": "문틀 너머로 보이는 방 안. 투톤 벽과 환풍기가 있는 천장이 보이며 붉은 비상등이 공간을 채우고 있음.",
        "entities": "미연은 체크 셔츠를 입고 있으나, 현우는 이전 샷에서 고정된 긴팔 셔츠 대신 민소매 티셔츠를 입고 있어 설정에 어긋남.",
        "hard_violations": [
         "이전 샷에서 고정(locked)된 현우의 의상(회색 셔츠)을 민소매로 임의 변경함."
        ],
        "physics": "무릎을 꿇거나 낮게 앉아 바닥에 체중을 지탱한 채 얼어붙은 포즈를 유지함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "열린 문과 붉은 섬광 속에서 천장을 올려다보는 두 사람의 상체 구도는 더 충실하지만, 현우의 회색 셔츠가 민소매로 바뀌어 의상 연속성이 깨집니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두 사람의 위쪽 시선과 의상은 잘 맞지만, 열린 출입문을 닫힌 정면 문과 별도의 측면 개구부로 바꾼 공간 구성이 장소 고정을 위반합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 턱을 들고 카메라 위쪽을, 현우는 그보다 더 높은 오른쪽 위를 바라봅니다. 두 사람 모두 서로나 렌즈가 아니라 머리 위 공간에 반응하며, 화면 상단에 천장 아랫면이 보여 시선의 목적지가 성립합니다. 겨누는 물건이나 이동하는 물체는 없습니다.",
        "built_space": "두 사람을 둘러싼 금속 문틀 한 개와 왼쪽으로 열린 문짝 한 개가 보입니다. 밝은 상부와 회색 하부로 나뉜 벽, 문틀의 마모는 이전 장면과 대체로 이어집니다. 왼쪽 문 가장자리의 둥근 돌출부는 보이지만 손잡이가 부러진 상태인지는 명확하지 않습니다. 위에는 천장 패널, 길쭉한 조명 한 개, 사각 천장형 설비 한 개가 보이며 이전 사진에서는 이 영역을 확인할 수 없었습니다. 두 사람은 열린 출입구 안쪽에 낮게 자리하고 있습니다. 반사상이나 경광등 자체는 없습니다.",
        "entities": "인물은 미연과 현우에 해당하는 두 명뿐입니다. 미연은 중년의 동아시아계 여성으로 검은 단발과 갈색 체크 셔츠가 이전 장면에 가깝고 얼굴에 멍과 부기가 보입니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리와 얼굴 상처가 맞지만, 이전 장면의 긴소매 회색 셔츠 대신 회색 민소매를 입었습니다. 국적은 외관만으로 판별할 수 없습니다. 다리 부상과 신발 속 카드는 이 상체 구도로 확인할 수 없으며, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 아래쪽으로 자연스럽게 이어지고 현우의 굽힌 무릎이 화면 오른쪽 아래에 들어옵니다. 바닥 접점은 잘렸지만 낮게 앉거나 무릎을 세운 자세로 해석되며, 공중에 떠 있다는 징후는 없습니다. 팔은 아래로 내려가고 손은 대부분 화면 밖입니다. 천장 설비는 천장에 고정되어 있으며 지지 없는 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "미연과 현우 모두 고개를 들고 화면 오른쪽 위의 천장 방향을 바라봅니다. 특히 현우의 눈동자와 턱 방향이 위쪽 반응을 분명하게 보여 줍니다. 노출된 천장 아랫면이 실제 시선 목적지를 제공하며, 무기나 지시하는 물건은 없습니다.",
        "built_space": "두 사람 뒤 중앙에는 문틀 안에 닫힌 금속 문짝 한 개가 있고, 오른쪽 가장자리에는 별도의 문틀 또는 개구부 한 개가 더 보입니다. 중앙 문에는 손잡이 자리로 보이는 구멍과 아래쪽 금속 부속이 있습니다. 따라서 이전 장면의 열린 출입구를 유지한 배치로 읽히지 않습니다. 벽의 밝은 상부와 회색 하부는 유사하지만, 상단은 노출 콘크리트와 여러 배관, 길쭉한 조명 한 개로 구성됩니다. 두 사람은 닫힌 문 앞의 실내에 나란히 있으며 반사상은 없습니다.",
        "entities": "중년의 동아시아계 여성 한 명과 젊은 동아시아계 남성 한 명만 보입니다. 미연의 검은 머리와 체크 셔츠, 현우의 검은 머리와 긴소매 회색 셔츠가 이전 장면에 부합합니다. 현우의 볼에는 선명한 상처가 있고 미연의 얼굴에도 붉은 멍이 보이지만, 심한 부기는 상대적으로 약하게 표현됩니다. 외투나 추가 인물은 없고, 다리 부상과 신발 속 카드는 프레임 밖이라 확인 대상이 아닙니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "계속 열려 있어야 하는 출입문을 두 사람 뒤의 닫힌 문으로 표현하고 오른쪽에 별도 개구부까지 배치하여, 고정된 출입구 상태와 공간 구성을 바꿨습니다."
        ],
        "physics": "두 사람의 상체는 조금 기울어져 있지만 골반과 하체가 화면 아래로 이어져 자연스럽게 지지되는 자세로 읽힙니다. 발이나 무릎의 접점은 상체 프레이밍 밖이며, 부유하거나 불가능하게 꺾인 신체는 보이지 않습니다. 천장의 배관과 조명은 구조물에 부착되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "열린 문과 붉은 섬광 속에서 천장을 올려다보는 두 사람의 상체 구도는 더 충실하지만, 현우의 회색 셔츠가 민소매로 바뀌어 의상 연속성이 깨집니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 사람의 위쪽 시선과 의상은 잘 맞지만, 열린 출입문을 닫힌 정면 문과 별도의 측면 개구부로 바꾼 공간 구성이 장소 고정을 위반합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 턱을 들고 카메라 위쪽을, 현우는 그보다 더 높은 오른쪽 위를 바라봅니다. 두 사람 모두 서로나 렌즈가 아니라 머리 위 공간에 반응하며, 화면 상단에 천장 아랫면이 보여 시선의 목적지가 성립합니다. 겨누는 물건이나 이동하는 물체는 없습니다.",
        "built_space": "두 사람을 둘러싼 금속 문틀 한 개와 왼쪽으로 열린 문짝 한 개가 보입니다. 밝은 상부와 회색 하부로 나뉜 벽, 문틀의 마모는 이전 장면과 대체로 이어집니다. 왼쪽 문 가장자리의 둥근 돌출부는 보이지만 손잡이가 부러진 상태인지는 명확하지 않습니다. 위에는 천장 패널, 길쭉한 조명 한 개, 사각 천장형 설비 한 개가 보이며 이전 사진에서는 이 영역을 확인할 수 없었습니다. 두 사람은 열린 출입구 안쪽에 낮게 자리하고 있습니다. 반사상이나 경광등 자체는 없습니다.",
        "entities": "인물은 미연과 현우에 해당하는 두 명뿐입니다. 미연은 중년의 동아시아계 여성으로 검은 단발과 갈색 체크 셔츠가 이전 장면에 가깝고 얼굴에 멍과 부기가 보입니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리와 얼굴 상처가 맞지만, 이전 장면의 긴소매 회색 셔츠 대신 회색 민소매를 입었습니다. 국적은 외관만으로 판별할 수 없습니다. 다리 부상과 신발 속 카드는 이 상체 구도로 확인할 수 없으며, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 아래쪽으로 자연스럽게 이어지고 현우의 굽힌 무릎이 화면 오른쪽 아래에 들어옵니다. 바닥 접점은 잘렸지만 낮게 앉거나 무릎을 세운 자세로 해석되며, 공중에 떠 있다는 징후는 없습니다. 팔은 아래로 내려가고 손은 대부분 화면 밖입니다. 천장 설비는 천장에 고정되어 있으며 지지 없는 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "미연과 현우 모두 고개를 들고 화면 오른쪽 위의 천장 방향을 바라봅니다. 특히 현우의 눈동자와 턱 방향이 위쪽 반응을 분명하게 보여 줍니다. 노출된 천장 아랫면이 실제 시선 목적지를 제공하며, 무기나 지시하는 물건은 없습니다.",
        "built_space": "두 사람 뒤 중앙에는 문틀 안에 닫힌 금속 문짝 한 개가 있고, 오른쪽 가장자리에는 별도의 문틀 또는 개구부 한 개가 더 보입니다. 중앙 문에는 손잡이 자리로 보이는 구멍과 아래쪽 금속 부속이 있습니다. 따라서 이전 장면의 열린 출입구를 유지한 배치로 읽히지 않습니다. 벽의 밝은 상부와 회색 하부는 유사하지만, 상단은 노출 콘크리트와 여러 배관, 길쭉한 조명 한 개로 구성됩니다. 두 사람은 닫힌 문 앞의 실내에 나란히 있으며 반사상은 없습니다.",
        "entities": "중년의 동아시아계 여성 한 명과 젊은 동아시아계 남성 한 명만 보입니다. 미연의 검은 머리와 체크 셔츠, 현우의 검은 머리와 긴소매 회색 셔츠가 이전 장면에 부합합니다. 현우의 볼에는 선명한 상처가 있고 미연의 얼굴에도 붉은 멍이 보이지만, 심한 부기는 상대적으로 약하게 표현됩니다. 외투나 추가 인물은 없고, 다리 부상과 신발 속 카드는 프레임 밖이라 확인 대상이 아닙니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "계속 열려 있어야 하는 출입문을 두 사람 뒤의 닫힌 문으로 표현하고 오른쪽에 별도 개구부까지 배치하여, 고정된 출입구 상태와 공간 구성을 바꿨습니다."
        ],
        "physics": "두 사람의 상체는 조금 기울어져 있지만 골반과 하체가 화면 아래로 이어져 자연스럽게 지지되는 자세로 읽힙니다. 발이나 무릎의 접점은 상체 프레이밍 밖이며, 부유하거나 불가능하게 꺾인 신체는 보이지 않습니다. 천장의 배관과 조명은 구조물에 부착되어 있습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 이전 샷에서 고정(locked)된 현우의 의상(회색 셔츠)을 민소매로 임의 변경함."
    ],
    "A": [
     "[gpt-high] 계속 열려 있어야 하는 출입문을 두 사람 뒤의 닫힌 문으로 표현하고 오른쪽에 별도 개구부까지 배치하여, 고정된 출입구 상태와 공간 구성을 바꿨습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "이전 샷의 의상(현우의 회색 셔츠)과 인물 외형을 정확히 유지했으며, 지시된 미디엄 샷과 붉은 조명 아래 천장을 올려다보는 앵글을 훌륭하게 구현했습니다.  ★위반: [gpt-high] 계속 열려 있어야 하는 출입문을 두 사람 뒤의 닫힌 문으로 표현하고 오른쪽에 별도 개구부까지 배치하여, 고정된 출입구 상태와 공간 구성을 바꿨습니다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "이전 샷과 완벽히 일치해야 하는 의상 조건(현우의 긴팔 셔츠)을 무시하고 민소매로 임의 변경하여 심각한 위반을 범했습니다.  ★위반: [gemini-pro] 이전 샷에서 고정(locked)된 현우의 의상(회색 셔츠)을 민소매로 임의 변경함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh7_sel.png",
    "asset_id": "806c032c-a12d-43d6-b0ad-fc2da3556355",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-5085-7f06-840a-571a66699a6b",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S34sh7"
  }
 },
 "S34sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:33:56.123900+00:00",
  "fingerprint": "fc1e8893e43a535802baac4968cac5e4a053caa2cbb00ed3e39c25396adbccb1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S34sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S34sh10_sel.png",
  "source_sha256": "8a6aaaa20b3bb20a27357488b6d8c4678a01f840d06a6b0bafeb1bd732d28acf",
  "file": "S34sh10_cine.png",
  "staged_sha256": "bffc3a5c3b607eace3016ec6ab5964b740162fa25ac3a9042bc5029faf071a6e",
  "latency_ms": 9039
 },
 "S35sh1::signage": {
  "fp": "516d3e716cf35998",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::aaf3aac37588d4aa": {
  "subjects": [],
  "subject_text": "해일이 덮치는 인천 난민촌 판잣집 구역\n거대한 콘크리트 방벽 아래 낮은 판잣집과 낡은 임시 주거 시설이 조밀하게 모여 있는 구역. 건물 사이로 좁은 길이 이어진다.",
  "identity": "canonical",
  "scope_id": "L192",
  "scope_role": "location_exterior",
  "scope_sha": "f93d49212003ab29"
 },
 "groupbg::camp_inundation_street": {
  "input_fingerprint": "252c428074a19472",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "camp_inundation_street",
    "tags": [
     "S35sh1",
     "S35sh3"
    ]
   },
   "context_sig": "d17b8ea2372b19fa"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해일이 덮치는 인천 난민촌 판잣집 구역: 거대한 바닷물 장벽이 허술한 주거지구 구조물들을 쓸어버리는 재난 순간. (특징: 집어삼킬 듯 밀려오는 검푸른 바닷물; 부서지며 떠오르는 판잣집과 컨테이너 파편들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 바닷물이 밀려들고, 맨몸으로 대피하는 사람들! 빠르게 덮치는 쓰나미.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해일이 덮치는 인천 난민촌 판잣집 구역: 거대한 바닷물 장벽이 허술한 주거지구 구조물들을 쓸어버리는 재난 순간. (특징: 집어삼킬 듯 밀려오는 검푸른 바닷물; 부서지며 떠오르는 판잣집과 컨테이너 파편들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 바닷물이 밀려들고, 맨몸으로 대피하는 사람들! 빠르게 덮치는 쓰나미.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_inundation_street_5076fe.png",
  "asset_id": "19f46218-6a77-42f8-9095-640b2dac045d",
  "input_asset_ids": [
   "d7134b9e-61cf-401a-932e-ef0971017649"
  ],
  "origin_tag": "S35sh1",
  "place_text": "On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.",
  "origin_inputs": {
   "place_text": "On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.",
   "time_of_day_en": "night",
   "conti_asset_id": "d7134b9e-61cf-401a-932e-ef0971017649"
  }
 },
 "S35sh1::bgfirst_bg": {
  "input_fingerprint": "eb9e5b46b5acf54d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1__bgfirst_bg.png",
  "asset_id": "c4a1bc33-f30d-4b6f-ae75-66d7dbdc2b52",
  "input_asset_ids": [
   "d7134b9e-61cf-401a-932e-ef0971017649",
   "19f46218-6a77-42f8-9095-640b2dac045d"
  ]
 },
 "S35sh1": {
  "input_fingerprint": "71cf5c9fa68d83cf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall breach remains open and seawater is advancing into the nighttime settlement. The ground is beginning to shake.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall breach remains open and seawater is advancing into the nighttime settlement. The ground is beginning to shake.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall breach remains open and seawater is advancing into the nighttime settlement. The ground is beginning to shake.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1__bgfirst_bg.png",
     "asset_id": "c4a1bc33-f30d-4b6f-ae75-66d7dbdc2b52",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S35sh1.png",
     "asset_id": "d7134b9e-61cf-401a-932e-ef0971017649",
     "role": "conti_light"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_inundation_street_5076fe.png",
     "asset_id": "19f46218-6a77-42f8-9095-640b2dac045d",
     "role": "bgfirst_group_bg"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물들의 시선이 모두 허공과 파도를 향해 위로 향해 있습니다.",
    "built_space": "판자촌과 갈라진 바닥은 구현되었으나, 전경 인물들의 크기가 주변 컨테이너 건물에 비해 과도하게 큽니다.",
    "entities": "다민족 설정에 비해 인물들의 외형과 복장이 단조롭고 작위적입니다.",
    "hard_violations": [
     "[gemini-pro] 주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
     "[gemini-pro] 우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
    ],
    "physics": "발은 땅에 있으나 자세가 부자연스럽고, 현장에 없는 강한 백색광이 인물들을 비추고 있습니다."
   },
   {
    "label": "B",
    "direction": "인물들이 위쪽 하늘과 밀려오는 파도를 향해 시선을 고정하고 있습니다.",
    "built_space": "참고 이미지의 판자촌 골목과 비상등을 정확히 재현했으며, 인물들이 구조물과 올바른 비율로 배치되었습니다.",
    "entities": "지정된 대로 다양한 인종과 연령대의 사람들이 다채로운 복장을 입고 공포에 질린 표정을 짓고 있습니다.",
    "hard_violations": [],
    "physics": "인물들이 바닥을 단단히 딛거나 벽에 기대어 몸을 지탱하고 있으며, 붉은 조명이 물리적으로 자연스럽게 반사됩니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "다양한 인종의 난민들이 공포에 질려 허공을 바라보는 모습을 지정된 환경 조명과 정확한 공간 비율 내에서 사실적으로 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경 구조물 대비 인물의 크기가 비정상적으로 크며, 공간에 존재하지 않는 인위적인 조명과 해부학적 오류가 있어 사실성이 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물들의 시선이 모두 허공과 파도를 향해 위로 향해 있습니다.",
        "built_space": "판자촌과 갈라진 바닥은 구현되었으나, 전경 인물들의 크기가 주변 컨테이너 건물에 비해 과도하게 큽니다.",
        "entities": "다민족 설정에 비해 인물들의 외형과 복장이 단조롭고 작위적입니다.",
        "hard_violations": [
         "주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
         "우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
        ],
        "physics": "발은 땅에 있으나 자세가 부자연스럽고, 현장에 없는 강한 백색광이 인물들을 비추고 있습니다."
       },
       {
        "label": "B",
        "direction": "인물들이 위쪽 하늘과 밀려오는 파도를 향해 시선을 고정하고 있습니다.",
        "built_space": "참고 이미지의 판자촌 골목과 비상등을 정확히 재현했으며, 인물들이 구조물과 올바른 비율로 배치되었습니다.",
        "entities": "지정된 대로 다양한 인종과 연령대의 사람들이 다채로운 복장을 입고 공포에 질린 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "인물들이 바닥을 단단히 딛거나 벽에 기대어 몸을 지탱하고 있으며, 붉은 조명이 물리적으로 자연스럽게 반사됩니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "다양한 인종의 난민들이 공포에 질려 허공을 바라보는 모습을 지정된 환경 조명과 정확한 공간 비율 내에서 사실적으로 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경 구조물 대비 인물의 크기가 비정상적으로 크며, 공간에 존재하지 않는 인위적인 조명과 해부학적 오류가 있어 사실성이 떨어집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물들의 시선이 모두 허공과 파도를 향해 위로 향해 있습니다.",
        "built_space": "판자촌과 갈라진 바닥은 구현되었으나, 전경 인물들의 크기가 주변 컨테이너 건물에 비해 과도하게 큽니다.",
        "entities": "다민족 설정에 비해 인물들의 외형과 복장이 단조롭고 작위적입니다.",
        "hard_violations": [
         "주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
         "우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
        ],
        "physics": "발은 땅에 있으나 자세가 부자연스럽고, 현장에 없는 강한 백색광이 인물들을 비추고 있습니다."
       },
       {
        "label": "B",
        "direction": "인물들이 위쪽 하늘과 밀려오는 파도를 향해 시선을 고정하고 있습니다.",
        "built_space": "참고 이미지의 판자촌 골목과 비상등을 정확히 재현했으며, 인물들이 구조물과 올바른 비율로 배치되었습니다.",
        "entities": "지정된 대로 다양한 인종과 연령대의 사람들이 다채로운 복장을 입고 공포에 질린 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "인물들이 바닥을 단단히 딛거나 벽에 기대어 몸을 지탱하고 있으며, 붉은 조명이 물리적으로 자연스럽게 반사됩니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간의 공포와 서로 의지하는 주민들은 잘 표현했지만, 사람들이 길 양옆에 정렬되어 흔들리는 바닥과의 관계가 약하고 전경 전봇대 등 장소 배치도 참조와 차이가 난다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "참조 장소의 구조를 유지하면서 넓은 전경 바닥, 불균형하게 버티는 발, 허공을 향한 시선을 한 와이드 숏에 연결해 핵심 순간을 더 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 앞의 두 남성과 여성 무리, 오른쪽 앞의 가족은 고개를 들어 화면 위쪽의 보이지 않는 대상을 바라본다. 다만 중앙 후방의 일부 주민은 정면이나 주변 사람을 보는 듯하여 상향 시선이 군중 전체에 일관되지는 않는다. 무기나 특정 대상을 겨누는 소품은 없다.",
        "built_space": "골목 양쪽에 골판금·목재·방수포 가옥이 있고, 중앙 통로가 깊숙이 열린다. 붉은 경고등은 일곱 개가 식별되며 전선과 전봇대에 연결되어 있다. 왼쪽 전경의 굵은 전봇대에 두 남성이 기대고, 나머지 주민 대부분은 길 양옆에 모여 있다. 참조의 재료와 방벽 배경은 유지되지만 전경 전봇대와 가까운 가옥의 배치가 달라 정확한 장소 일치도는 떨어진다. 젖은 바닥의 붉은 반사는 가능한 배치다.",
        "entities": "성인 남녀와 어린이로 구성된 다인종 주민 무리가 보이며, 동아시아계로 보이는 인물이 다수이고 오른쪽에는 흑인으로 보이는 남성도 있다. 개별 국적은 확인할 수 없다. 낡은 일상복, 임시 가옥, 전선, 경고등, 갈라진 젖은 도로, 방벽 너머의 큰 파도가 보인다. 눈과 얼굴은 정상적인 인간 형태다. 전봇대에 작은 표식이 있으나 읽을 수 있는 문구는 식별되지 않는다. 방벽의 열린 파손부와 바닷물의 전진 자체는 명확히 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞쪽 주민들의 발은 노면이나 돌무더기에 닿아 있고, 왼쪽 남성들은 손으로 전봇대 또는 벽을 짚는다. 여성들은 서로 팔을 붙잡고 오른쪽 가족도 아이를 감싸며 지지한다. 뒤쪽 주민 몇 명은 무릎을 굽히고 몸을 낮춘다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다. 다만 앞쪽 여러 사람의 비교적 곧은 자세는 지면 진동에 대응하는 불균형을 약하게 전달한다."
       },
       {
        "label": "B",
        "direction": "왼쪽 앞 남성과 그 손을 잡은 아이는 화면 위쪽을 올려다본다. 중앙 남성은 얼굴 앞에 팔을 들고 위를 보며, 아기를 안은 여성과 오른쪽 남성도 허공을 향해 고개를 든다. 뒤쪽의 웅크린 주민들은 시선이 덜 분명하지만, 주요 인물들의 시선 대상은 일관되게 화면 밖 상공이다. 팔은 균형을 잡거나 얼굴을 가리는 방향으로 뻗어 있다.",
        "built_space": "양쪽 임시 가옥 사이의 넓은 골목, 왼쪽 앞 파란 드럼통 한 개, 오른쪽 수레 한 대, 머리 위 전선, 오른쪽 후방 방벽이 참조와 잘 대응한다. 붉은 경고등은 일곱 개가 보이며 전경부터 원경까지 거리감에 맞춰 작아진다. 주민들은 길 위 여러 깊이에 분산되어 있고, 발밑에서 화면 하단까지 갈라지고 젖은 노면이 넓게 이어진다. 낮은 와이드 시점에서 사람과 바닥의 관계가 명확하며 물웅덩이의 반사도 광원 위치와 양립한다.",
        "entities": "성인 다섯 명과 어린이·아기 네 명으로 보이는 주민 아홉 명이 있다. 여러 피부색의 남녀가 보이며 국적이나 언어는 이미지로 확정할 수 없다. 낡은 셔츠와 바지, 치마, 품에 안긴 아기, 임시 가옥, 경고등, 균열과 물웅덩이, 방벽 너머 파도가 확인된다. 사람들은 구체적인 얼굴과 정상적인 눈을 가진 실사 인물로 표현되어 있다. 읽을 수 있는 글자는 보이지 않는다. 열린 방벽 파손부와 유입수의 이동 방향은 명확히 확인되지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 남성은 양발을 넓게 디디고 아이의 손을 잡으며, 아이도 기울어진 몸을 지면에 닿은 발과 잡은 손으로 지탱한다. 중앙 남성은 무릎을 굽히고 발을 벌려 버티며 뒤쪽 인물들도 웅크리거나 중심을 낮춘다. 여성은 두 발로 서서 팔과 몸통으로 아기를 받친다. 오른쪽 남성 역시 벌린 발과 뻗은 팔로 균형을 잡는다. 지지 없는 부유나 불가능한 자세는 보이지 않으며, 흔들리기 시작한 지면에 반응하는 순간으로 읽힌다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "야간의 공포와 서로 의지하는 주민들은 잘 표현했지만, 사람들이 길 양옆에 정렬되어 흔들리는 바닥과의 관계가 약하고 전경 전봇대 등 장소 배치도 참조와 차이가 난다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "참조 장소의 구조를 유지하면서 넓은 전경 바닥, 불균형하게 버티는 발, 허공을 향한 시선을 한 와이드 숏에 연결해 핵심 순간을 더 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 앞의 두 남성과 여성 무리, 오른쪽 앞의 가족은 고개를 들어 화면 위쪽의 보이지 않는 대상을 바라본다. 다만 중앙 후방의 일부 주민은 정면이나 주변 사람을 보는 듯하여 상향 시선이 군중 전체에 일관되지는 않는다. 무기나 특정 대상을 겨누는 소품은 없다.",
        "built_space": "골목 양쪽에 골판금·목재·방수포 가옥이 있고, 중앙 통로가 깊숙이 열린다. 붉은 경고등은 일곱 개가 식별되며 전선과 전봇대에 연결되어 있다. 왼쪽 전경의 굵은 전봇대에 두 남성이 기대고, 나머지 주민 대부분은 길 양옆에 모여 있다. 참조의 재료와 방벽 배경은 유지되지만 전경 전봇대와 가까운 가옥의 배치가 달라 정확한 장소 일치도는 떨어진다. 젖은 바닥의 붉은 반사는 가능한 배치다.",
        "entities": "성인 남녀와 어린이로 구성된 다인종 주민 무리가 보이며, 동아시아계로 보이는 인물이 다수이고 오른쪽에는 흑인으로 보이는 남성도 있다. 개별 국적은 확인할 수 없다. 낡은 일상복, 임시 가옥, 전선, 경고등, 갈라진 젖은 도로, 방벽 너머의 큰 파도가 보인다. 눈과 얼굴은 정상적인 인간 형태다. 전봇대에 작은 표식이 있으나 읽을 수 있는 문구는 식별되지 않는다. 방벽의 열린 파손부와 바닷물의 전진 자체는 명확히 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞쪽 주민들의 발은 노면이나 돌무더기에 닿아 있고, 왼쪽 남성들은 손으로 전봇대 또는 벽을 짚는다. 여성들은 서로 팔을 붙잡고 오른쪽 가족도 아이를 감싸며 지지한다. 뒤쪽 주민 몇 명은 무릎을 굽히고 몸을 낮춘다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다. 다만 앞쪽 여러 사람의 비교적 곧은 자세는 지면 진동에 대응하는 불균형을 약하게 전달한다."
       },
       {
        "label": "A",
        "direction": "왼쪽 앞 남성과 그 손을 잡은 아이는 화면 위쪽을 올려다본다. 중앙 남성은 얼굴 앞에 팔을 들고 위를 보며, 아기를 안은 여성과 오른쪽 남성도 허공을 향해 고개를 든다. 뒤쪽의 웅크린 주민들은 시선이 덜 분명하지만, 주요 인물들의 시선 대상은 일관되게 화면 밖 상공이다. 팔은 균형을 잡거나 얼굴을 가리는 방향으로 뻗어 있다.",
        "built_space": "양쪽 임시 가옥 사이의 넓은 골목, 왼쪽 앞 파란 드럼통 한 개, 오른쪽 수레 한 대, 머리 위 전선, 오른쪽 후방 방벽이 참조와 잘 대응한다. 붉은 경고등은 일곱 개가 보이며 전경부터 원경까지 거리감에 맞춰 작아진다. 주민들은 길 위 여러 깊이에 분산되어 있고, 발밑에서 화면 하단까지 갈라지고 젖은 노면이 넓게 이어진다. 낮은 와이드 시점에서 사람과 바닥의 관계가 명확하며 물웅덩이의 반사도 광원 위치와 양립한다.",
        "entities": "성인 다섯 명과 어린이·아기 네 명으로 보이는 주민 아홉 명이 있다. 여러 피부색의 남녀가 보이며 국적이나 언어는 이미지로 확정할 수 없다. 낡은 셔츠와 바지, 치마, 품에 안긴 아기, 임시 가옥, 경고등, 균열과 물웅덩이, 방벽 너머 파도가 확인된다. 사람들은 구체적인 얼굴과 정상적인 눈을 가진 실사 인물로 표현되어 있다. 읽을 수 있는 글자는 보이지 않는다. 열린 방벽 파손부와 유입수의 이동 방향은 명확히 확인되지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 남성은 양발을 넓게 디디고 아이의 손을 잡으며, 아이도 기울어진 몸을 지면에 닿은 발과 잡은 손으로 지탱한다. 중앙 남성은 무릎을 굽히고 발을 벌려 버티며 뒤쪽 인물들도 웅크리거나 중심을 낮춘다. 여성은 두 발로 서서 팔과 몸통으로 아기를 받친다. 오른쪽 남성 역시 벌린 발과 뻗은 팔로 균형을 잡는다. 지지 없는 부유나 불가능한 자세는 보이지 않으며, 흔들리기 시작한 지면에 반응하는 순간으로 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.778
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.778
   },
   "violations": {
    "A": [
     "[gemini-pro] 주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
     "[gemini-pro] 우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1778,
   "A": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1778,
    "verdict_ko": "다양한 인종의 난민들이 공포에 질려 허공을 바라보는 모습을 지정된 환경 조명과 정확한 공간 비율 내에서 사실적으로 잘 구현했습니다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "배경 구조물 대비 인물의 크기가 비정상적으로 크며, 공간에 존재하지 않는 인위적인 조명과 해부학적 오류가 있어 사실성이 떨어집니다.  ★위반: [gemini-pro] 주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류 / [gemini-pro] 우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_inundation_street_5076fe.png",
    "asset_id": "19f46218-6a77-42f8-9095-640b2dac045d",
    "role": "bgfirst_group_bg"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-5234-741b-8494-3708ef009ca8",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1__bgfirst_bg.png",
   "bg_asset_id": "c4a1bc33-f30d-4b6f-ae75-66d7dbdc2b52",
   "bg_record_key": "S35sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "camp_inundation_street",
   "groupbg_asset_id": "19f46218-6a77-42f8-9095-640b2dac045d"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S35sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:59:42.512754+00:00",
  "fingerprint": "7664191713986e8cf0e08f7faa7bb861399605fe0c829d5ff5864dc92baa0dbd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S35sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S35sh1_sel.png",
  "source_sha256": "f872bc579aec4c4894567c2071326fbabddaeb2eeac92a6f5654c4d5dbca5cfc",
  "file": "S35sh1_cine.png",
  "staged_sha256": "dda543c4253d6a91851566eceae3ec1f26d182b924d3d798d1ae27d06223df09",
  "latency_ms": 13800
 },
 "S35sh3::signage": {
  "fp": "aa08dc17650491ac",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S35sh3": {
  "input_fingerprint": "1a43543febe4bbc9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 판잣집 지붕들 위로 솟구쳐 오른 거대한 검은 바닷물이 하늘을 완전히 뒤덮은 압도적인 찰나.\n\nLOCATION (lock): Above the roofs of the refugee settlement's shacks, where a towering nighttime surge rises over the dwellings. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Shack roofline retained as scale reference in the lower-center of the frame, midground; Towering seawater above the roofs in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Shack roofline (Below the rising tsunami) — Roof edges and partial sloping faces are seen sharply from below; used as Remain along the bottom edge as the essential scale reference; Towering seawater (Rising above the roofs and covering the sky); used as Occupy the field above the roofline as a continuous environmental threat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the black seawater within a restrained nighttime tonal range, keeping its advancing contours distinguishable from the roof silhouettes without introducing a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rapidly advancing tsunami is engulfing the settlement following the seawall breach. The flooding has not subsided.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 판잣집 지붕들 위로 솟구쳐 오른 거대한 검은 바닷물이 하늘을 완전히 뒤덮은 압도적인 찰나.\n\nLOCATION (lock): Above the roofs of the refugee settlement's shacks, where a towering nighttime surge rises over the dwellings. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Shack roofline retained as scale reference in the lower-center of the frame, midground; Towering seawater above the roofs in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Shack roofline (Below the rising tsunami) — Roof edges and partial sloping faces are seen sharply from below; used as Remain along the bottom edge as the essential scale reference; Towering seawater (Rising above the roofs and covering the sky); used as Occupy the field above the roofline as a continuous environmental threat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the black seawater within a restrained nighttime tonal range, keeping its advancing contours distinguishable from the roof silhouettes without introducing a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rapidly advancing tsunami is engulfing the settlement following the seawall breach. The flooding has not subsided.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 판잣집 지붕들 위로 솟구쳐 오른 거대한 검은 바닷물이 하늘을 완전히 뒤덮은 압도적인 찰나.\n\nLOCATION (lock): Above the roofs of the refugee settlement's shacks, where a towering nighttime surge rises over the dwellings. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Shack roofline retained as scale reference in the lower-center of the frame, midground; Towering seawater above the roofs in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Shack roofline (Below the rising tsunami) — Roof edges and partial sloping faces are seen sharply from below; used as Remain along the bottom edge as the essential scale reference; Towering seawater (Rising above the roofs and covering the sky); used as Occupy the field above the roofline as a continuous environmental threat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the black seawater within a restrained nighttime tonal range, keeping its advancing contours distinguishable from the roof silhouettes without introducing a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rapidly advancing tsunami is engulfing the settlement following the seawall breach. The flooding has not subsided.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "해당 사항 없음.",
    "built_space": "판잣촌 골목길 시점으로 양옆에 판잣집들이 있고 중앙 하단은 지붕이 아닌 골목길 바닥이 보임.",
    "entities": "거대한 검은 파도, 판잣집 외벽, 붉은 조명. 지시된 대로 인물은 등장하지 않음.",
    "hard_violations": [
     "[gemini-pro] 명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
    ],
    "physics": "거대한 파도가 골목길 끝에서 솟구쳐 오르는 모습."
   },
   {
    "label": "B",
    "direction": "해당 사항 없음 (인물 및 조준 방향 없음).",
    "built_space": "카메라가 판잣집 지붕들 위로 설정되어 프레임 하단을 지붕과 방수포들이 채우고 있으며, 레퍼런스와 일치하는 붉은 조명들이 배치됨.",
    "entities": "프레임 상단을 뒤덮은 거대한 검은 파도, 하단의 판잣집 지붕들. 지시된 대로 인물은 등장하지 않음.",
    "hard_violations": [
     "[gpt-high] 장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
    ],
    "physics": "거대한 파도가 지붕들 너머로 솟구쳐 오르는 모습이 물리적으로 타당하게 묘사됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "'지붕들 위'라는 카메라 위치와 프레임 하단에 지붕선을 배치하라는 구도 지시를 완벽하게 수행했으며, 지시된 대로 사람이 없는 상태의 위압적인 파도를 잘 묘사했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "사람을 제거하라는 지시는 따랐으나, 명시된 '지붕들 위' 시점과 하단 지붕선 구도를 무시하고 레퍼런스 이미지의 골목길 바닥 시점을 그대로 답습했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "해당 사항 없음 (인물 및 조준 방향 없음).",
        "built_space": "카메라가 판잣집 지붕들 위로 설정되어 프레임 하단을 지붕과 방수포들이 채우고 있으며, 레퍼런스와 일치하는 붉은 조명들이 배치됨.",
        "entities": "프레임 상단을 뒤덮은 거대한 검은 파도, 하단의 판잣집 지붕들. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "거대한 파도가 지붕들 너머로 솟구쳐 오르는 모습이 물리적으로 타당하게 묘사됨."
       },
       {
        "label": "A",
        "direction": "해당 사항 없음.",
        "built_space": "판잣촌 골목길 시점으로 양옆에 판잣집들이 있고 중앙 하단은 지붕이 아닌 골목길 바닥이 보임.",
        "entities": "거대한 검은 파도, 판잣집 외벽, 붉은 조명. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [
         "명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
        ],
        "physics": "거대한 파도가 골목길 끝에서 솟구쳐 오르는 모습."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "'지붕들 위'라는 카메라 위치와 프레임 하단에 지붕선을 배치하라는 구도 지시를 완벽하게 수행했으며, 지시된 대로 사람이 없는 상태의 위압적인 파도를 잘 묘사했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "사람을 제거하라는 지시는 따랐으나, 명시된 '지붕들 위' 시점과 하단 지붕선 구도를 무시하고 레퍼런스 이미지의 골목길 바닥 시점을 그대로 답습했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "해당 사항 없음 (인물 및 조준 방향 없음).",
        "built_space": "카메라가 판잣집 지붕들 위로 설정되어 프레임 하단을 지붕과 방수포들이 채우고 있으며, 레퍼런스와 일치하는 붉은 조명들이 배치됨.",
        "entities": "프레임 상단을 뒤덮은 거대한 검은 파도, 하단의 판잣집 지붕들. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "거대한 파도가 지붕들 너머로 솟구쳐 오르는 모습이 물리적으로 타당하게 묘사됨."
       },
       {
        "label": "A",
        "direction": "해당 사항 없음.",
        "built_space": "판잣촌 골목길 시점으로 양옆에 판잣집들이 있고 중앙 하단은 지붕이 아닌 골목길 바닥이 보임.",
        "entities": "거대한 검은 파도, 판잣집 외벽, 붉은 조명. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [
         "명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
        ],
        "physics": "거대한 파도가 골목길 끝에서 솟구쳐 오르는 모습."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "검은 해일의 규모는 전달하지만, 지붕을 위에서 넓게 내려다보는 구도로 지정된 아래쪽 지붕선 배치를 벗어나며 참조에 없는 위성접시들을 추가했다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "아래에서 올려다보는 지붕 처마와 참조의 골목 재질·조명을 더 충실히 유지하지만, 양옆 건물이 과도하게 높이 들어오고 해일 위에 하늘이 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "거대한 파면이 뒤쪽에서 전면의 판잣집 지붕들을 향해 밀려오는 방향으로 보인다. 사람의 시선이나 무기는 없다. 물결의 마루 위로 어두운 하늘이 남아 있어 바닷물이 하늘을 완전히 덮은 순간은 아니다.",
        "built_space": "중앙의 좁은 통로 양쪽에 골함석 판잣집들이 있고, 지붕 윗면이 화면 아래 약 40%를 차지한다. 카메라는 처마 아래가 아니라 지붕 높이 이상에서 내려다본다. 붉은 경고등 네 개와 여러 전선·기둥이 보인다. 참조의 함석과 방수포 재질은 이어지지만, 지붕을 하단 중경의 규모 기준으로만 남기는 배치와 다르다.",
        "entities": "검은 바닷물, 젖은 골함석 지붕, 방수포, 붉은 경고등은 요구된 환경과 대체로 맞는다. 사람이나 얼굴, 판독 가능한 글자는 보이지 않는다. 참조에서 확인되지 않는 위성접시 여러 개가 지붕 위의 뚜렷한 소품으로 추가되어 있다.",
        "hard_violations": [
         "장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
        ],
        "physics": "지붕은 판잣집 벽체에, 접시들은 지붕의 거치대에, 전선은 기둥에 지지되어 있다. 방수포와 무게추는 지붕 표면에 놓여 있다. 해일은 화면 아래로 이어지는 연속된 수괴이며 마루에서 포말이 날린다. 지지 없이 떠 있는 독립 물체나 신체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "뒤쪽의 거대한 물벽이 골목과 양옆 판잣집을 향해 전진하는 것으로 읽힌다. 카메라는 처마 너머 물벽을 올려다본다. 사람의 시선이나 조준 대상은 없다. 특히 왼쪽 위에 하늘이 드러나므로 하늘을 완전히 덮었다는 조건에는 미달한다.",
        "built_space": "골목 양쪽에 함석 벽체와 방수포를 두른 판잣집들이 늘어서고, 가까운 처마가 좌우에서 안쪽으로 돌출된다. 중앙 왼쪽의 큰 전신주 한 개와 골목을 따라 이어지는 작은 기둥들, 여러 붉은 경고등과 따뜻한 실내등이 보인다. 참조의 골목 구조와 낡은 재질을 더 잘 유지하며 처마를 아래에서 보는 시점도 맞는다. 다만 양옆 지붕과 벽이 화면 높이의 상당 부분을 차지해 하단 중앙에 지붕선만 남기는 구성과는 차이가 있다.",
        "entities": "거대한 검푸른 해일, 골함석 판잣집, 방수포, 전신주와 전선, 붉은 경고등이 보인다. 밤의 색조와 기존 조명의 성격이 참조에 가깝다. 식별 가능한 사람·얼굴과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "처마는 벽체와 지붕 구조에 연결되고 방수포는 건물에 걸려 있다. 전선은 기둥 사이에 연결되어 있다. 물벽은 하부가 건물 뒤에 가려진 연속된 수괴로 보이며 포말도 파면과 이어진다. 지지 없이 공중에 멈춘 물체나 신체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "검은 해일의 규모는 전달하지만, 지붕을 위에서 넓게 내려다보는 구도로 지정된 아래쪽 지붕선 배치를 벗어나며 참조에 없는 위성접시들을 추가했다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "아래에서 올려다보는 지붕 처마와 참조의 골목 재질·조명을 더 충실히 유지하지만, 양옆 건물이 과도하게 높이 들어오고 해일 위에 하늘이 남는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "거대한 파면이 뒤쪽에서 전면의 판잣집 지붕들을 향해 밀려오는 방향으로 보인다. 사람의 시선이나 무기는 없다. 물결의 마루 위로 어두운 하늘이 남아 있어 바닷물이 하늘을 완전히 덮은 순간은 아니다.",
        "built_space": "중앙의 좁은 통로 양쪽에 골함석 판잣집들이 있고, 지붕 윗면이 화면 아래 약 40%를 차지한다. 카메라는 처마 아래가 아니라 지붕 높이 이상에서 내려다본다. 붉은 경고등 네 개와 여러 전선·기둥이 보인다. 참조의 함석과 방수포 재질은 이어지지만, 지붕을 하단 중경의 규모 기준으로만 남기는 배치와 다르다.",
        "entities": "검은 바닷물, 젖은 골함석 지붕, 방수포, 붉은 경고등은 요구된 환경과 대체로 맞는다. 사람이나 얼굴, 판독 가능한 글자는 보이지 않는다. 참조에서 확인되지 않는 위성접시 여러 개가 지붕 위의 뚜렷한 소품으로 추가되어 있다.",
        "hard_violations": [
         "장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
        ],
        "physics": "지붕은 판잣집 벽체에, 접시들은 지붕의 거치대에, 전선은 기둥에 지지되어 있다. 방수포와 무게추는 지붕 표면에 놓여 있다. 해일은 화면 아래로 이어지는 연속된 수괴이며 마루에서 포말이 날린다. 지지 없이 떠 있는 독립 물체나 신체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "뒤쪽의 거대한 물벽이 골목과 양옆 판잣집을 향해 전진하는 것으로 읽힌다. 카메라는 처마 너머 물벽을 올려다본다. 사람의 시선이나 조준 대상은 없다. 특히 왼쪽 위에 하늘이 드러나므로 하늘을 완전히 덮었다는 조건에는 미달한다.",
        "built_space": "골목 양쪽에 함석 벽체와 방수포를 두른 판잣집들이 늘어서고, 가까운 처마가 좌우에서 안쪽으로 돌출된다. 중앙 왼쪽의 큰 전신주 한 개와 골목을 따라 이어지는 작은 기둥들, 여러 붉은 경고등과 따뜻한 실내등이 보인다. 참조의 골목 구조와 낡은 재질을 더 잘 유지하며 처마를 아래에서 보는 시점도 맞는다. 다만 양옆 지붕과 벽이 화면 높이의 상당 부분을 차지해 하단 중앙에 지붕선만 남기는 구성과는 차이가 있다.",
        "entities": "거대한 검푸른 해일, 골함석 판잣집, 방수포, 전신주와 전선, 붉은 경고등이 보인다. 밤의 색조와 기존 조명의 성격이 참조에 가깝다. 식별 가능한 사람·얼굴과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "처마는 벽체와 지붕 구조에 연결되고 방수포는 건물에 걸려 있다. 전선은 기둥 사이에 연결되어 있다. 물벽은 하부가 건물 뒤에 가려진 연속된 수괴로 보이며 포말도 파면과 이어진다. 지지 없이 공중에 멈춘 물체나 신체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.321
   },
   "violations": {
    "A": [
     "[gemini-pro] 명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
    ],
    "B": [
     "[gpt-high] 장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1321,
   "A": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "'지붕들 위'라는 카메라 위치와 프레임 하단에 지붕선을 배치하라는 구도 지시를 완벽하게 수행했으며, 지시된 대로 사람이 없는 상태의 위압적인 파도를 잘 묘사했습니다.  ★위반: [gpt-high] 장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "사람을 제거하라는 지시는 따랐으나, 명시된 '지붕들 위' 시점과 하단 지붕선 구도를 무시하고 레퍼런스 이미지의 골목길 바닥 시점을 그대로 답습했습니다.  ★위반: [gemini-pro] 명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1_sel.png",
    "asset_id": "5e49bfe7-5f64-4f29-81c7-b25afbfd5938",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-56d8-7b2e-b787-a8dec32c49d0",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S35sh1"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S35sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:00:39.330668+00:00",
  "fingerprint": "876a27f13ac0d9cabef4a866d5a179859b0c7a9a35443e17978253bcbc70e507",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S35sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S35sh3_sel.png",
  "source_sha256": "aafbb6d42b1474c5b617ac63045189bf2767fbe6833d98b37ce4c64004181de9",
  "file": "S35sh3_cine.png",
  "staged_sha256": "5720db55e95852943285ddf22836668f78f2a4c8e1391b89f524a537b29c188b",
  "latency_ms": 11811
 },
 "S36sh4::signage": {
  "fp": "77b6793b97b3620d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::43286f194755c581": {
  "subjects": [],
  "subject_text": "바닷물에 침수된 인천 난민촌 지하 하수도\n맨홀과 연결된 좁고 어두운 지하 배수 통로. 길게 이어지는 벽과 천장 아래로 바닷물이 가득 차 통로의 윤곽만 드러난다.",
  "identity": "canonical",
  "scope_id": "L176",
  "scope_role": "location_interior",
  "scope_sha": "3095aa32a1f7998e"
 },
 "S36sh4::bgfirst_bg": {
  "input_fingerprint": "3a741e19ed84fdf6",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh4__bgfirst_bg.png",
  "asset_id": "9ca46561-613d-4a43-8f4a-8931c01b07eb",
  "input_asset_ids": [
   "6c12b889-9578-4c38-95ec-c46ab7596741",
   "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b"
  ]
 },
 "S36sh4": {
  "input_fingerprint": "61949a541e5b27f2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Seawater bursts into the underground sewer passages and begins filling them. Charlie retains his old coat, hat and radio as the flood reaches the escape route. 구도환: He is running through the sewer escape route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Seawater bursts into the underground sewer passages and begins filling them. Charlie retains his old coat, hat and radio as the flood reaches the escape route. 구도환: He is running through the sewer escape route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Seawater bursts into the underground sewer passages and begins filling them. Charlie retains his old coat, hat and radio as the flood reaches the escape route. 구도환: He is running through the sewer escape route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh4__bgfirst_bg.png",
     "asset_id": "9ca46561-613d-4a43-8f4a-8931c01b07eb",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S36sh4.png",
     "asset_id": "6c12b889-9578-4c38-95ec-c46ab7596741",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1023014>",
     "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
     "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1023014>",
     "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물은 정면을 향해 달리며 카메라 렌즈 쪽을 바라봄.",
    "built_space": "우측에 사다리가 있는 하수구 터널. 인물이 화면 정중앙에 서 있어 터널 중심부를 가림.",
    "entities": "기준과 다르게 모자를 쓰고 무전기를 든 구도환, 등 뒤의 바닷물.",
    "hard_violations": [
     "[gemini-pro] 인물을 중앙 우측에 배치하여 터널 중앙 통로를 비우라는 프레이밍 지시를 무시하고 정중앙에 배치함",
     "[gpt-high] 인물을 지정된 오른쪽 위치가 아니라 중앙 전경에 크게 세워, 반드시 열려 있어야 할 중앙 통로와 홍수 분출 지점을 몸으로 가렸다.",
     "[gpt-high] 이 장면의 허용 인물인 구도환에게 지정되지 않은 모자와 무전기를 추가하여 찰리의 소품을 혼입했다."
    ],
    "physics": "한 발로 물을 딛고 뛰는 자연스러운 달리기 자세."
   },
   {
    "label": "B",
    "direction": "인물은 화면 좌측 전방을 향해 달리며 시선도 진행 방향을 향함.",
    "built_space": "우측에 사다리가 있는 하수구 터널. 인물이 우측에 배치되어 터널 중앙 해일이 막힘없이 노출됨.",
    "entities": "기준 이미지의 얼굴, 헤어, 의상(밧줄 벨트 등)과 일치하는 구도환, 배경 중앙에서 밀려오는 바닷물.",
    "hard_violations": [
     "[gpt-high] 장소 참조와 지시에 없는 상자·통·판재 등의 잡동사니 더미를 양쪽 전경에 대량으로 만들어 넣었다."
    ],
    "physics": "한 발이 바닥에 닿아 체중을 지탱하는 달리기 자세."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "터널 중앙을 비우라는 구도 지시를 어기고 인물을 정중앙에 배치했으며, 기준에 없는 모자와 무전기를 추가해 감점됨."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물을 우측에 배치해 터널 중앙의 해일을 시원하게 보여주며, 기준 이미지의 얼굴과 의상 디테일을 정확히 반영함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 정면을 향해 달리며 카메라 렌즈 쪽을 바라봄.",
        "built_space": "우측에 사다리가 있는 하수구 터널. 인물이 화면 정중앙에 서 있어 터널 중심부를 가림.",
        "entities": "기준과 다르게 모자를 쓰고 무전기를 든 구도환, 등 뒤의 바닷물.",
        "hard_violations": [
         "인물을 중앙 우측에 배치하여 터널 중앙 통로를 비우라는 프레이밍 지시를 무시하고 정중앙에 배치함"
        ],
        "physics": "한 발로 물을 딛고 뛰는 자연스러운 달리기 자세."
       },
       {
        "label": "B",
        "direction": "인물은 화면 좌측 전방을 향해 달리며 시선도 진행 방향을 향함.",
        "built_space": "우측에 사다리가 있는 하수구 터널. 인물이 우측에 배치되어 터널 중앙 해일이 막힘없이 노출됨.",
        "entities": "기준 이미지의 얼굴, 헤어, 의상(밧줄 벨트 등)과 일치하는 구도환, 배경 중앙에서 밀려오는 바닷물.",
        "hard_violations": [],
        "physics": "한 발이 바닥에 닿아 체중을 지탱하는 달리기 자세."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "터널 중앙을 비우라는 구도 지시를 어기고 인물을 정중앙에 배치했으며, 기준에 없는 모자와 무전기를 추가해 감점됨."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물을 우측에 배치해 터널 중앙의 해일을 시원하게 보여주며, 기준 이미지의 얼굴과 의상 디테일을 정확히 반영함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 정면을 향해 달리며 카메라 렌즈 쪽을 바라봄.",
        "built_space": "우측에 사다리가 있는 하수구 터널. 인물이 화면 정중앙에 서 있어 터널 중심부를 가림.",
        "entities": "기준과 다르게 모자를 쓰고 무전기를 든 구도환, 등 뒤의 바닷물.",
        "hard_violations": [
         "인물을 중앙 우측에 배치하여 터널 중앙 통로를 비우라는 프레이밍 지시를 무시하고 정중앙에 배치함"
        ],
        "physics": "한 발로 물을 딛고 뛰는 자연스러운 달리기 자세."
       },
       {
        "label": "B",
        "direction": "인물은 화면 좌측 전방을 향해 달리며 시선도 진행 방향을 향함.",
        "built_space": "우측에 사다리가 있는 하수구 터널. 인물이 우측에 배치되어 터널 중앙 해일이 막힘없이 노출됨.",
        "entities": "기준 이미지의 얼굴, 헤어, 의상(밧줄 벨트 등)과 일치하는 구도환, 배경 중앙에서 밀려오는 바닷물.",
        "hard_violations": [],
        "physics": "한 발이 바닥에 닿아 체중을 지탱하는 달리기 자세."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "오른쪽의 구도환 뒤로 중앙 통로와 폭발하는 물을 드러내고 인물도 잘 맞지만, 원본 장소에 없는 전경의 상자·용기·잡동사니를 대량 추가했다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "구도환을 중앙에 크게 배치해 홍수 분출 지점을 가렸으며, 참조와 다른 복장에 찰리의 모자·무전기까지 혼입했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "구도환은 화면 오른쪽에서 카메라 쪽 전경으로 달리며 시선은 전방 탈출 방향을 향한다. 물은 그의 뒤쪽 중앙 터널 끝에서 솟구쳐 중앙과 오른쪽 바닥을 따라 전경으로 밀려온다. 몸통이 주요 분출 지점을 가리지 않는다.",
        "built_space": "낮고 평평한 콘크리트 천장, 왼쪽 측면 개구부 하나, 뒤쪽 중앙 통로 하나, 오른쪽 사다리 하나와 그 위 원형 천장 구멍 하나가 보인다. 오른쪽 측면 개구부 부근은 인물과 물에 가려져 전체 윤곽을 확인하기 어렵다. 주요 구조와 오염된 재질은 참조에 가깝고 인물은 사다리 앞쪽 통로에서 달린다. 다만 양쪽 전경에 참조에 없는 상자와 용기 더미가 추가되어 장소를 바꾸었다.",
        "entities": "사람은 한 명이며 검은 머리의 중년 동아시아계 남성으로, 구도환 참조의 얼굴과 체격에 가깝다. 털 안감이 있는 낡은 짙은 녹색 외투, 검은 목도리, 허리의 끈, 바지와 부츠가 참조와 잘 맞는다. 뒤쪽의 거대한 포말과 바닥의 물은 실제 액체처럼 보이나 염분 자체는 영상으로 확인할 수 없다. 찰리나 다른 인물은 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "장소 참조와 지시에 없는 상자·통·판재 등의 잡동사니 더미를 양쪽 전경에 대량으로 만들어 넣었다."
        ],
        "physics": "앞쪽 부츠는 젖은 바닥에 닿아 체중을 받고, 반대쪽 다리는 뒤로 접혀 달리기의 추진 동작을 이룬다. 팔의 교차 움직임도 달리기와 맞으며 부유하는 인체는 없다. 분출한 물과 물방울은 뒤쪽 수류의 충격으로 상승하고 바닥의 흐름으로 이어진다. 잡동사니는 바닥에 놓이거나 물에 부분적으로 잠겨 있어 지지 없는 부유물로 보이지는 않는다."
       },
       {
        "label": "B",
        "direction": "남성은 화면 중앙에서 거의 카메라 정면을 향해 달리고 전방을 바라본다. 물도 뒤쪽에서 카메라 방향으로 흐르지만, 인물의 몸통과 벌어진 외투가 중앙 터널 끝의 분출을 상당 부분 가린다. 손에 쥔 무전기는 가슴 앞에 있으며, 화면을 읽는 동작은 아니다.",
        "built_space": "평평한 콘크리트 천장, 좌우 측면 개구부 각 하나, 중앙 후방 통로 하나, 오른쪽 벽 사다리 하나, 그 위 원형 천장 구멍 하나가 보인다. 고정 구조의 종류와 수량, 젖고 얼룩진 콘크리트 재질은 참조에 가깝다. 그러나 인물이 중앙 전경을 크게 차지해 요구된 오른쪽 배치와 비어 있는 중앙 홍수 통로가 성립하지 않는다.",
        "entities": "인물은 한 명이고 중년 동아시아계 남성으로 보이지만, 참조 얼굴과의 일치는 A보다 낮으며 머리는 모자에 가려져 있다. 외투에는 참조의 털 안감 후드와 허리끈이 보이지 않고, 검은 목도리 대신 목이 드러난 상의를 입었다. 모자와 가슴 앞 무전기는 구도환 참조에 없으며 찰리에게 지정된 소품에 해당한다. 물과 콘크리트는 실물 재질로 읽히고, 명확히 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "인물을 지정된 오른쪽 위치가 아니라 중앙 전경에 크게 세워, 반드시 열려 있어야 할 중앙 통로와 홍수 분출 지점을 몸으로 가렸다.",
         "이 장면의 허용 인물인 구도환에게 지정되지 않은 모자와 무전기를 추가하여 찰리의 소품을 혼입했다."
        ],
        "physics": "앞쪽 발은 수면을 뚫고 바닥을 딛는 위치에 있고 주변에 물이 튄다. 반대쪽 무릎과 발은 뒤로 접혀 있어 달리는 보폭으로 설명된다. 몸 전체가 근거 없이 떠 있는 상태는 아니다. 무전기는 손으로 잡고 있으며 외투 자락은 달리기에 따라 벌어진다. 물은 후방과 측면 통로에서 유입되어 전경으로 이어지므로 흐름의 물리적 연결은 보인다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "오른쪽의 구도환 뒤로 중앙 통로와 폭발하는 물을 드러내고 인물도 잘 맞지만, 원본 장소에 없는 전경의 상자·용기·잡동사니를 대량 추가했다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "구도환을 중앙에 크게 배치해 홍수 분출 지점을 가렸으며, 참조와 다른 복장에 찰리의 모자·무전기까지 혼입했다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "구도환은 화면 오른쪽에서 카메라 쪽 전경으로 달리며 시선은 전방 탈출 방향을 향한다. 물은 그의 뒤쪽 중앙 터널 끝에서 솟구쳐 중앙과 오른쪽 바닥을 따라 전경으로 밀려온다. 몸통이 주요 분출 지점을 가리지 않는다.",
        "built_space": "낮고 평평한 콘크리트 천장, 왼쪽 측면 개구부 하나, 뒤쪽 중앙 통로 하나, 오른쪽 사다리 하나와 그 위 원형 천장 구멍 하나가 보인다. 오른쪽 측면 개구부 부근은 인물과 물에 가려져 전체 윤곽을 확인하기 어렵다. 주요 구조와 오염된 재질은 참조에 가깝고 인물은 사다리 앞쪽 통로에서 달린다. 다만 양쪽 전경에 참조에 없는 상자와 용기 더미가 추가되어 장소를 바꾸었다.",
        "entities": "사람은 한 명이며 검은 머리의 중년 동아시아계 남성으로, 구도환 참조의 얼굴과 체격에 가깝다. 털 안감이 있는 낡은 짙은 녹색 외투, 검은 목도리, 허리의 끈, 바지와 부츠가 참조와 잘 맞는다. 뒤쪽의 거대한 포말과 바닥의 물은 실제 액체처럼 보이나 염분 자체는 영상으로 확인할 수 없다. 찰리나 다른 인물은 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "장소 참조와 지시에 없는 상자·통·판재 등의 잡동사니 더미를 양쪽 전경에 대량으로 만들어 넣었다."
        ],
        "physics": "앞쪽 부츠는 젖은 바닥에 닿아 체중을 받고, 반대쪽 다리는 뒤로 접혀 달리기의 추진 동작을 이룬다. 팔의 교차 움직임도 달리기와 맞으며 부유하는 인체는 없다. 분출한 물과 물방울은 뒤쪽 수류의 충격으로 상승하고 바닥의 흐름으로 이어진다. 잡동사니는 바닥에 놓이거나 물에 부분적으로 잠겨 있어 지지 없는 부유물로 보이지는 않는다."
       },
       {
        "label": "A",
        "direction": "남성은 화면 중앙에서 거의 카메라 정면을 향해 달리고 전방을 바라본다. 물도 뒤쪽에서 카메라 방향으로 흐르지만, 인물의 몸통과 벌어진 외투가 중앙 터널 끝의 분출을 상당 부분 가린다. 손에 쥔 무전기는 가슴 앞에 있으며, 화면을 읽는 동작은 아니다.",
        "built_space": "평평한 콘크리트 천장, 좌우 측면 개구부 각 하나, 중앙 후방 통로 하나, 오른쪽 벽 사다리 하나, 그 위 원형 천장 구멍 하나가 보인다. 고정 구조의 종류와 수량, 젖고 얼룩진 콘크리트 재질은 참조에 가깝다. 그러나 인물이 중앙 전경을 크게 차지해 요구된 오른쪽 배치와 비어 있는 중앙 홍수 통로가 성립하지 않는다.",
        "entities": "인물은 한 명이고 중년 동아시아계 남성으로 보이지만, 참조 얼굴과의 일치는 A보다 낮으며 머리는 모자에 가려져 있다. 외투에는 참조의 털 안감 후드와 허리끈이 보이지 않고, 검은 목도리 대신 목이 드러난 상의를 입었다. 모자와 가슴 앞 무전기는 구도환 참조에 없으며 찰리에게 지정된 소품에 해당한다. 물과 콘크리트는 실물 재질로 읽히고, 명확히 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "인물을 지정된 오른쪽 위치가 아니라 중앙 전경에 크게 세워, 반드시 열려 있어야 할 중앙 통로와 홍수 분출 지점을 몸으로 가렸다.",
         "이 장면의 허용 인물인 구도환에게 지정되지 않은 모자와 무전기를 추가하여 찰리의 소품을 혼입했다."
        ],
        "physics": "앞쪽 발은 수면을 뚫고 바닥을 딛는 위치에 있고 주변에 물이 튄다. 반대쪽 무릎과 발은 뒤로 접혀 있어 달리는 보폭으로 설명된다. 몸 전체가 근거 없이 떠 있는 상태는 아니다. 무전기는 손으로 잡고 있으며 외투 자락은 달리기에 따라 벌어진다. 물은 후방과 측면 통로에서 유입되어 전경으로 이어지므로 흐름의 물리적 연결은 보인다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.929,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.679,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 인물을 중앙 우측에 배치하여 터널 중앙 통로를 비우라는 프레이밍 지시를 무시하고 정중앙에 배치함",
     "[gpt-high] 인물을 지정된 오른쪽 위치가 아니라 중앙 전경에 크게 세워, 반드시 열려 있어야 할 중앙 통로와 홍수 분출 지점을 몸으로 가렸다.",
     "[gpt-high] 이 장면의 허용 인물인 구도환에게 지정되지 않은 모자와 무전기를 추가하여 찰리의 소품을 혼입했다."
    ],
    "B": [
     "[gpt-high] 장소 참조와 지시에 없는 상자·통·판재 등의 잡동사니 더미를 양쪽 전경에 대량으로 만들어 넣었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 679,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 679,
    "verdict_ko": "터널 중앙을 비우라는 구도 지시를 어기고 인물을 정중앙에 배치했으며, 기준에 없는 모자와 무전기를 추가해 감점됨.  ★위반: [gemini-pro] 인물을 중앙 우측에 배치하여 터널 중앙 통로를 비우라는 프레이밍 지시를 무시하고 정중앙에 배치함 / [gpt-high] 인물을 지정된 오른쪽 위치가 아니라 중앙 전경에 크게 세워, 반드시 열려 있어야 할 중앙 통로와 홍수 분출 지점을 몸으로 가렸다. / [gpt-high] 이 장면의 허용 인물인 구도환에게 지정되지 않은 모자와 무전기를 추가하여 찰리의 소품을 혼입했다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "인물을 우측에 배치해 터널 중앙의 해일을 시원하게 보여주며, 기준 이미지의 얼굴과 의상 디테일을 정확히 반영함.  ★위반: [gpt-high] 장소 참조와 지시에 없는 상자·통·판재 등의 잡동사니 더미를 양쪽 전경에 대량으로 만들어 넣었다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
    "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1023014>",
    "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-5876-7a1e-8082-8bf6412d8a29",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh4__bgfirst_bg.png",
   "bg_asset_id": "9ca46561-613d-4a43-8f4a-8931c01b07eb",
   "bg_record_key": "S36sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S36sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:35:00.758960+00:00",
  "fingerprint": "7d2f1449c2fad57b4f0152973491b49c2ca816a2d7eb51dc6016021de5687ee8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S36sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S36sh4_sel.png",
  "source_sha256": "55a5aa702084d170e1f8ad2564540a5e52314b93720f2d295b2058ba713d1891",
  "file": "S36sh4_cine.png",
  "staged_sha256": "1ac165a066984bae52005c64702b61d278b4fc35c25e6d4262a7f64d86e35261",
  "latency_ms": 9818
 },
 "S36sh7::signage": {
  "fp": "965f03999b11a1e1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S36sh7::bgfirst_bg": {
  "input_fingerprint": "824ac1b7a3bd52bc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh7__bgfirst_bg.png",
  "asset_id": "8484c4a2-ba85-42fe-b5be-7c3cd2e4cc63",
  "input_asset_ids": [
   "e8a72824-fe1d-4ef0-b1b3-784cdbd25fdc",
   "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b"
  ]
 },
 "S36sh7": {
  "input_fingerprint": "9a1be4e79fdbc46b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer is inundated, and the flood has forced the escape into open water. Charlie is submerged and swimming, with his worn metal body and old disguise now soaked. 앰버: She has been swept out of the sewer and is submerged, struggling in the seawater with soaked clothing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer is inundated, and the flood has forced the escape into open water. Charlie is submerged and swimming, with his worn metal body and old disguise now soaked. 앰버: She has been swept out of the sewer and is submerged, struggling in the seawater with soaked clothing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer is inundated, and the flood has forced the escape into open water. Charlie is submerged and swimming, with his worn metal body and old disguise now soaked. 앰버: She has been swept out of the sewer and is submerged, struggling in the seawater with soaked clothing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh7__bgfirst_bg.png",
     "asset_id": "8484c4a2-ba85-42fe-b5be-7c3cd2e4cc63",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S36sh7.png",
     "asset_id": "e8a72824-fe1d-4ef0-b1b3-784cdbd25fdc",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
     "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "찰리의 팔은 화면 위에서 중앙 아래로 뻗고, 금속 손가락은 앰버의 손목 둘레로 굽어 있다. 앰버의 팔은 왼쪽 아래에서 오른쪽 위로 향하며, 앰버도 금속 손을 붙잡는 모양이다. 얼굴과 시선, 수영 진행 방향은 보이지 않는다.",
    "built_space": "양쪽 콘크리트 벽과 천장이 둘러싼 하수구 내부가 보인다. 측면 개구부는 좌우 하나씩, 오른쪽 사다리는 하나, 그 위 원형 천장 구멍은 하나가 보이며 중앙 후방 개구부는 팔에 일부 가려진다. 참조의 내부 구조는 따르지만 요구된 출구 밖 바닷물이 아니다. 물결치는 수면이 손 아래로 펼쳐져 손과 카메라가 물속에 있다는 공간 관계도 읽히지 않는다.",
    "entities": "보이는 인물 부분은 작은 맨손·팔과 큰 금속 손·팔뿐이다. 금속은 마모된 샌드 베이지 장갑과 관절을 갖춰 찰리와 대체로 맞는다. 작은 손은 어린이로 읽힐 수 있지만 얼굴이 없어 앰버의 성별과 혼혈 정체성은 확인할 수 없다. 파란 긴소매는 참조의 남색 반소매와 갈색 멜빵옷과 다르다. 읽을 수 있는 글자나 추가 인물은 없다.",
    "hard_violations": [
     "인물의 손을 출구 밖이 아니라 벽과 천장으로 둘러싸인 하수구 내부에 배치했다."
    ],
    "physics": "금속 손은 위쪽 전완에 연결되어 있고, 앰버의 손목은 굽힌 금속 손가락과 접촉한다. 앰버의 팔도 화면 밖 몸으로 이어져 지지 없이 분리된 물체는 없다. 다만 아래쪽에 드러난 수면 때문에 잠긴 채 수영하는 순간보다는 물 위에서 서로 붙잡는 순간으로 보인다."
   },
   {
    "label": "A",
    "direction": "찰리의 팔은 오른쪽 위 몸통에서 왼쪽 아래의 앰버 손목으로 뻗고, 금속 손가락이 손목을 둘러싼다. 앰버는 왼쪽에서 오른쪽으로 팔을 내밀며 머리도 접촉 지점 쪽으로 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없고, 두 사람의 이동 방향 역시 단정할 수 없다.",
    "built_space": "머리 위에 수면의 아랫면이 보이고 그 아래에 두 인물의 부분 신체가 놓여 있어 수중 시점이 명확하다. 그러나 좌우 콘크리트 벽, 왼쪽 측면 개구부 하나, 중앙 후방 개구부 하나, 오른쪽 사다리 하나와 위쪽 원형 구멍 하나가 하수구 내부를 형성한다. 오른쪽 측면 개구부는 찰리에게 가려져 확인하기 어렵다. 요구된 출구 밖 개방 수역은 아니다.",
    "entities": "왼쪽의 금발 어린아이 부분과 오른쪽의 육중한 베이지 금속 팔·몸통은 각각 앰버와 찰리로 읽힌다. 앰버의 얼굴 정면이 없어 큰 눈과 둥근 얼굴, 혼혈 외모의 정확한 일치는 확인할 수 없다. 갈색 긴소매는 참조의 멜빵옷 아래 남색 반소매와 다르며, 찰리의 보이는 몸통에는 참조의 낡은 코트가 보이지 않는다. 금속의 마모와 피부 접촉은 구체적이다. 여러 기포가 추가되어 기포를 넣지 말라는 조건과 다르고, 읽을 수 있는 글자는 없다.",
    "hard_violations": [
     "잠긴 두 인물을 출구 밖 바닷물이 아니라 하수구 내부에 배치했다."
    ],
    "physics": "금속 손은 전완과 몸통으로 이어지고, 손가락과 엄지가 앰버 손목을 감싸 실제로 붙잡는 접촉을 만든다. 앰버의 팔도 어깨에 연결되어 있다. 물속으로 퍼진 머리카락과 머리 위 수면은 잠긴 상태를 뒷받침한다. 부분 신체만 보여 추진 동작은 확인할 수 없지만, 수중 부력과 손목 연결로 설명되는 배치이며 근거 없이 공중에 떠 있는 신체는 아니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "손목 접촉을 크게 잡았지만 두 손이 수면 위에 있는 듯 보이며, 하수구 밖 수중이라는 핵심 공간·행동 조건을 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "큰 금속 손이 작은 손목을 감싼 수중 순간은 더 명확하지만, 하수구 내부 배치와 불필요한 기포 때문에 완전한 충족은 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔은 화면 위에서 중앙 아래로 뻗고, 금속 손가락은 앰버의 손목 둘레로 굽어 있다. 앰버의 팔은 왼쪽 아래에서 오른쪽 위로 향하며, 앰버도 금속 손을 붙잡는 모양이다. 얼굴과 시선, 수영 진행 방향은 보이지 않는다.",
        "built_space": "양쪽 콘크리트 벽과 천장이 둘러싼 하수구 내부가 보인다. 측면 개구부는 좌우 하나씩, 오른쪽 사다리는 하나, 그 위 원형 천장 구멍은 하나가 보이며 중앙 후방 개구부는 팔에 일부 가려진다. 참조의 내부 구조는 따르지만 요구된 출구 밖 바닷물이 아니다. 물결치는 수면이 손 아래로 펼쳐져 손과 카메라가 물속에 있다는 공간 관계도 읽히지 않는다.",
        "entities": "보이는 인물 부분은 작은 맨손·팔과 큰 금속 손·팔뿐이다. 금속은 마모된 샌드 베이지 장갑과 관절을 갖춰 찰리와 대체로 맞는다. 작은 손은 어린이로 읽힐 수 있지만 얼굴이 없어 앰버의 성별과 혼혈 정체성은 확인할 수 없다. 파란 긴소매는 참조의 남색 반소매와 갈색 멜빵옷과 다르다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [
         "인물의 손을 출구 밖이 아니라 벽과 천장으로 둘러싸인 하수구 내부에 배치했다."
        ],
        "physics": "금속 손은 위쪽 전완에 연결되어 있고, 앰버의 손목은 굽힌 금속 손가락과 접촉한다. 앰버의 팔도 화면 밖 몸으로 이어져 지지 없이 분리된 물체는 없다. 다만 아래쪽에 드러난 수면 때문에 잠긴 채 수영하는 순간보다는 물 위에서 서로 붙잡는 순간으로 보인다."
       },
       {
        "label": "B",
        "direction": "찰리의 팔은 오른쪽 위 몸통에서 왼쪽 아래의 앰버 손목으로 뻗고, 금속 손가락이 손목을 둘러싼다. 앰버는 왼쪽에서 오른쪽으로 팔을 내밀며 머리도 접촉 지점 쪽으로 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없고, 두 사람의 이동 방향 역시 단정할 수 없다.",
        "built_space": "머리 위에 수면의 아랫면이 보이고 그 아래에 두 인물의 부분 신체가 놓여 있어 수중 시점이 명확하다. 그러나 좌우 콘크리트 벽, 왼쪽 측면 개구부 하나, 중앙 후방 개구부 하나, 오른쪽 사다리 하나와 위쪽 원형 구멍 하나가 하수구 내부를 형성한다. 오른쪽 측면 개구부는 찰리에게 가려져 확인하기 어렵다. 요구된 출구 밖 개방 수역은 아니다.",
        "entities": "왼쪽의 금발 어린아이 부분과 오른쪽의 육중한 베이지 금속 팔·몸통은 각각 앰버와 찰리로 읽힌다. 앰버의 얼굴 정면이 없어 큰 눈과 둥근 얼굴, 혼혈 외모의 정확한 일치는 확인할 수 없다. 갈색 긴소매는 참조의 멜빵옷 아래 남색 반소매와 다르며, 찰리의 보이는 몸통에는 참조의 낡은 코트가 보이지 않는다. 금속의 마모와 피부 접촉은 구체적이다. 여러 기포가 추가되어 기포를 넣지 말라는 조건과 다르고, 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "잠긴 두 인물을 출구 밖 바닷물이 아니라 하수구 내부에 배치했다."
        ],
        "physics": "금속 손은 전완과 몸통으로 이어지고, 손가락과 엄지가 앰버 손목을 감싸 실제로 붙잡는 접촉을 만든다. 앰버의 팔도 어깨에 연결되어 있다. 물속으로 퍼진 머리카락과 머리 위 수면은 잠긴 상태를 뒷받침한다. 부분 신체만 보여 추진 동작은 확인할 수 없지만, 수중 부력과 손목 연결로 설명되는 배치이며 근거 없이 공중에 떠 있는 신체는 아니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손목 접촉을 크게 잡았지만 두 손이 수면 위에 있는 듯 보이며, 하수구 밖 수중이라는 핵심 공간·행동 조건을 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "큰 금속 손이 작은 손목을 감싼 수중 순간은 더 명확하지만, 하수구 내부 배치와 불필요한 기포 때문에 완전한 충족은 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 팔은 화면 위에서 중앙 아래로 뻗고, 금속 손가락은 앰버의 손목 둘레로 굽어 있다. 앰버의 팔은 왼쪽 아래에서 오른쪽 위로 향하며, 앰버도 금속 손을 붙잡는 모양이다. 얼굴과 시선, 수영 진행 방향은 보이지 않는다.",
        "built_space": "양쪽 콘크리트 벽과 천장이 둘러싼 하수구 내부가 보인다. 측면 개구부는 좌우 하나씩, 오른쪽 사다리는 하나, 그 위 원형 천장 구멍은 하나가 보이며 중앙 후방 개구부는 팔에 일부 가려진다. 참조의 내부 구조는 따르지만 요구된 출구 밖 바닷물이 아니다. 물결치는 수면이 손 아래로 펼쳐져 손과 카메라가 물속에 있다는 공간 관계도 읽히지 않는다.",
        "entities": "보이는 인물 부분은 작은 맨손·팔과 큰 금속 손·팔뿐이다. 금속은 마모된 샌드 베이지 장갑과 관절을 갖춰 찰리와 대체로 맞는다. 작은 손은 어린이로 읽힐 수 있지만 얼굴이 없어 앰버의 성별과 혼혈 정체성은 확인할 수 없다. 파란 긴소매는 참조의 남색 반소매와 갈색 멜빵옷과 다르다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [
         "인물의 손을 출구 밖이 아니라 벽과 천장으로 둘러싸인 하수구 내부에 배치했다."
        ],
        "physics": "금속 손은 위쪽 전완에 연결되어 있고, 앰버의 손목은 굽힌 금속 손가락과 접촉한다. 앰버의 팔도 화면 밖 몸으로 이어져 지지 없이 분리된 물체는 없다. 다만 아래쪽에 드러난 수면 때문에 잠긴 채 수영하는 순간보다는 물 위에서 서로 붙잡는 순간으로 보인다."
       },
       {
        "label": "A",
        "direction": "찰리의 팔은 오른쪽 위 몸통에서 왼쪽 아래의 앰버 손목으로 뻗고, 금속 손가락이 손목을 둘러싼다. 앰버는 왼쪽에서 오른쪽으로 팔을 내밀며 머리도 접촉 지점 쪽으로 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없고, 두 사람의 이동 방향 역시 단정할 수 없다.",
        "built_space": "머리 위에 수면의 아랫면이 보이고 그 아래에 두 인물의 부분 신체가 놓여 있어 수중 시점이 명확하다. 그러나 좌우 콘크리트 벽, 왼쪽 측면 개구부 하나, 중앙 후방 개구부 하나, 오른쪽 사다리 하나와 위쪽 원형 구멍 하나가 하수구 내부를 형성한다. 오른쪽 측면 개구부는 찰리에게 가려져 확인하기 어렵다. 요구된 출구 밖 개방 수역은 아니다.",
        "entities": "왼쪽의 금발 어린아이 부분과 오른쪽의 육중한 베이지 금속 팔·몸통은 각각 앰버와 찰리로 읽힌다. 앰버의 얼굴 정면이 없어 큰 눈과 둥근 얼굴, 혼혈 외모의 정확한 일치는 확인할 수 없다. 갈색 긴소매는 참조의 멜빵옷 아래 남색 반소매와 다르며, 찰리의 보이는 몸통에는 참조의 낡은 코트가 보이지 않는다. 금속의 마모와 피부 접촉은 구체적이다. 여러 기포가 추가되어 기포를 넣지 말라는 조건과 다르고, 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "잠긴 두 인물을 출구 밖 바닷물이 아니라 하수구 내부에 배치했다."
        ],
        "physics": "금속 손은 전완과 몸통으로 이어지고, 손가락과 엄지가 앰버 손목을 감싸 실제로 붙잡는 접촉을 만든다. 앰버의 팔도 어깨에 연결되어 있다. 물속으로 퍼진 머리카락과 머리 위 수면은 잠긴 상태를 뒷받침한다. 부분 신체만 보여 추진 동작은 확인할 수 없지만, 수중 부력과 손목 연결로 설명되는 배치이며 근거 없이 공중에 떠 있는 신체는 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 5
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "손목 접촉을 크게 잡았지만 두 손이 수면 위에 있는 듯 보이며, 하수구 밖 수중이라는 핵심 공간·행동 조건을 충족하지 못한다."
   },
   {
    "label": "A",
    "score": 5,
    "verdict_ko": "큰 금속 손이 작은 손목을 감싼 수중 순간은 더 명확하지만, 하수구 내부 배치와 불필요한 기포 때문에 완전한 충족은 아니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
    "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-5baf-732b-8952-9cbb55632578",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh7__bgfirst_bg.png",
   "bg_asset_id": "8484c4a2-ba85-42fe-b5be-7c3cd2e4cc63",
   "bg_record_key": "S36sh7::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S36sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:36:26.985562+00:00",
  "fingerprint": "f74f31f79c78a5811f81912194bc7a25d3fec142364b4b4de1fa45a406a62305",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S36sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S36sh7_sel.png",
  "source_sha256": "87f058a9ded7c4b8c0e0f11e4f847410a75293d63beea65ae22148a8276c5d02",
  "file": "S36sh7_cine.png",
  "staged_sha256": "3b1e28a7cdea9e9f236cedca1cc3362d3c998f2fe8ddcf80b4816319bf3bf90e",
  "latency_ms": 10247
 },
 "S37sh4::signage": {
  "fp": "88e9bb51cd7b6166",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S37sh4::bgfirst_bg": {
  "input_fingerprint": "92fbdfb31bc64ace",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4__bgfirst_bg.png",
  "asset_id": "6b563ad8-6d4e-4a21-9db4-4cdc611ddc68",
  "input_asset_ids": [
   "eda28974-c427-4e36-a410-4567430cdb20",
   "09d33278-e815-4961-9ee0-fa14a5350447"
  ]
 },
 "S37sh4": {
  "input_fingerprint": "ffec54c34b35d7f0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management-office corridor windows are still intact as the approaching tidal wave fills the view outside. 미연: Her face remains badly swollen from the beating. She stands with both arms closed tightly against her chest in a protective embrace. 현우: He retains facial bruising and the untreated dog-bite wound on his leg, and his outer garment remains removed. Yoon's contact card is concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management-office corridor windows are still intact as the approaching tidal wave fills the view outside. 미연: Her face remains badly swollen from the beating. She stands with both arms closed tightly against her chest in a protective embrace. 현우: He retains facial bruising and the untreated dog-bite wound on his leg, and his outer garment remains removed. Yoon's contact card is concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management-office corridor windows are still intact as the approaching tidal wave fills the view outside. 미연: Her face remains badly swollen from the beating. She stands with both arms closed tightly against her chest in a protective embrace. 현우: He retains facial bruising and the untreated dog-bite wound on his leg, and his outer garment remains removed. Yoon's contact card is concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4__bgfirst_bg.png",
     "asset_id": "6b563ad8-6d4e-4a21-9db4-4cdc611ddc68",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S37sh4.png",
     "asset_id": "eda28974-c427-4e36-a410-4567430cdb20",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B01.png",
     "asset_id": "09d33278-e815-4961-9ee0-fa14a5350447",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:943817>",
     "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연은 두 눈을 감고 현우의 머리를 가슴팍으로 끌어안고 있음.",
    "built_space": "왼쪽의 창문, 오른쪽의 소화기와 계단 등 참조 이미지의 복도 구조가 정확히 배치됨. 창밖에 파도가 들이닥침.",
    "entities": "미연의 의상과 외형은 참조와 일치하나 붓기가 미미함. 현우는 지시와 달리 겉옷(회색 셔츠)을 입고 있음.",
    "hard_violations": [],
    "physics": "미연의 팔이 현우를 감싸 안고 있으며 자세와 지탱이 안정적임."
   },
   {
    "label": "B",
    "direction": "미연이 눈을 질끈 감은 채 현우를 자신의 가슴으로 강하게 끌어안고 있음.",
    "built_space": "창문, 소화기, 복도 형태가 위치에 맞게 렌더링되었으며 창밖으로 검은 파도가 보임.",
    "entities": "미연의 얼굴에 심한 붓기와 멍이 묘사됨. 현우는 프롬프트 지시대로 겉옷을 벗은 맨몸 상태임.",
    "hard_violations": [],
    "physics": "미연의 양손이 현우의 머리와 등을 단단히 쥐고 지탱하며 물리적 오류가 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 겉옷을 벗어야 한다는 지시를 누락하고 참조 이미지의 의상을 그대로 묘사했으며, 미연의 얼굴 붓기도 다소 약하게 표현되었습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우가 겉옷을 벗은 상태와 미연의 심하게 부어오른 얼굴 등 프롬프트의 구체적인 상태 변화 지시를 정확하고 충실하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 두 눈을 감고 현우의 머리를 가슴팍으로 끌어안고 있음.",
        "built_space": "왼쪽의 창문, 오른쪽의 소화기와 계단 등 참조 이미지의 복도 구조가 정확히 배치됨. 창밖에 파도가 들이닥침.",
        "entities": "미연의 의상과 외형은 참조와 일치하나 붓기가 미미함. 현우는 지시와 달리 겉옷(회색 셔츠)을 입고 있음.",
        "hard_violations": [],
        "physics": "미연의 팔이 현우를 감싸 안고 있으며 자세와 지탱이 안정적임."
       },
       {
        "label": "B",
        "direction": "미연이 눈을 질끈 감은 채 현우를 자신의 가슴으로 강하게 끌어안고 있음.",
        "built_space": "창문, 소화기, 복도 형태가 위치에 맞게 렌더링되었으며 창밖으로 검은 파도가 보임.",
        "entities": "미연의 얼굴에 심한 붓기와 멍이 묘사됨. 현우는 프롬프트 지시대로 겉옷을 벗은 맨몸 상태임.",
        "hard_violations": [],
        "physics": "미연의 양손이 현우의 머리와 등을 단단히 쥐고 지탱하며 물리적 오류가 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 겉옷을 벗어야 한다는 지시를 누락하고 참조 이미지의 의상을 그대로 묘사했으며, 미연의 얼굴 붓기도 다소 약하게 표현되었습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우가 겉옷을 벗은 상태와 미연의 심하게 부어오른 얼굴 등 프롬프트의 구체적인 상태 변화 지시를 정확하고 충실하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 두 눈을 감고 현우의 머리를 가슴팍으로 끌어안고 있음.",
        "built_space": "왼쪽의 창문, 오른쪽의 소화기와 계단 등 참조 이미지의 복도 구조가 정확히 배치됨. 창밖에 파도가 들이닥침.",
        "entities": "미연의 의상과 외형은 참조와 일치하나 붓기가 미미함. 현우는 지시와 달리 겉옷(회색 셔츠)을 입고 있음.",
        "hard_violations": [],
        "physics": "미연의 팔이 현우를 감싸 안고 있으며 자세와 지탱이 안정적임."
       },
       {
        "label": "B",
        "direction": "미연이 눈을 질끈 감은 채 현우를 자신의 가슴으로 강하게 끌어안고 있음.",
        "built_space": "창문, 소화기, 복도 형태가 위치에 맞게 렌더링되었으며 창밖으로 검은 파도가 보임.",
        "entities": "미연의 얼굴에 심한 붓기와 멍이 묘사됨. 현우는 프롬프트 지시대로 겉옷을 벗은 맨몸 상태임.",
        "hard_violations": [],
        "physics": "미연의 양손이 현우의 머리와 등을 단단히 쥐고 지탱하며 물리적 오류가 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "눈을 질끈 감은 표정과 부상·겉옷 제거는 잘 맞지만, 현우의 얼굴이 가슴보다 미연의 팔 쪽에 기대고 인물 크기도 작아 핵심 포옹의 클로즈업이 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우의 머리를 가슴팍에 빈틈없이 묻고 두 팔로 감싼 밀착 클로즈업이 더 정확하지만, 현우에게 회색 셔츠가 남아 있어 겉옷 제거 지시와 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 두 눈을 강하게 감고 얼굴을 거의 정면으로 향한다. 현우는 눈을 감고 왼쪽 아래로 얼굴을 돌려 미연의 팔 안쪽에 뺨을 댄다. 미연의 한 손은 뒤통수를, 다른 손은 등과 어깨를 끌어안지만, 현우의 얼굴이 가슴 안쪽으로 완전히 파묻힌 방향은 아니다.",
        "built_space": "왼쪽 전경에 중앙 금속틀로 나뉜 온전한 창유리 두 면과 창턱이 보인다. 뒤쪽에는 끝 창 하나, 올라가는 계단 한 구간과 난간, 오른쪽 벽의 붉은 매립함 하나와 소화기 하나가 보인다. 낡은 투톤 벽과 바닥은 장소 참고와 부합한다. 두 사람은 창 바로 옆 복도에 있으나, 배경과 현우의 등이 상당한 면적을 차지한다. 천장등 하나가 직접 보여 광원을 노출하지 말라는 지시에는 맞지 않는다.",
        "entities": "중년 한국계 여성과 검은 머리의 앳된 한국계 남성 두 명만 보인다. 미연은 참고와 유사한 검은 머리와 체크 셔츠, 손목의 끈 장식을 갖췄지만 모자는 없다. 얼굴의 붓기와 눈 주변 멍이 보인다. 현우는 헝클어진 검은 머리와 얼굴의 상처가 있고 상체가 벗겨져 있어 겉옷 제거가 명확하다. 다리 상처와 신발 속 명함은 프레임 밖이므로 판단하지 않는다. 창밖에는 어두운 거친 물결이 보이지만 밀려오는 거대한 파도의 윤곽은 덜 명확하다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연의 한 손이 현우의 뒤통수를 받치고 다른 팔과 손이 그의 등을 감싼다. 현우의 머리는 미연의 팔과 몸에 실제로 접촉하며, 팔꿈치와 손목의 연결도 가능한 자세다. 하체는 잘려 있어 발의 지지는 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 소화기는 벽걸이 지점에 매달려 있다."
       },
       {
        "label": "B",
        "direction": "미연은 두 눈을 질끈 감고 고개를 현우의 정수리 쪽으로 숙인다. 현우는 얼굴을 미연의 가슴 안쪽으로 향해 묻고 있다. 미연의 손과 팔은 뒤통수와 등을 자기 몸 쪽으로 당기므로 머리를 가슴팍에 밀착시키라는 동작의 방향이 정확하다.",
        "built_space": "왼쪽에 온전한 금속틀 창과 창턱이 있고, 뒤로 같은 벽의 창 개구부 두 곳이 이어진다. 복도 끝에는 창 하나와 위로 올라가는 계단 한 구간, 난간이 보인다. 오른쪽에는 붉은 매립함 하나, 벽걸이 소화기 하나와 가까운 문틀이 있다. 노출 배관과 투톤 벽, 닳은 바닥이 참고 장소에 부합한다. 두 사람은 창 옆을 차지하고 얼굴과 팔이 전경을 크게 채우며 복도는 비스듬한 주변 맥락으로 남는다. 다만 천장등 하나가 직접 노출된다.",
        "entities": "검은 머리의 중년 한국계 여성과 앳된 한국계 남성 두 명만 보인다. 미연의 얼굴 윤곽과 체크 셔츠는 참고에 가깝고, 눈과 뺨 주변의 멍과 부기가 보인다. 참고의 모자는 없다. 현우는 검은 머리와 젊은 체격이 맞지만 얼굴 대부분이 가려져 얼굴 정체성과 멍은 확인할 수 없다. 이는 가슴에 얼굴을 묻는 동작에 자연스럽다. 다만 참고의 겉옷과 유사한 회색 칼라 셔츠가 남아 있다. 다리 상처와 신발 속 명함은 프레임 밖이다. 창밖에는 밤의 거대한 파도 능선이 명확하고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "미연의 한 손이 현우의 뒤통수에 닿아 머리를 붙잡고, 반대 팔과 손이 등을 감싸 몸 쪽으로 당긴다. 현우의 머리와 상체가 미연의 가슴과 팔에 밀착되어 지지 관계가 분명하다. 목을 숙이고 상체를 기울여 만들 수 있는 자세이며, 하체가 잘린 것을 부유로 볼 근거는 없다. 소화기는 벽의 고정 장치에 걸려 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "눈을 질끈 감은 표정과 부상·겉옷 제거는 잘 맞지만, 현우의 얼굴이 가슴보다 미연의 팔 쪽에 기대고 인물 크기도 작아 핵심 포옹의 클로즈업이 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우의 머리를 가슴팍에 빈틈없이 묻고 두 팔로 감싼 밀착 클로즈업이 더 정확하지만, 현우에게 회색 셔츠가 남아 있어 겉옷 제거 지시와 어긋난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 두 눈을 강하게 감고 얼굴을 거의 정면으로 향한다. 현우는 눈을 감고 왼쪽 아래로 얼굴을 돌려 미연의 팔 안쪽에 뺨을 댄다. 미연의 한 손은 뒤통수를, 다른 손은 등과 어깨를 끌어안지만, 현우의 얼굴이 가슴 안쪽으로 완전히 파묻힌 방향은 아니다.",
        "built_space": "왼쪽 전경에 중앙 금속틀로 나뉜 온전한 창유리 두 면과 창턱이 보인다. 뒤쪽에는 끝 창 하나, 올라가는 계단 한 구간과 난간, 오른쪽 벽의 붉은 매립함 하나와 소화기 하나가 보인다. 낡은 투톤 벽과 바닥은 장소 참고와 부합한다. 두 사람은 창 바로 옆 복도에 있으나, 배경과 현우의 등이 상당한 면적을 차지한다. 천장등 하나가 직접 보여 광원을 노출하지 말라는 지시에는 맞지 않는다.",
        "entities": "중년 한국계 여성과 검은 머리의 앳된 한국계 남성 두 명만 보인다. 미연은 참고와 유사한 검은 머리와 체크 셔츠, 손목의 끈 장식을 갖췄지만 모자는 없다. 얼굴의 붓기와 눈 주변 멍이 보인다. 현우는 헝클어진 검은 머리와 얼굴의 상처가 있고 상체가 벗겨져 있어 겉옷 제거가 명확하다. 다리 상처와 신발 속 명함은 프레임 밖이므로 판단하지 않는다. 창밖에는 어두운 거친 물결이 보이지만 밀려오는 거대한 파도의 윤곽은 덜 명확하다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연의 한 손이 현우의 뒤통수를 받치고 다른 팔과 손이 그의 등을 감싼다. 현우의 머리는 미연의 팔과 몸에 실제로 접촉하며, 팔꿈치와 손목의 연결도 가능한 자세다. 하체는 잘려 있어 발의 지지는 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 소화기는 벽걸이 지점에 매달려 있다."
       },
       {
        "label": "A",
        "direction": "미연은 두 눈을 질끈 감고 고개를 현우의 정수리 쪽으로 숙인다. 현우는 얼굴을 미연의 가슴 안쪽으로 향해 묻고 있다. 미연의 손과 팔은 뒤통수와 등을 자기 몸 쪽으로 당기므로 머리를 가슴팍에 밀착시키라는 동작의 방향이 정확하다.",
        "built_space": "왼쪽에 온전한 금속틀 창과 창턱이 있고, 뒤로 같은 벽의 창 개구부 두 곳이 이어진다. 복도 끝에는 창 하나와 위로 올라가는 계단 한 구간, 난간이 보인다. 오른쪽에는 붉은 매립함 하나, 벽걸이 소화기 하나와 가까운 문틀이 있다. 노출 배관과 투톤 벽, 닳은 바닥이 참고 장소에 부합한다. 두 사람은 창 옆을 차지하고 얼굴과 팔이 전경을 크게 채우며 복도는 비스듬한 주변 맥락으로 남는다. 다만 천장등 하나가 직접 노출된다.",
        "entities": "검은 머리의 중년 한국계 여성과 앳된 한국계 남성 두 명만 보인다. 미연의 얼굴 윤곽과 체크 셔츠는 참고에 가깝고, 눈과 뺨 주변의 멍과 부기가 보인다. 참고의 모자는 없다. 현우는 검은 머리와 젊은 체격이 맞지만 얼굴 대부분이 가려져 얼굴 정체성과 멍은 확인할 수 없다. 이는 가슴에 얼굴을 묻는 동작에 자연스럽다. 다만 참고의 겉옷과 유사한 회색 칼라 셔츠가 남아 있다. 다리 상처와 신발 속 명함은 프레임 밖이다. 창밖에는 밤의 거대한 파도 능선이 명확하고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "미연의 한 손이 현우의 뒤통수에 닿아 머리를 붙잡고, 반대 팔과 손이 등을 감싸 몸 쪽으로 당긴다. 현우의 머리와 상체가 미연의 가슴과 팔에 밀착되어 지지 관계가 분명하다. 목을 숙이고 상체를 기울여 만들 수 있는 자세이며, 하체가 잘린 것을 부유로 볼 근거는 없다. 소화기는 벽의 고정 장치에 걸려 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1571,
   "B": 1875
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "현우가 겉옷을 벗어야 한다는 지시를 누락하고 참조 이미지의 의상을 그대로 묘사했으며, 미연의 얼굴 붓기도 다소 약하게 표현되었습니다."
   },
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "현우가 겉옷을 벗은 상태와 미연의 심하게 부어오른 얼굴 등 프롬프트의 구체적인 상태 변화 지시를 정확하고 충실하게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B01.png",
    "asset_id": "09d33278-e815-4961-9ee0-fa14a5350447",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-5efd-7c26-acb2-c6f67e835015",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4__bgfirst_bg.png",
   "bg_asset_id": "6b563ad8-6d4e-4a21-9db4-4cdc611ddc68",
   "bg_record_key": "S37sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S37sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:37:35.702795+00:00",
  "fingerprint": "5406d56b5b1ab86d90e761d0ab70a472f5397111233936334ac187b0c9d7eeb3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S37sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S37sh4_sel.png",
  "source_sha256": "310a77a7761c2b38b5dc1af8412e802472c164fea6f5d32b9d471e6869c08f5d",
  "file": "S37sh4_cine.png",
  "staged_sha256": "681fa00869c15f72c41d8c71a19448a35c5081c13e15c2fa7ad02bf4465b25ae",
  "latency_ms": 10408
 },
 "S37sh5::signage": {
  "fp": "1942506ba6b1cd01",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S37sh5": {
  "input_fingerprint": "017a9a57bf96cce1",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 산산조각 나 흩어지는 복도 창문 파편들과 함께 거대한 해일이 두 사람을 덮치는 폭발적인 찰나.\n\nLOCATION (lock): Inside the upper-floor corridor of the refugee administration building, at the windows bursting inward under the nighttime wave. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Breaking corridor window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Corridor window (Breaking into scattered fragments as the wave enters) — Viewed obliquely from inside the corridor, with the approaching water visible through the breaking opening; used as Upper-right impact boundary above the embracing pair; Giant wave (Crashing through the window and engulfing the pair) — The advancing water enters from the window side toward the corridor interior; used as Connects the window break to the human figures without eliminating spatial context; Upper-floor corridor (Being overtaken by the wave) — The interior extends around and behind the pair; used as Provides scale and defines the camera's retreat away from the window.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime tonal treatment through the impact, preserving readable separation between the embracing figures, water, and breaking glass.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The corridor windows shatter, scattering glass as seawater surges into the management office. 미연: Her face remains swollen from the beating, and her arms remain tightly closed in the protective embrace as the water strikes. 현우: He still has facial bruises and the untreated leg wound, with his outer garment removed. Yoon's contact card remains concealed inside his shoe through the flood.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 산산조각 나 흩어지는 복도 창문 파편들과 함께 거대한 해일이 두 사람을 덮치는 폭발적인 찰나.\n\nLOCATION (lock): Inside the upper-floor corridor of the refugee administration building, at the windows bursting inward under the nighttime wave. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Breaking corridor window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Corridor window (Breaking into scattered fragments as the wave enters) — Viewed obliquely from inside the corridor, with the approaching water visible through the breaking opening; used as Upper-right impact boundary above the embracing pair; Giant wave (Crashing through the window and engulfing the pair) — The advancing water enters from the window side toward the corridor interior; used as Connects the window break to the human figures without eliminating spatial context; Upper-floor corridor (Being overtaken by the wave) — The interior extends around and behind the pair; used as Provides scale and defines the camera's retreat away from the window.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime tonal treatment through the impact, preserving readable separation between the embracing figures, water, and breaking glass.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The corridor windows shatter, scattering glass as seawater surges into the management office. 미연: Her face remains swollen from the beating, and her arms remain tightly closed in the protective embrace as the water strikes. 현우: He still has facial bruises and the untreated leg wound, with his outer garment removed. Yoon's contact card remains concealed inside his shoe through the flood.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 산산조각 나 흩어지는 복도 창문 파편들과 함께 거대한 해일이 두 사람을 덮치는 폭발적인 찰나.\n\nLOCATION (lock): Inside the upper-floor corridor of the refugee administration building, at the windows bursting inward under the nighttime wave. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Breaking corridor window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Corridor window (Breaking into scattered fragments as the wave enters) — Viewed obliquely from inside the corridor, with the approaching water visible through the breaking opening; used as Upper-right impact boundary above the embracing pair; Giant wave (Crashing through the window and engulfing the pair) — The advancing water enters from the window side toward the corridor interior; used as Connects the window break to the human figures without eliminating spatial context; Upper-floor corridor (Being overtaken by the wave) — The interior extends around and behind the pair; used as Provides scale and defines the camera's retreat away from the window.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime tonal treatment through the impact, preserving readable separation between the embracing figures, water, and breaking glass.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The corridor windows shatter, scattering glass as seawater surges into the management office. 미연: Her face remains swollen from the beating, and her arms remain tightly closed in the protective embrace as the water strikes. 현우: He still has facial bruises and the untreated leg wound, with his outer garment removed. Yoon's contact card remains concealed inside his shoe through the flood.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "파도가 오른쪽 창문에서 안쪽으로 밀려듦.",
    "built_space": "복도 내부. 우측에 창문, 좌측에 소화전 위치.",
    "entities": "미연(체크 셔츠), 현우(상의 탈의). 미연의 팔이 3개 이상 나타남.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (미연의 팔이 3개 이상임)"
    ],
    "physics": "유리 파편과 물보라가 공중에 흩날림. 두 사람은 바닥을 딛고 섬."
   },
   {
    "label": "B",
    "direction": "파도가 우측 창문을 깨고 들어옴.",
    "built_space": "지정된 복도 내부. 카메라가 뒤로 물러난 넓은 샷.",
    "entities": "미연(체크 셔츠), 현우(상의 탈의 지시와 달리 셔츠 착용).",
    "hard_violations": [],
    "physics": "충격으로 유리가 날아가며 파도가 밀려옴. 두 사람은 바닥을 딛고 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미연의 팔이 3개 이상으로 그려지는 치명적인 해부학적 오류가 발생했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우가 상의를 입고 있어 복장 지시를 어겼으나, 지정된 구도와 파도가 덮치는 상황 연출은 안정적입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "파도가 오른쪽 창문에서 안쪽으로 밀려듦.",
        "built_space": "복도 내부. 우측에 창문, 좌측에 소화전 위치.",
        "entities": "미연(체크 셔츠), 현우(상의 탈의). 미연의 팔이 3개 이상 나타남.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (미연의 팔이 3개 이상임)"
        ],
        "physics": "유리 파편과 물보라가 공중에 흩날림. 두 사람은 바닥을 딛고 섬."
       },
       {
        "label": "B",
        "direction": "파도가 우측 창문을 깨고 들어옴.",
        "built_space": "지정된 복도 내부. 카메라가 뒤로 물러난 넓은 샷.",
        "entities": "미연(체크 셔츠), 현우(상의 탈의 지시와 달리 셔츠 착용).",
        "hard_violations": [],
        "physics": "충격으로 유리가 날아가며 파도가 밀려옴. 두 사람은 바닥을 딛고 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미연의 팔이 3개 이상으로 그려지는 치명적인 해부학적 오류가 발생했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우가 상의를 입고 있어 복장 지시를 어겼으나, 지정된 구도와 파도가 덮치는 상황 연출은 안정적입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "파도가 오른쪽 창문에서 안쪽으로 밀려듦.",
        "built_space": "복도 내부. 우측에 창문, 좌측에 소화전 위치.",
        "entities": "미연(체크 셔츠), 현우(상의 탈의). 미연의 팔이 3개 이상 나타남.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (미연의 팔이 3개 이상임)"
        ],
        "physics": "유리 파편과 물보라가 공중에 흩날림. 두 사람은 바닥을 딛고 섬."
       },
       {
        "label": "B",
        "direction": "파도가 우측 창문을 깨고 들어옴.",
        "built_space": "지정된 복도 내부. 카메라가 뒤로 물러난 넓은 샷.",
        "entities": "미연(체크 셔츠), 현우(상의 탈의 지시와 달리 셔츠 착용).",
        "hard_violations": [],
        "physics": "충격으로 유리가 날아가며 파도가 밀려옴. 두 사람은 바닥을 딛고 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "우상단 창문과 복도의 공간감은 구현했지만, 물이 두 사람을 덮치는 정도가 약하고 현우의 등에 회색 상의를 남겨 이전 장면의 맨상체 상태를 어겼습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "우상단의 깨진 창에서 밀려든 해수가 포옹한 두 사람에게 직접 부딪히는 찰나를 복도 전경과 함께 보여주며, 현우의 맨상체와 미연의 보호 자세도 유지합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 눈을 감고 현우의 머리를 감싸며, 현우는 얼굴을 왼쪽 아래로 숙입니다. 유리와 물은 오른쪽 창에서 왼쪽 복도 안으로 들어옵니다. 물의 선두가 두 사람의 창가 쪽 몸에 접근하지만, 주된 물덩어리는 아직 창턱 부근에 있어 두 사람을 덮치는 충격은 약합니다. 무기나 방향을 판별할 휴대 물건은 없습니다.",
        "built_space": "오른쪽에 크게 깨진 창 구획 하나와 그 뒤로 이어지는 창 구획들이 있고, 왼쪽에는 붉은 매립함 하나와 소화기 하나가 보입니다. 밝은 직사각형 천장등 하나와 먼 천장의 어두운 등기구가 보이며, 복도는 왼쪽 뒤로 깊어집니다. 두 사람은 창턱 바로 안쪽에 있습니다. 낡은 이색 벽면, 벗겨진 창턱, 금속 창틀은 참고 장소와 대체로 맞습니다. 깨진 창은 우상단에 있지만 배경보다는 가까운 측면 공간을 크게 차지합니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "사람은 두 명뿐입니다. 미연은 중년 동아시아계 여성으로 검은 머리, 체크 셔츠, 얼굴의 부기와 멍이 참고와 대체로 맞습니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리를 유지하지만, 어깨 아래 등 대부분에 회색 상의가 걸쳐 있어 이전 장면의 맨상체와 다릅니다. 현우의 얼굴 상처는 각도상 선명하지 않습니다. 다리 상처와 신발 속 명함은 프레임 밖이라 확인할 수 없습니다. 해수와 흩어진 유리 조각은 보이며 판독 가능한 글자는 없습니다.",
        "hard_violations": [],
        "physics": "두 사람의 하체는 화면 아래로 이어지고 서로의 몸을 팔과 손으로 붙잡고 있습니다. 발은 잘렸지만 몸이 공중에 떠 있다는 징후는 없습니다. 미연의 손은 현우의 머리와 등에 접촉합니다. 날아가는 파편과 물에는 창밖 해수의 충격이라는 발생 원인이 있고, 일부 물은 창턱을 넘어 바닥으로 떨어집니다. 명백한 무지지 부유물이나 불가능한 관절은 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 감고 현우를 감싸고, 현우는 얼굴을 아래로 숙여 미연 쪽에 붙입니다. 물과 유리 조각은 오른쪽 위의 깨진 창에서 복도 안쪽과 왼쪽으로 뻗습니다. 물보라가 두 사람의 어깨와 등, 옆구리에 실제로 닿아 창 파괴와 인물 피격의 방향이 연결됩니다. 무기나 방향을 판별할 휴대 물건은 없습니다.",
        "built_space": "오른쪽에 크게 깨진 창 구획 하나와 뒤로 이어지는 금속 창들이 있고, 왼쪽 벽에는 붉은 매립함 하나와 소화기 하나가 있습니다. 직사각형 천장등 두 개, 복도 끝의 계단 한 곳과 난간이 보입니다. 두 사람은 창턱 안쪽에 붙어 있고 그 왼쪽으로 복도 바닥과 통로가 길게 남습니다. 참고의 낡은 이색 벽, 벗겨진 창턱, 바닥 재질과 차가운 야간 조명을 대체로 유지합니다. 깨진 창은 우상단이지만 엄밀한 먼 배경보다는 인물 바로 뒤의 측면입니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "사람은 미연과 현우에 해당하는 두 명뿐입니다. 미연은 검은 머리의 중년 동아시아계 여성으로 체크 셔츠와 어두운 앞치마를 입고 있으며, 얼굴의 부기는 옆얼굴과 물보라 때문에 덜 선명합니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 얼굴의 멍과 상처, 상의를 벗은 상태가 보입니다. 두 사람의 외모와 의상은 이전 장면에 대체로 부합합니다. 다리 상처와 신발 속 명함은 화면 밖이라 판단하지 않습니다. 거대한 물결, 해수와 유리 파편이 보이고 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "두 사람은 상체를 낮추고 서로를 팔로 붙들며 충격을 버팁니다. 손은 상대의 어깨와 등에 닿아 있고, 하체는 화면 아래로 이어져 부유하는 자세로 보이지 않습니다. 물은 창밖 파도에서 끊김 없이 들어와 몸에 부딪히고 바닥으로 쏟아집니다. 공중의 유리 조각도 창 파괴의 충격으로 설명되며, 지지나 운동 원인이 없는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "우상단 창문과 복도의 공간감은 구현했지만, 물이 두 사람을 덮치는 정도가 약하고 현우의 등에 회색 상의를 남겨 이전 장면의 맨상체 상태를 어겼습니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "우상단의 깨진 창에서 밀려든 해수가 포옹한 두 사람에게 직접 부딪히는 찰나를 복도 전경과 함께 보여주며, 현우의 맨상체와 미연의 보호 자세도 유지합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 눈을 감고 현우의 머리를 감싸며, 현우는 얼굴을 왼쪽 아래로 숙입니다. 유리와 물은 오른쪽 창에서 왼쪽 복도 안으로 들어옵니다. 물의 선두가 두 사람의 창가 쪽 몸에 접근하지만, 주된 물덩어리는 아직 창턱 부근에 있어 두 사람을 덮치는 충격은 약합니다. 무기나 방향을 판별할 휴대 물건은 없습니다.",
        "built_space": "오른쪽에 크게 깨진 창 구획 하나와 그 뒤로 이어지는 창 구획들이 있고, 왼쪽에는 붉은 매립함 하나와 소화기 하나가 보입니다. 밝은 직사각형 천장등 하나와 먼 천장의 어두운 등기구가 보이며, 복도는 왼쪽 뒤로 깊어집니다. 두 사람은 창턱 바로 안쪽에 있습니다. 낡은 이색 벽면, 벗겨진 창턱, 금속 창틀은 참고 장소와 대체로 맞습니다. 깨진 창은 우상단에 있지만 배경보다는 가까운 측면 공간을 크게 차지합니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "사람은 두 명뿐입니다. 미연은 중년 동아시아계 여성으로 검은 머리, 체크 셔츠, 얼굴의 부기와 멍이 참고와 대체로 맞습니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리를 유지하지만, 어깨 아래 등 대부분에 회색 상의가 걸쳐 있어 이전 장면의 맨상체와 다릅니다. 현우의 얼굴 상처는 각도상 선명하지 않습니다. 다리 상처와 신발 속 명함은 프레임 밖이라 확인할 수 없습니다. 해수와 흩어진 유리 조각은 보이며 판독 가능한 글자는 없습니다.",
        "hard_violations": [],
        "physics": "두 사람의 하체는 화면 아래로 이어지고 서로의 몸을 팔과 손으로 붙잡고 있습니다. 발은 잘렸지만 몸이 공중에 떠 있다는 징후는 없습니다. 미연의 손은 현우의 머리와 등에 접촉합니다. 날아가는 파편과 물에는 창밖 해수의 충격이라는 발생 원인이 있고, 일부 물은 창턱을 넘어 바닥으로 떨어집니다. 명백한 무지지 부유물이나 불가능한 관절은 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "미연은 눈을 감고 현우를 감싸고, 현우는 얼굴을 아래로 숙여 미연 쪽에 붙입니다. 물과 유리 조각은 오른쪽 위의 깨진 창에서 복도 안쪽과 왼쪽으로 뻗습니다. 물보라가 두 사람의 어깨와 등, 옆구리에 실제로 닿아 창 파괴와 인물 피격의 방향이 연결됩니다. 무기나 방향을 판별할 휴대 물건은 없습니다.",
        "built_space": "오른쪽에 크게 깨진 창 구획 하나와 뒤로 이어지는 금속 창들이 있고, 왼쪽 벽에는 붉은 매립함 하나와 소화기 하나가 있습니다. 직사각형 천장등 두 개, 복도 끝의 계단 한 곳과 난간이 보입니다. 두 사람은 창턱 안쪽에 붙어 있고 그 왼쪽으로 복도 바닥과 통로가 길게 남습니다. 참고의 낡은 이색 벽, 벗겨진 창턱, 바닥 재질과 차가운 야간 조명을 대체로 유지합니다. 깨진 창은 우상단이지만 엄밀한 먼 배경보다는 인물 바로 뒤의 측면입니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "사람은 미연과 현우에 해당하는 두 명뿐입니다. 미연은 검은 머리의 중년 동아시아계 여성으로 체크 셔츠와 어두운 앞치마를 입고 있으며, 얼굴의 부기는 옆얼굴과 물보라 때문에 덜 선명합니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 얼굴의 멍과 상처, 상의를 벗은 상태가 보입니다. 두 사람의 외모와 의상은 이전 장면에 대체로 부합합니다. 다리 상처와 신발 속 명함은 화면 밖이라 판단하지 않습니다. 거대한 물결, 해수와 유리 파편이 보이고 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "두 사람은 상체를 낮추고 서로를 팔로 붙들며 충격을 버팁니다. 손은 상대의 어깨와 등에 닿아 있고, 하체는 화면 아래로 이어져 부유하는 자세로 보이지 않습니다. 물은 창밖 파도에서 끊김 없이 들어와 몸에 부딪히고 바닥으로 쏟아집니다. 공중의 유리 조각도 창 파괴의 충격으로 설명되며, 지지나 운동 원인이 없는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.667
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (미연의 팔이 3개 이상임)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1667
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "미연의 팔이 3개 이상으로 그려지는 치명적인 해부학적 오류가 발생했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (미연의 팔이 3개 이상임)"
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "현우가 상의를 입고 있어 복장 지시를 어겼으나, 지정된 구도와 파도가 덮치는 상황 연출은 안정적입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4_sel.png",
    "asset_id": "b7c434fe-5350-4b8a-9cb6-8a985ffe6c1b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-6251-7244-a431-c7367bc0b605",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S37sh4"
  }
 },
 "S37sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:39:15.143412+00:00",
  "fingerprint": "3f6b1713ef72c448dff6a05dcb786eb7355f630b5d0c302a1a4515d18edb2eac",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S37sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S37sh5_sel.png",
  "source_sha256": "60971a247465ab90fa6ee4f88d10125a7894c4ed1f445ad01204e2d588443f00",
  "file": "S37sh5_cine.png",
  "staged_sha256": "e25a96dd27c82c8048427e8a63c2d62b9e62ec4f85598809fcca253c61cb88af",
  "latency_ms": 8058
 },
 "S38sh2::signage": {
  "fp": "160a517829789c7a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5c8660f181826b50": {
  "subjects": [],
  "subject_text": "수몰된 인천 난민촌 수면과 떠다니는 컨테이너\n넓게 펼쳐진 바닷물 위에 철제 컨테이너와 판자가 떠 있는 공간. 담벼락 일부가 수면 위로 드러나고 희미한 새벽빛이 물에 비친다.",
  "identity": "canonical",
  "scope_id": "L193",
  "scope_role": "location_exterior",
  "scope_sha": "a3674d85c9b2a6df"
 },
 "groupbg::flood_raft_water": {
  "input_fingerprint": "dba0c01f7f8f09a9",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "flood_raft_water",
    "tags": [
     "S38sh13",
     "S38sh2",
     "S38sh7",
     "S63sh16"
    ]
   },
   "context_sig": "70ee00269da35a94"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 저 멀리 수면 위에 컨테이너 판자 위에 누워있는 미연.\n- / INSERT(회상F.B): 앰버를... 하면서 숨을 거두기 직전의 엄마 모습 /\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 저 멀리 수면 위에 컨테이너 판자 위에 누워있는 미연.\n- / INSERT(회상F.B): 앰버를... 하면서 숨을 거두기 직전의 엄마 모습 /\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_raft_water_456400.png",
  "asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448",
  "input_asset_ids": [
   "316d049e-c00b-4b04-9191-f52cdd862974"
  ],
  "origin_tag": "S38sh2",
  "place_text": "On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.",
  "origin_inputs": {
   "place_text": "On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.",
   "time_of_day_en": "dawn",
   "conti_asset_id": "316d049e-c00b-4b04-9191-f52cdd862974"
  }
 },
 "S38sh2::bgfirst_bg": {
  "input_fingerprint": "dc101a1a466114c0",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2__bgfirst_bg.png",
  "asset_id": "6a473823-e079-4c39-92bf-e7c364af89d6",
  "input_asset_ids": [
   "316d049e-c00b-4b04-9191-f52cdd862974",
   "0797d468-af7c-4934-b8f4-036bf4cc5448"
  ]
 },
 "S38sh2": {
  "input_fingerprint": "51d05d0480a69d7a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): At dawn, separate container panels float on the sea amid the flood wreckage. Storage boxes remain inside the container that has served as overnight shelter. 미연: She lies soaked and motionless on a floating container panel, with her face still swollen. Her abdominal puncture wound is already bleeding, although it is discovered only upon closer approach.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): At dawn, separate container panels float on the sea amid the flood wreckage. Storage boxes remain inside the container that has served as overnight shelter. 미연: She lies soaked and motionless on a floating container panel, with her face still swollen. Her abdominal puncture wound is already bleeding, although it is discovered only upon closer approach.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): At dawn, separate container panels float on the sea amid the flood wreckage. Storage boxes remain inside the container that has served as overnight shelter. 미연: She lies soaked and motionless on a floating container panel, with her face still swollen. Her abdominal puncture wound is already bleeding, although it is discovered only upon closer approach.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2__bgfirst_bg.png",
     "asset_id": "6a473823-e079-4c39-92bf-e7c364af89d6",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S38sh2.png",
     "asset_id": "316d049e-c00b-4b04-9191-f52cdd862974",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1338232>",
     "asset_id": "29c79c02-d280-452c-b70e-07da4af44bfa",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_raft_water_456400.png",
     "asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1338232>",
     "asset_id": "29c79c02-d280-452c-b70e-07da4af44bfa",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "카메라(현우의 시점)가 멀리 수면 위에 떠 있는 컨테이너 판자와 그 위의 미연을 향하고 있음.",
    "built_space": "참조 이미지와 동일한 공간. 좌측 하단의 근경 구조물과 우측 원경의 컨테이너 판자가 올바른 위치에 1개씩 존재함.",
    "entities": "지시된 모자와 작업복을 착용하고 피를 흘리며 원경의 판자 위에 누워 있는 미연의 실루엣이 확인됨.",
    "hard_violations": [],
    "physics": "미연의 전신이 컨테이너 판자 표면에 중력에 맞게 완전히 지지되어 누워 있음."
   },
   {
    "label": "B",
    "direction": "카메라가 수면 중앙에 가깝게 배치된 컨테이너 판자 위의 미연을 향함.",
    "built_space": "좌측 하단 구조물과 우측 원경의 기존 판자가 있는 상태에서, 수면 중앙에 원래 로케이션에 없는 거대한 컨테이너 판자가 추가로 생성됨.",
    "entities": "미연이 누워있으나 캐릭터 참조에 있는 캡 모자가 누락됨.",
    "hard_violations": [
     "[gemini-pro] 로케이션 참조에 존재하지 않는 구조물(수면 중앙의 컨테이너 판자)을 화면에 임의로 추가 생성함."
    ],
    "physics": "미연의 몸이 중앙에 생성된 판자 위에 지지되어 누워 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "원경의 프레이밍과 로케이션 지시를 완벽히 준수하여, 참조 사진에 존재하는 멀리 떨어진 판자 위에 미연을 정확히 배치함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "로케이션 참조 이미지에 없는 새로운 판자를 수면 중앙에 임의로 생성하는 공간적 오류를 범했으며 캐릭터의 모자가 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라(현우의 시점)가 멀리 수면 위에 떠 있는 컨테이너 판자와 그 위의 미연을 향하고 있음.",
        "built_space": "참조 이미지와 동일한 공간. 좌측 하단의 근경 구조물과 우측 원경의 컨테이너 판자가 올바른 위치에 1개씩 존재함.",
        "entities": "지시된 모자와 작업복을 착용하고 피를 흘리며 원경의 판자 위에 누워 있는 미연의 실루엣이 확인됨.",
        "hard_violations": [],
        "physics": "미연의 전신이 컨테이너 판자 표면에 중력에 맞게 완전히 지지되어 누워 있음."
       },
       {
        "label": "B",
        "direction": "카메라가 수면 중앙에 가깝게 배치된 컨테이너 판자 위의 미연을 향함.",
        "built_space": "좌측 하단 구조물과 우측 원경의 기존 판자가 있는 상태에서, 수면 중앙에 원래 로케이션에 없는 거대한 컨테이너 판자가 추가로 생성됨.",
        "entities": "미연이 누워있으나 캐릭터 참조에 있는 캡 모자가 누락됨.",
        "hard_violations": [
         "로케이션 참조에 존재하지 않는 구조물(수면 중앙의 컨테이너 판자)을 화면에 임의로 추가 생성함."
        ],
        "physics": "미연의 몸이 중앙에 생성된 판자 위에 지지되어 누워 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "원경의 프레이밍과 로케이션 지시를 완벽히 준수하여, 참조 사진에 존재하는 멀리 떨어진 판자 위에 미연을 정확히 배치함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "로케이션 참조 이미지에 없는 새로운 판자를 수면 중앙에 임의로 생성하는 공간적 오류를 범했으며 캐릭터의 모자가 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라(현우의 시점)가 멀리 수면 위에 떠 있는 컨테이너 판자와 그 위의 미연을 향하고 있음.",
        "built_space": "참조 이미지와 동일한 공간. 좌측 하단의 근경 구조물과 우측 원경의 컨테이너 판자가 올바른 위치에 1개씩 존재함.",
        "entities": "지시된 모자와 작업복을 착용하고 피를 흘리며 원경의 판자 위에 누워 있는 미연의 실루엣이 확인됨.",
        "hard_violations": [],
        "physics": "미연의 전신이 컨테이너 판자 표면에 중력에 맞게 완전히 지지되어 누워 있음."
       },
       {
        "label": "B",
        "direction": "카메라가 수면 중앙에 가깝게 배치된 컨테이너 판자 위의 미연을 향함.",
        "built_space": "좌측 하단 구조물과 우측 원경의 기존 판자가 있는 상태에서, 수면 중앙에 원래 로케이션에 없는 거대한 컨테이너 판자가 추가로 생성됨.",
        "entities": "미연이 누워있으나 캐릭터 참조에 있는 캡 모자가 누락됨.",
        "hard_violations": [
         "로케이션 참조에 존재하지 않는 구조물(수면 중앙의 컨테이너 판자)을 화면에 임의로 추가 생성함."
        ],
        "physics": "미연의 몸이 중앙에 생성된 판자 위에 지지되어 누워 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "전신의 지지와 차분한 새벽빛은 적절하지만, 미연의 판자가 화면 중앙에 너무 크고 가깝게 놓여 핵심인 먼 거리와 넓은 수면 간격을 약화한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "넓은 수면 너머 오른쪽의 작은 판자에 미연의 전신을 배치해 현우의 원거리 시점에 더 충실하며, 다만 판자 윗면의 가시성과 절제된 새벽 색조는 다소 부족하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 수면 가까이에서 중앙 판자의 미연을 비스듬히 바라본다. 미연은 머리가 왼쪽, 발이 오른쪽이며 얼굴은 위를 향한다. 현우의 몸이나 별도의 인물은 보이지 않아 주관 시점으로 읽히지만, 바라보는 대상이 요구보다 가깝다. 조준하거나 이동하는 물체는 없다.",
        "built_space": "왼쪽에 기울어진 큰 컨테이너, 오른쪽에 침수된 컨테이너와 비스듬한 골조, 수평선에 도시 윤곽이 있다. 왼쪽 아래 전경 패널 하나, 중앙의 미연 지지 패널 하나, 그 뒤 오른쪽의 긴 패널 하나가 구별된다. 주변 잔해와 녹슨 금속 재질은 장소 참조와 부합한다. 미연의 패널은 윗면과 측면이 잘 보이지만 화면 폭의 절반 이상을 차지하여 작은 원거리 지지물이라는 구성에서 벗어난다.",
        "entities": "검은 머리의 중년 여성 한 명만 보이며, 창백한 얼굴과 체크 셔츠, 갈색 앞치마는 미연의 참조와 대체로 맞는다. 참조의 모자는 보이지 않는다. 복부와 옷에 짙은 혈흔이 보이지만 천공 자체와 얼굴 부종의 정도는 확정하기 어렵다. 전신이 구체적인 사람으로 표현되어 있고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 몸통, 골반, 다리는 금속 패널 위에 놓여 있고 가까운 팔과 손도 패널에 내려앉아 있다. 신발 뒤쪽도 패널에 닿아 있어 무지지 부유나 의도적으로 들어 올린 팔다리는 보이지 않는다. 패널의 하부는 물에 잠겨 부력을 받는 배치이며 주변 잔해도 수면에 접한다."
       },
       {
        "label": "B",
        "direction": "카메라는 넓은 수면을 가로질러 오른쪽 중경의 미연을 바라본다. 미연은 머리가 왼쪽, 발이 오른쪽이고 얼굴은 위쪽을 향한다. 현우를 화면에 넣지 않고 관찰 대상만 보여 주므로 요구한 주관 시점에 부합한다. 별도의 응시 대상이나 조준 물체, 이동 동작은 없다.",
        "built_space": "왼쪽의 기울어진 컨테이너, 오른쪽 침수 컨테이너와 골조, 먼 도시 윤곽이 장소 참조와 대응한다. 왼쪽 아래에 전경 패널 하나, 오른쪽 중경에 미연이 누운 패널 하나가 보이며 사이에는 넓은 열린 수면이 이어진다. 지지 패널은 화면 폭의 약 3분의 1로 A보다 작고 멀리 놓였다. 낮은 시점 때문에 좁은 측면은 명확하지만 몸 아래 윗면은 매우 얇게 드러난다. 낮은 해와 수면 반사는 서로 맞는 방향이다.",
        "entities": "모자를 쓴 검은 머리의 여성 한 명이 체크 셔츠와 갈색 앞치마를 입고 누워 있어 미연의 참조 복장에 부합한다. 옆얼굴은 창백하게 보이지만 이 거리에서 정확한 연령, 한국인 정체성, 얼굴 일치도와 부종을 세밀하게 검증하기는 어렵다. 복부는 어둡고 얼룩져 있으나 출혈 부위는 뚜렷하지 않다. 신체는 전신의 사람으로 식별되며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리, 몸통, 골반과 다리가 패널에 누워 있고 가까운 손도 가장자리의 윗면에 내려져 있다. 발끝은 위로 향하지만 뒤꿈치 쪽이 지지되어 다리 전체가 공중에 떠 있는 자세는 아니다. 머리 왼쪽의 돌출부는 모자 챙으로 보인다. 패널은 하단이 수면에 잠긴 채 몸을 받치며, 지지 없는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "전신의 지지와 차분한 새벽빛은 적절하지만, 미연의 판자가 화면 중앙에 너무 크고 가깝게 놓여 핵심인 먼 거리와 넓은 수면 간격을 약화한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "넓은 수면 너머 오른쪽의 작은 판자에 미연의 전신을 배치해 현우의 원거리 시점에 더 충실하며, 다만 판자 윗면의 가시성과 절제된 새벽 색조는 다소 부족하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 수면 가까이에서 중앙 판자의 미연을 비스듬히 바라본다. 미연은 머리가 왼쪽, 발이 오른쪽이며 얼굴은 위를 향한다. 현우의 몸이나 별도의 인물은 보이지 않아 주관 시점으로 읽히지만, 바라보는 대상이 요구보다 가깝다. 조준하거나 이동하는 물체는 없다.",
        "built_space": "왼쪽에 기울어진 큰 컨테이너, 오른쪽에 침수된 컨테이너와 비스듬한 골조, 수평선에 도시 윤곽이 있다. 왼쪽 아래 전경 패널 하나, 중앙의 미연 지지 패널 하나, 그 뒤 오른쪽의 긴 패널 하나가 구별된다. 주변 잔해와 녹슨 금속 재질은 장소 참조와 부합한다. 미연의 패널은 윗면과 측면이 잘 보이지만 화면 폭의 절반 이상을 차지하여 작은 원거리 지지물이라는 구성에서 벗어난다.",
        "entities": "검은 머리의 중년 여성 한 명만 보이며, 창백한 얼굴과 체크 셔츠, 갈색 앞치마는 미연의 참조와 대체로 맞는다. 참조의 모자는 보이지 않는다. 복부와 옷에 짙은 혈흔이 보이지만 천공 자체와 얼굴 부종의 정도는 확정하기 어렵다. 전신이 구체적인 사람으로 표현되어 있고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 몸통, 골반, 다리는 금속 패널 위에 놓여 있고 가까운 팔과 손도 패널에 내려앉아 있다. 신발 뒤쪽도 패널에 닿아 있어 무지지 부유나 의도적으로 들어 올린 팔다리는 보이지 않는다. 패널의 하부는 물에 잠겨 부력을 받는 배치이며 주변 잔해도 수면에 접한다."
       },
       {
        "label": "A",
        "direction": "카메라는 넓은 수면을 가로질러 오른쪽 중경의 미연을 바라본다. 미연은 머리가 왼쪽, 발이 오른쪽이고 얼굴은 위쪽을 향한다. 현우를 화면에 넣지 않고 관찰 대상만 보여 주므로 요구한 주관 시점에 부합한다. 별도의 응시 대상이나 조준 물체, 이동 동작은 없다.",
        "built_space": "왼쪽의 기울어진 컨테이너, 오른쪽 침수 컨테이너와 골조, 먼 도시 윤곽이 장소 참조와 대응한다. 왼쪽 아래에 전경 패널 하나, 오른쪽 중경에 미연이 누운 패널 하나가 보이며 사이에는 넓은 열린 수면이 이어진다. 지지 패널은 화면 폭의 약 3분의 1로 A보다 작고 멀리 놓였다. 낮은 시점 때문에 좁은 측면은 명확하지만 몸 아래 윗면은 매우 얇게 드러난다. 낮은 해와 수면 반사는 서로 맞는 방향이다.",
        "entities": "모자를 쓴 검은 머리의 여성 한 명이 체크 셔츠와 갈색 앞치마를 입고 누워 있어 미연의 참조 복장에 부합한다. 옆얼굴은 창백하게 보이지만 이 거리에서 정확한 연령, 한국인 정체성, 얼굴 일치도와 부종을 세밀하게 검증하기는 어렵다. 복부는 어둡고 얼룩져 있으나 출혈 부위는 뚜렷하지 않다. 신체는 전신의 사람으로 식별되며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리, 몸통, 골반과 다리가 패널에 누워 있고 가까운 손도 가장자리의 윗면에 내려져 있다. 발끝은 위로 향하지만 뒤꿈치 쪽이 지지되어 다리 전체가 공중에 떠 있는 자세는 아니다. 머리 왼쪽의 돌출부는 모자 챙으로 보인다. 패널은 하단이 수면에 잠긴 채 몸을 받치며, 지지 없는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.179
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.929
   },
   "violations": {
    "B": [
     "[gemini-pro] 로케이션 참조에 존재하지 않는 구조물(수면 중앙의 컨테이너 판자)을 화면에 임의로 추가 생성함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "원경의 프레이밍과 로케이션 지시를 완벽히 준수하여, 참조 사진에 존재하는 멀리 떨어진 판자 위에 미연을 정확히 배치함."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "로케이션 참조 이미지에 없는 새로운 판자를 수면 중앙에 임의로 생성하는 공간적 오류를 범했으며 캐릭터의 모자가 누락됨.  ★위반: [gemini-pro] 로케이션 참조에 존재하지 않는 구조물(수면 중앙의 컨테이너 판자)을 화면에 임의로 추가 생성함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_raft_water_456400.png",
    "asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1338232>",
    "asset_id": "29c79c02-d280-452c-b70e-07da4af44bfa",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-63f3-76cd-aaeb-5a01060980fe",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2__bgfirst_bg.png",
   "bg_asset_id": "6a473823-e079-4c39-92bf-e7c364af89d6",
   "bg_record_key": "S38sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "flood_raft_water",
   "groupbg_asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S38sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:40:13.080441+00:00",
  "fingerprint": "27b328d7310c370c30ad8035f100185f759b220750ce38044272ea8c7e89bf7e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S38sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S38sh2_sel.png",
  "source_sha256": "ce0f7322b05404edb881975726f6b328c9c4ddf1f90c51e1fe66d8993e1a92a4",
  "file": "S38sh2_cine.png",
  "staged_sha256": "18f3f7a3ec6f19d3e13a431144015de3fdc52df1929bddfb07112aeb353a6e03",
  "latency_ms": 11159
 },
 "S38sh7::signage": {
  "fp": "7ea4e02b84278fc1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S38sh7": {
  "input_fingerprint": "41fd12bf0e0b6175",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 숨을 거둔 미연의 몸을 끌어안은 채 새벽하늘을 향해 입을 크게 벌리고 절규하는 현우의 얼굴.\n\nLOCATION (lock): On the floating container panel supporting the injured mother, surrounded by the flooded settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container board (Floating beneath the pair) — Only a small portion of its supporting face is visible at the lower edge; used as Maintains the physical basis of 현우's seated embrace; Sea (Surrounding the floating board); used as A peripheral strip preserves the exposed setting without distracting from the faces; Dawn sky (Visible above the sea at daybreak); used as Provides uncluttered space above 현우's upward cry.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft early-dawn ambient illumination preserves facial detail and restrained contrast, with no change in light treatment to signal 미연's death.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the floating container panel, surrounding floodwater, and dawn coloration. Exclude unrelated floating platforms and do not import the reference's human figure as part of the setting.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon's motionless body is held against Hyunwoo in his embrace on the floating container panel after her death. The source does not specify how far he lifts her torso, how her head rests, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container panels and surviving containers remain afloat on the dawn sea. 미연: She is now dead, with closed eyes, a swollen face and a bleeding abdominal puncture wound. Her body remains on the floating container panel. 현우: He is on the floating panel, crying with his arms closed in an embrace. His facial bruises and leg wound remain, and the contact card is still hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 숨을 거둔 미연의 몸을 끌어안은 채 새벽하늘을 향해 입을 크게 벌리고 절규하는 현우의 얼굴.\n\nLOCATION (lock): On the floating container panel supporting the injured mother, surrounded by the flooded settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container board (Floating beneath the pair) — Only a small portion of its supporting face is visible at the lower edge; used as Maintains the physical basis of 현우's seated embrace; Sea (Surrounding the floating board); used as A peripheral strip preserves the exposed setting without distracting from the faces; Dawn sky (Visible above the sea at daybreak); used as Provides uncluttered space above 현우's upward cry.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft early-dawn ambient illumination preserves facial detail and restrained contrast, with no change in light treatment to signal 미연's death.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the floating container panel, surrounding floodwater, and dawn coloration. Exclude unrelated floating platforms and do not import the reference's human figure as part of the setting.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon's motionless body is held against Hyunwoo in his embrace on the floating container panel after her death. The source does not specify how far he lifts her torso, how her head rests, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container panels and surviving containers remain afloat on the dawn sea. 미연: She is now dead, with closed eyes, a swollen face and a bleeding abdominal puncture wound. Her body remains on the floating container panel. 현우: He is on the floating panel, crying with his arms closed in an embrace. His facial bruises and leg wound remain, and the contact card is still hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 숨을 거둔 미연의 몸을 끌어안은 채 새벽하늘을 향해 입을 크게 벌리고 절규하는 현우의 얼굴.\n\nLOCATION (lock): On the floating container panel supporting the injured mother, surrounded by the flooded settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container board (Floating beneath the pair) — Only a small portion of its supporting face is visible at the lower edge; used as Maintains the physical basis of 현우's seated embrace; Sea (Surrounding the floating board); used as A peripheral strip preserves the exposed setting without distracting from the faces; Dawn sky (Visible above the sea at daybreak); used as Provides uncluttered space above 현우's upward cry.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft early-dawn ambient illumination preserves facial detail and restrained contrast, with no change in light treatment to signal 미연's death.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the floating container panel, surrounding floodwater, and dawn coloration. Exclude unrelated floating platforms and do not import the reference's human figure as part of the setting.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon's motionless body is held against Hyunwoo in his embrace on the floating container panel after her death. The source does not specify how far he lifts her torso, how her head rests, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container panels and surviving containers remain afloat on the dawn sea. 미연: She is now dead, with closed eyes, a swollen face and a bleeding abdominal puncture wound. Her body remains on the floating container panel. 현우: He is on the floating panel, crying with his arms closed in an embrace. His facial bruises and leg wound remain, and the contact card is still hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우가 위를 올려다보며 새벽하늘을 향해 소리치고 있습니다.",
    "built_space": "하단 좌측에 컨테이너 패널이 보이고 주변으로 바다와 여명의 하늘이 둘러싸고 있습니다.",
    "entities": "현우의 모습은 대체로 일치하나, 미연의 머리에 있어야 할 모자와 복부 상처가 없고, 미연의 녹색 시계가 현우의 왼쪽 손목에 위치하는 오류가 있습니다.",
    "hard_violations": [],
    "physics": "현우가 미연의 상체를 안고 컨테이너 위에 앉아 있어 자세의 물리적 지지는 정상적이나, 죽은 미연의 늘어진 느낌보다는 다소 뻣뻣하게 안겨 있는 형태입니다."
   },
   {
    "label": "B",
    "direction": "현우의 시선과 입이 위쪽 새벽하늘을 향해 정확히 열려 절규하고 있습니다.",
    "built_space": "화면 하단에 두 사람을 지지하는 컨테이너 패널이 있으며, 뒤로 바다와 새벽하늘이 적절한 비율로 펼쳐져 있습니다.",
    "entities": "현우의 얼굴 상처와 회색 셔츠, 미연의 청색 모자, 앞치마, 복부의 핏자국, 왼쪽 손목의 녹색 시계가 레퍼런스와 정확히 일치합니다.",
    "hard_violations": [],
    "physics": "현우가 패널 위에 앉아 미연을 단단히 끌어안고 있으며, 미연의 양팔은 중력에 순응해 완전히 축 늘어져 지지 조건과 무활력 상태를 충족합니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미연의 모자, 복부 출혈 자국, 시계 위치 등을 정확히 재현했으며, 축 늘어진 자세와 절규하는 현우의 감정 표현이 지시문에 완벽히 부합합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "절규하는 표정은 좋으나 미연의 지정된 모자와 복부 상처가 누락되었고, 미연의 시계가 현우의 손목에 잘못 위치해 레퍼런스 고정 조건을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선과 입이 위쪽 새벽하늘을 향해 정확히 열려 절규하고 있습니다.",
        "built_space": "화면 하단에 두 사람을 지지하는 컨테이너 패널이 있으며, 뒤로 바다와 새벽하늘이 적절한 비율로 펼쳐져 있습니다.",
        "entities": "현우의 얼굴 상처와 회색 셔츠, 미연의 청색 모자, 앞치마, 복부의 핏자국, 왼쪽 손목의 녹색 시계가 레퍼런스와 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "현우가 패널 위에 앉아 미연을 단단히 끌어안고 있으며, 미연의 양팔은 중력에 순응해 완전히 축 늘어져 지지 조건과 무활력 상태를 충족합니다."
       },
       {
        "label": "A",
        "direction": "현우가 위를 올려다보며 새벽하늘을 향해 소리치고 있습니다.",
        "built_space": "하단 좌측에 컨테이너 패널이 보이고 주변으로 바다와 여명의 하늘이 둘러싸고 있습니다.",
        "entities": "현우의 모습은 대체로 일치하나, 미연의 머리에 있어야 할 모자와 복부 상처가 없고, 미연의 녹색 시계가 현우의 왼쪽 손목에 위치하는 오류가 있습니다.",
        "hard_violations": [],
        "physics": "현우가 미연의 상체를 안고 컨테이너 위에 앉아 있어 자세의 물리적 지지는 정상적이나, 죽은 미연의 늘어진 느낌보다는 다소 뻣뻣하게 안겨 있는 형태입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미연의 모자, 복부 출혈 자국, 시계 위치 등을 정확히 재현했으며, 축 늘어진 자세와 절규하는 현우의 감정 표현이 지시문에 완벽히 부합합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "절규하는 표정은 좋으나 미연의 지정된 모자와 복부 상처가 누락되었고, 미연의 시계가 현우의 손목에 잘못 위치해 레퍼런스 고정 조건을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선과 입이 위쪽 새벽하늘을 향해 정확히 열려 절규하고 있습니다.",
        "built_space": "화면 하단에 두 사람을 지지하는 컨테이너 패널이 있으며, 뒤로 바다와 새벽하늘이 적절한 비율로 펼쳐져 있습니다.",
        "entities": "현우의 얼굴 상처와 회색 셔츠, 미연의 청색 모자, 앞치마, 복부의 핏자국, 왼쪽 손목의 녹색 시계가 레퍼런스와 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "현우가 패널 위에 앉아 미연을 단단히 끌어안고 있으며, 미연의 양팔은 중력에 순응해 완전히 축 늘어져 지지 조건과 무활력 상태를 충족합니다."
       },
       {
        "label": "A",
        "direction": "현우가 위를 올려다보며 새벽하늘을 향해 소리치고 있습니다.",
        "built_space": "하단 좌측에 컨테이너 패널이 보이고 주변으로 바다와 여명의 하늘이 둘러싸고 있습니다.",
        "entities": "현우의 모습은 대체로 일치하나, 미연의 머리에 있어야 할 모자와 복부 상처가 없고, 미연의 녹색 시계가 현우의 왼쪽 손목에 위치하는 오류가 있습니다.",
        "hard_violations": [],
        "physics": "현우가 미연의 상체를 안고 컨테이너 위에 앉아 있어 자세의 물리적 지지는 정상적이나, 죽은 미연의 늘어진 느낌보다는 다소 뻣뻣하게 안겨 있는 형태입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "미연의 모자와 포옹은 잘 유지했지만, 현우의 절규가 하늘보다 정면을 향하고 허리와 무릎까지 담아 얼굴 클로즈업 지시에서 크게 벗어납니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "하늘을 향한 현우의 시선과 더 크게 잡힌 얼굴이 핵심 순간에 가깝지만, 여전히 구도가 넓고 미연의 고정된 모자가 빠져 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 눈을 거의 감고 입을 크게 벌렸으며 턱을 조금 들었지만, 얼굴과 입의 방향은 주로 카메라 정면입니다. 새벽하늘을 향한 절규라는 방향성은 약합니다. 미연은 눈을 감고 머리를 화면 왼쪽으로 늘어뜨렸습니다.",
        "built_space": "두 사람 아래에 녹슨 컨테이너 패널 하나가 있고 좌우에 모서리 금속 부속이 하나씩 보입니다. 현우는 미연 뒤에 앉아 양 무릎 사이로 몸을 받칩니다. 패널 윗면과 바다가 하단의 상당 부분을 차지해, 작은 패널 조각과 주변부 바다만 보여야 한다는 구도보다 넓습니다. 침수 정착지의 구조물은 식별되지 않으며, 새벽빛은 참고의 주황빛보다 차갑습니다.",
        "entities": "인물은 현우와 미연 두 명뿐입니다. 현우는 앳된 동아시아계 남성으로 검은 흐트러진 머리, 회색 셔츠, 얼굴 멍이 참고와 대체로 맞습니다. 미연은 중년 동아시아계 여성으로 검은 머리, 모자, 체크 셔츠, 갈색 앞치마와 감긴 눈을 유지합니다. 복부 앞치마에 혈흔은 보이지만 얼굴 부종은 뚜렷하지 않습니다. 다리 상처와 신발 속 카드는 프레임 밖이므로 판단하지 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "현우의 앉은 몸은 패널에 놓이고, 미연의 몸통은 현우의 가슴과 무릎 및 가슴을 감싼 팔에 지지됩니다. 미연의 머리는 옆으로 처져 어깨 쪽에 기대고 양팔은 아래로 떨어집니다. 보이는 범위에서 사체가 스스로 팔다리를 들거나 지지 없이 떠 있는 부분은 없습니다."
       },
       {
        "label": "B",
        "direction": "현우는 턱을 들고 눈과 얼굴을 화면 왼쪽 위의 새벽하늘로 향한 채 입을 벌려 울부짖습니다. 절규의 목표 방향이 A보다 명확합니다. 미연은 눈을 감고 얼굴이 위쪽으로 기울어진 채 현우 품에 누워 있습니다.",
        "built_space": "녹슨 컨테이너 패널 하나와 왼쪽의 모서리 부속 하나가 보입니다. 현우는 패널에 앉아 미연을 무릎 위에 비스듬히 눕혀 안습니다. 얼굴은 A보다 크게 보이지만 미연의 긴 몸통, 현우의 무릎, 넓은 패널 윗면까지 포함되어 요청한 얼굴 클로즈업은 아닙니다. 바다도 주변의 좁은 띠보다 넓게 보이며, 참고 장소의 침수 구조물은 식별되지 않습니다. 부드러운 새벽 조명은 유지하지만 참고보다 푸른 색조입니다.",
        "entities": "현우와 미연 두 명만 보입니다. 현우의 젊은 동아시아계 얼굴, 검은 머리, 회색 셔츠와 얼굴 멍은 참고에 대체로 부합합니다. 미연의 중년 동아시아계 얼굴, 검은 머리, 체크 셔츠, 갈색 앞치마와 손목시계는 대응하지만 참고와 이전 장면의 모자가 없습니다. 눈은 감겨 있고 뺨에 변색이 보이지만 부종은 약하며, 보이는 복부 앞치마에서 출혈은 뚜렷하지 않습니다. 신발 속 카드와 다리 상처는 확인할 수 없는 구도입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "현우의 몸은 패널에 앉아 지지되고, 미연의 몸통은 그의 무릎과 품에 걸쳐 있습니다. 현우의 손이 미연의 위팔을 감싸고 팔이 어깨 뒤로 이어져 상체를 받치는 배치입니다. 미연의 머리는 뒤로 기울어지고 머리카락은 아래로 떨어지며, 한 손은 복부 위에 힘없이 놓여 있습니다. 보이는 자세만으로 지지 없는 공중 부양이나 명백히 불가능한 신체 배치는 확인되지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "미연의 모자와 포옹은 잘 유지했지만, 현우의 절규가 하늘보다 정면을 향하고 허리와 무릎까지 담아 얼굴 클로즈업 지시에서 크게 벗어납니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "하늘을 향한 현우의 시선과 더 크게 잡힌 얼굴이 핵심 순간에 가깝지만, 여전히 구도가 넓고 미연의 고정된 모자가 빠져 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 눈을 거의 감고 입을 크게 벌렸으며 턱을 조금 들었지만, 얼굴과 입의 방향은 주로 카메라 정면입니다. 새벽하늘을 향한 절규라는 방향성은 약합니다. 미연은 눈을 감고 머리를 화면 왼쪽으로 늘어뜨렸습니다.",
        "built_space": "두 사람 아래에 녹슨 컨테이너 패널 하나가 있고 좌우에 모서리 금속 부속이 하나씩 보입니다. 현우는 미연 뒤에 앉아 양 무릎 사이로 몸을 받칩니다. 패널 윗면과 바다가 하단의 상당 부분을 차지해, 작은 패널 조각과 주변부 바다만 보여야 한다는 구도보다 넓습니다. 침수 정착지의 구조물은 식별되지 않으며, 새벽빛은 참고의 주황빛보다 차갑습니다.",
        "entities": "인물은 현우와 미연 두 명뿐입니다. 현우는 앳된 동아시아계 남성으로 검은 흐트러진 머리, 회색 셔츠, 얼굴 멍이 참고와 대체로 맞습니다. 미연은 중년 동아시아계 여성으로 검은 머리, 모자, 체크 셔츠, 갈색 앞치마와 감긴 눈을 유지합니다. 복부 앞치마에 혈흔은 보이지만 얼굴 부종은 뚜렷하지 않습니다. 다리 상처와 신발 속 카드는 프레임 밖이므로 판단하지 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "현우의 앉은 몸은 패널에 놓이고, 미연의 몸통은 현우의 가슴과 무릎 및 가슴을 감싼 팔에 지지됩니다. 미연의 머리는 옆으로 처져 어깨 쪽에 기대고 양팔은 아래로 떨어집니다. 보이는 범위에서 사체가 스스로 팔다리를 들거나 지지 없이 떠 있는 부분은 없습니다."
       },
       {
        "label": "A",
        "direction": "현우는 턱을 들고 눈과 얼굴을 화면 왼쪽 위의 새벽하늘로 향한 채 입을 벌려 울부짖습니다. 절규의 목표 방향이 A보다 명확합니다. 미연은 눈을 감고 얼굴이 위쪽으로 기울어진 채 현우 품에 누워 있습니다.",
        "built_space": "녹슨 컨테이너 패널 하나와 왼쪽의 모서리 부속 하나가 보입니다. 현우는 패널에 앉아 미연을 무릎 위에 비스듬히 눕혀 안습니다. 얼굴은 A보다 크게 보이지만 미연의 긴 몸통, 현우의 무릎, 넓은 패널 윗면까지 포함되어 요청한 얼굴 클로즈업은 아닙니다. 바다도 주변의 좁은 띠보다 넓게 보이며, 참고 장소의 침수 구조물은 식별되지 않습니다. 부드러운 새벽 조명은 유지하지만 참고보다 푸른 색조입니다.",
        "entities": "현우와 미연 두 명만 보입니다. 현우의 젊은 동아시아계 얼굴, 검은 머리, 회색 셔츠와 얼굴 멍은 참고에 대체로 부합합니다. 미연의 중년 동아시아계 얼굴, 검은 머리, 체크 셔츠, 갈색 앞치마와 손목시계는 대응하지만 참고와 이전 장면의 모자가 없습니다. 눈은 감겨 있고 뺨에 변색이 보이지만 부종은 약하며, 보이는 복부 앞치마에서 출혈은 뚜렷하지 않습니다. 신발 속 카드와 다리 상처는 확인할 수 없는 구도입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "현우의 몸은 패널에 앉아 지지되고, 미연의 몸통은 그의 무릎과 품에 걸쳐 있습니다. 현우의 손이 미연의 위팔을 감싸고 팔이 어깨 뒤로 이어져 상체를 받치는 배치입니다. 미연의 머리는 뒤로 기울어지고 머리카락은 아래로 떨어지며, 한 손은 복부 위에 힘없이 놓여 있습니다. 보이는 자세만으로 지지 없는 공중 부양이나 명백히 불가능한 신체 배치는 확인되지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.833
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.833
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1833,
   "A": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1833,
    "verdict_ko": "미연의 모자, 복부 출혈 자국, 시계 위치 등을 정확히 재현했으며, 축 늘어진 자세와 절규하는 현우의 감정 표현이 지시문에 완벽히 부합합니다."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "절규하는 표정은 좋으나 미연의 지정된 모자와 복부 상처가 누락되었고, 미연의 시계가 현우의 손목에 잘못 위치해 레퍼런스 고정 조건을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2_sel.png",
    "asset_id": "88b0fadf-f97d-43e5-ad24-21e20dd67b05",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1140853>",
    "asset_id": "74c747e9-e2b0-4b77-9fb2-2d2691633179",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-6897-75f2-aec2-1fe4419d74d9",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S38sh2"
  }
 },
 "S38sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:41:19.660010+00:00",
  "fingerprint": "d9e7df7350880c066378fa32e27e60e5724332bf1b7bb2679b3ef0ef74b3036d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S38sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S38sh7_sel.png",
  "source_sha256": "182b16686c3bd9dd14e4c01c2836e00c2e77906fda730fc052d465d3cf663902",
  "file": "S38sh7_cine.png",
  "staged_sha256": "c80686a93e38a82e3c7927335e3b3bf322aac5365e23c3b341f87afff91ff4dc",
  "latency_ms": 9148
 },
 "S38sh13::signage": {
  "fp": "f4fe3ea44c2a1244",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S38sh13": {
  "input_fingerprint": "9d4ab08a9a36a7b6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 오열하는 앰버를 굽어보며 어깨를 축 늘어뜨린 채 슬픈 표정을 짓는 찰리의 낡은 금속 상체.\n\nLOCATION (lock): On floating container wreckage beside the family's raft-like panel, above the inundated refugee settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Container top (Supporting 찰리 and 앰버 after their escape) — A limited portion of the upper supporting face is visible beneath the figures; used as Connects their different frame heights within the same physical space; Sea (Surrounding the containers); used as Peripheral background preserves the aftermath setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the early-dawn illumination subdued and continuous, with the final fade treated as an editorial darkening rather than a change of physical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same floating panel, immediately surrounding floodwater, and dawn light. Exclude the separate enclosed container interior and its storage boxes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The floating containers and container panel have drawn close together on the dawn sea. 앰버: She has climbed onto a container and reached the nearby floating panel. She is crying out in grief. 찰리: He stands on the container with a sorrowful expression, retaining his worn metal body and old coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 오열하는 앰버를 굽어보며 어깨를 축 늘어뜨린 채 슬픈 표정을 짓는 찰리의 낡은 금속 상체.\n\nLOCATION (lock): On floating container wreckage beside the family's raft-like panel, above the inundated refugee settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Container top (Supporting 찰리 and 앰버 after their escape) — A limited portion of the upper supporting face is visible beneath the figures; used as Connects their different frame heights within the same physical space; Sea (Surrounding the containers); used as Peripheral background preserves the aftermath setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the early-dawn illumination subdued and continuous, with the final fade treated as an editorial darkening rather than a change of physical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same floating panel, immediately surrounding floodwater, and dawn light. Exclude the separate enclosed container interior and its storage boxes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The floating containers and container panel have drawn close together on the dawn sea. 앰버: She has climbed onto a container and reached the nearby floating panel. She is crying out in grief. 찰리: He stands on the container with a sorrowful expression, retaining his worn metal body and old coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 오열하는 앰버를 굽어보며 어깨를 축 늘어뜨린 채 슬픈 표정을 짓는 찰리의 낡은 금속 상체.\n\nLOCATION (lock): On floating container wreckage beside the family's raft-like panel, above the inundated refugee settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Container top (Supporting 찰리 and 앰버 after their escape) — A limited portion of the upper supporting face is visible beneath the figures; used as Connects their different frame heights within the same physical space; Sea (Surrounding the containers); used as Peripheral background preserves the aftermath setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the early-dawn illumination subdued and continuous, with the final fade treated as an editorial darkening rather than a change of physical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same floating panel, immediately surrounding floodwater, and dawn light. Exclude the separate enclosed container interior and its storage boxes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The floating containers and container panel have drawn close together on the dawn sea. 앰버: She has climbed onto a container and reached the nearby floating panel. She is crying out in grief. 찰리: He stands on the container with a sorrowful expression, retaining his worn metal body and old coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 고개를 숙여 아래에 있는 앰버를 향해 시선을 두고 있으며, 앰버는 위를 향해 울부짖고 있다.",
    "built_space": "바다 위에 떠 있는 두 개의 컨테이너 상단이 보이며, 뒤로는 빈 바다와 수평선이 배경을 채우고 있다.",
    "entities": "찰리는 레퍼런스에 명시된 로봇 외형, 코트, 모자를 정확히 착용하고 있으며, 앰버 또한 금발 소녀의 외형을 잘 반영하고 있다.",
    "hard_violations": [],
    "physics": "찰리는 컨테이너 가장자리에 발을 딛고 서서 체중을 지탱하고 있으며, 앰버는 컨테이너 사이 틈새 공간에서 팔을 지붕 표면에 얹고 몸을 지지하고 있다."
   },
   {
    "label": "B",
    "direction": "찰리는 바닥에 엎드린 앰버를 내려다보고 있고, 앰버는 바닥에 얼굴을 묻은 채 시선을 아래로 향하고 있다.",
    "built_space": "넓은 컨테이너 상단 표면이 프레임을 채우고 있으며, 배경에는 이전 숏에 존재하지 않았던 해상 판자촌 구조물들이 추가되어 있다.",
    "entities": "찰리의 외형은 레퍼런스와 일치하나, 앰버는 금발 소녀의 머리를 하고 있음에도 이전 숏에 등장한 성인 여성의 의상을 입고 있다.",
    "hard_violations": [
     "[gemini-pro] 이전 숏 레퍼런스의 금지된 요소(성인 여성의 체크무늬 셔츠와 갈색 앞치마)가 앰버의 캐릭터에 누출됨",
     "[gpt-high] 왼쪽 컨테이너 문에 식별 가능한 문자·숫자 표기가 있어 읽을 수 있는 글자를 금지한 지시를 위반한다."
    ],
    "physics": "찰리는 양손과 두 발로 컨테이너 표면을 짚고 엎드려 자세를 지탱하고 있으며, 앰버는 표면 위에 몸을 완전히 눕혀 체중을 싣고 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 찰리의 축 처진 어깨와 앰버를 굽어보는 앵글을 훌륭하게 구현했으며, 레퍼런스의 금지 사항들을 모두 충실히 준수했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 숏 레퍼런스에 등장했던 인물의 특정 의상을 앰버에게 그대로 적용하여 지시사항을 정면으로 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개를 숙여 아래에 있는 앰버를 향해 시선을 두고 있으며, 앰버는 위를 향해 울부짖고 있다.",
        "built_space": "바다 위에 떠 있는 두 개의 컨테이너 상단이 보이며, 뒤로는 빈 바다와 수평선이 배경을 채우고 있다.",
        "entities": "찰리는 레퍼런스에 명시된 로봇 외형, 코트, 모자를 정확히 착용하고 있으며, 앰버 또한 금발 소녀의 외형을 잘 반영하고 있다.",
        "hard_violations": [],
        "physics": "찰리는 컨테이너 가장자리에 발을 딛고 서서 체중을 지탱하고 있으며, 앰버는 컨테이너 사이 틈새 공간에서 팔을 지붕 표면에 얹고 몸을 지지하고 있다."
       },
       {
        "label": "B",
        "direction": "찰리는 바닥에 엎드린 앰버를 내려다보고 있고, 앰버는 바닥에 얼굴을 묻은 채 시선을 아래로 향하고 있다.",
        "built_space": "넓은 컨테이너 상단 표면이 프레임을 채우고 있으며, 배경에는 이전 숏에 존재하지 않았던 해상 판자촌 구조물들이 추가되어 있다.",
        "entities": "찰리의 외형은 레퍼런스와 일치하나, 앰버는 금발 소녀의 머리를 하고 있음에도 이전 숏에 등장한 성인 여성의 의상을 입고 있다.",
        "hard_violations": [
         "이전 숏 레퍼런스의 금지된 요소(성인 여성의 체크무늬 셔츠와 갈색 앞치마)가 앰버의 캐릭터에 누출됨"
        ],
        "physics": "찰리는 양손과 두 발로 컨테이너 표면을 짚고 엎드려 자세를 지탱하고 있으며, 앰버는 표면 위에 몸을 완전히 눕혀 체중을 싣고 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 찰리의 축 처진 어깨와 앰버를 굽어보는 앵글을 훌륭하게 구현했으며, 레퍼런스의 금지 사항들을 모두 충실히 준수했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 숏 레퍼런스에 등장했던 인물의 특정 의상을 앰버에게 그대로 적용하여 지시사항을 정면으로 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개를 숙여 아래에 있는 앰버를 향해 시선을 두고 있으며, 앰버는 위를 향해 울부짖고 있다.",
        "built_space": "바다 위에 떠 있는 두 개의 컨테이너 상단이 보이며, 뒤로는 빈 바다와 수평선이 배경을 채우고 있다.",
        "entities": "찰리는 레퍼런스에 명시된 로봇 외형, 코트, 모자를 정확히 착용하고 있으며, 앰버 또한 금발 소녀의 외형을 잘 반영하고 있다.",
        "hard_violations": [],
        "physics": "찰리는 컨테이너 가장자리에 발을 딛고 서서 체중을 지탱하고 있으며, 앰버는 컨테이너 사이 틈새 공간에서 팔을 지붕 표면에 얹고 몸을 지지하고 있다."
       },
       {
        "label": "B",
        "direction": "찰리는 바닥에 엎드린 앰버를 내려다보고 있고, 앰버는 바닥에 얼굴을 묻은 채 시선을 아래로 향하고 있다.",
        "built_space": "넓은 컨테이너 상단 표면이 프레임을 채우고 있으며, 배경에는 이전 숏에 존재하지 않았던 해상 판자촌 구조물들이 추가되어 있다.",
        "entities": "찰리의 외형은 레퍼런스와 일치하나, 앰버는 금발 소녀의 머리를 하고 있음에도 이전 숏에 등장한 성인 여성의 의상을 입고 있다.",
        "hard_violations": [
         "이전 숏 레퍼런스의 금지된 요소(성인 여성의 체크무늬 셔츠와 갈색 앞치마)가 앰버의 캐릭터에 누출됨"
        ],
        "physics": "찰리는 양손과 두 발로 컨테이너 표면을 짚고 엎드려 자세를 지탱하고 있으며, 앰버는 표면 위에 몸을 완전히 눕혀 체중을 싣고 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리가 앰버를 내려다보지만, 앰버의 전신과 넓은 컨테이너 지붕까지 담아 상체 중심 미디엄 숏을 벗어나며 컨테이너의 문자 표기도 노출된다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "찰리의 낡은 금속 상체와 처진 팔, 앰버를 향한 아래 시선 및 앰버의 오열을 미디엄 숏으로 담아 핵심 지시에 가장 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개를 정면 아래로 숙여 바로 앞에 웅크린 앰버 쪽을 바라본다. 앰버는 얼굴을 손과 패널 가장자리에 묻고 아래를 향하므로 눈과 입의 오열 연기는 거의 보이지 않는다. 무기나 방향을 판정할 휴대 물체는 없다.",
        "built_space": "중앙에 넓은 골판 컨테이너 상면 하나, 오른쪽 전경에 테두리가 있는 별도 패널 하나가 보인다. 좌우 중경에는 다른 컨테이너들이 있고 멀리 침수된 건물들이 있다. 찰리는 중앙 상면 뒤쪽에서 몸을 굽히고 앰버는 같은 상면 앞쪽에 누워 패널 가장자리에 얼굴을 댄다. 공간 연결은 이해되지만 상면과 주변 구조물이 너무 넓게 드러나며, 이전 사진의 단순한 바다 배경보다 복잡하다. 실내나 수납 상자는 없다.",
        "entities": "등장 개체는 찰리와 앰버 두 명뿐이다. 찰리의 흰 기계 마스크, 샌드 베이지 장갑, 육중한 긴 팔, 낡은 코트와 모자는 참조에 대체로 맞는다. 앰버는 금발의 어린 여자아이지만 얼굴이 대부분 가려져 큰 눈과 정확한 얼굴 동일성은 확인하기 어렵다. 체크 셔츠와 갈색 앞치마형 옷은 앰버 참조의 남색 상의 대신 이전 장면 여성의 의상을 옮겨온 모습에 가깝다. 왼쪽 컨테이너 문에는 흰 문자와 숫자 표기가 노출된다.",
        "hard_violations": [
         "왼쪽 컨테이너 문에 식별 가능한 문자·숫자 표기가 있어 읽을 수 있는 글자를 금지한 지시를 위반한다."
        ],
        "physics": "앰버의 옆구리와 굽힌 다리는 컨테이너 상면에 놓여 있고 손과 얼굴은 패널 가장자리에 닿아 있어 지지가 분명하다. 찰리의 발은 앰버와 몸통에 가려져 접촉점을 확인할 수 없지만, 굽힌 다리는 같은 상면으로 이어지는 자세다. 공중에 떠 있는 신체나 손 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리의 머리와 마스크는 화면 오른쪽 아래의 앰버 얼굴을 향한다. 앰버는 눈을 감고 입을 크게 벌려 울며 얼굴을 약간 위로 든다. 따라서 내려다보는 찰리와 오열하는 앰버의 방향 관계가 명확하다. 무기나 방향성 있는 휴대 물체는 없다.",
        "built_space": "전경의 녹슨 골판 상면 하나와 오른쪽의 테두리 있는 패널 하나가 보이며, 뒤쪽은 거의 전부 바다다. 두 상면의 가장자리가 인물 하체를 가린다. 찰리는 왼쪽의 높은 위치, 앰버는 오른쪽의 낮은 위치에 배치되어 높이 차이가 읽힌다. 실제 발판 접촉부는 가려져 있지만 구조물의 중복이나 불가능한 반사는 없다. 녹슨 금속, 패널 테두리, 수평선과 새벽 바다가 이전 장소의 특징을 잘 유지한다.",
        "entities": "찰리와 앰버만 등장한다. 찰리의 흰 각진 마스크, 베이지 장갑판, 가슴의 원형 부품, 긴 팔, 낡은 코트와 모자는 참조와 잘 맞는다. 금발과 둥근 얼굴의 어린 여자아이는 앰버의 연령대와 주요 외형에 부합하지만, 눈을 감고 울어 눈의 형태는 비교하기 어렵다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 앰버의 겉옷은 참조의 단순한 남색 상의와 차이가 있다. 읽을 수 있는 글자나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 팔은 어깨 관절에서 자연스럽게 아래로 늘어지고 상체는 앰버 쪽으로 기울어 있다. 두 인물의 하체와 발판 접촉점은 전경 금속면 뒤에 가려져 직접 확인되지 않는다. 상체만으로 무지지 부유나 불가능한 자세라고 판단할 근거는 없다. 모자는 머리에 놓여 있고 코트는 몸에서 자연스럽게 늘어진다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리가 앰버를 내려다보지만, 앰버의 전신과 넓은 컨테이너 지붕까지 담아 상체 중심 미디엄 숏을 벗어나며 컨테이너의 문자 표기도 노출된다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찰리의 낡은 금속 상체와 처진 팔, 앰버를 향한 아래 시선 및 앰버의 오열을 미디엄 숏으로 담아 핵심 지시에 가장 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 고개를 정면 아래로 숙여 바로 앞에 웅크린 앰버 쪽을 바라본다. 앰버는 얼굴을 손과 패널 가장자리에 묻고 아래를 향하므로 눈과 입의 오열 연기는 거의 보이지 않는다. 무기나 방향을 판정할 휴대 물체는 없다.",
        "built_space": "중앙에 넓은 골판 컨테이너 상면 하나, 오른쪽 전경에 테두리가 있는 별도 패널 하나가 보인다. 좌우 중경에는 다른 컨테이너들이 있고 멀리 침수된 건물들이 있다. 찰리는 중앙 상면 뒤쪽에서 몸을 굽히고 앰버는 같은 상면 앞쪽에 누워 패널 가장자리에 얼굴을 댄다. 공간 연결은 이해되지만 상면과 주변 구조물이 너무 넓게 드러나며, 이전 사진의 단순한 바다 배경보다 복잡하다. 실내나 수납 상자는 없다.",
        "entities": "등장 개체는 찰리와 앰버 두 명뿐이다. 찰리의 흰 기계 마스크, 샌드 베이지 장갑, 육중한 긴 팔, 낡은 코트와 모자는 참조에 대체로 맞는다. 앰버는 금발의 어린 여자아이지만 얼굴이 대부분 가려져 큰 눈과 정확한 얼굴 동일성은 확인하기 어렵다. 체크 셔츠와 갈색 앞치마형 옷은 앰버 참조의 남색 상의 대신 이전 장면 여성의 의상을 옮겨온 모습에 가깝다. 왼쪽 컨테이너 문에는 흰 문자와 숫자 표기가 노출된다.",
        "hard_violations": [
         "왼쪽 컨테이너 문에 식별 가능한 문자·숫자 표기가 있어 읽을 수 있는 글자를 금지한 지시를 위반한다."
        ],
        "physics": "앰버의 옆구리와 굽힌 다리는 컨테이너 상면에 놓여 있고 손과 얼굴은 패널 가장자리에 닿아 있어 지지가 분명하다. 찰리의 발은 앰버와 몸통에 가려져 접촉점을 확인할 수 없지만, 굽힌 다리는 같은 상면으로 이어지는 자세다. 공중에 떠 있는 신체나 손 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리의 머리와 마스크는 화면 오른쪽 아래의 앰버 얼굴을 향한다. 앰버는 눈을 감고 입을 크게 벌려 울며 얼굴을 약간 위로 든다. 따라서 내려다보는 찰리와 오열하는 앰버의 방향 관계가 명확하다. 무기나 방향성 있는 휴대 물체는 없다.",
        "built_space": "전경의 녹슨 골판 상면 하나와 오른쪽의 테두리 있는 패널 하나가 보이며, 뒤쪽은 거의 전부 바다다. 두 상면의 가장자리가 인물 하체를 가린다. 찰리는 왼쪽의 높은 위치, 앰버는 오른쪽의 낮은 위치에 배치되어 높이 차이가 읽힌다. 실제 발판 접촉부는 가려져 있지만 구조물의 중복이나 불가능한 반사는 없다. 녹슨 금속, 패널 테두리, 수평선과 새벽 바다가 이전 장소의 특징을 잘 유지한다.",
        "entities": "찰리와 앰버만 등장한다. 찰리의 흰 각진 마스크, 베이지 장갑판, 가슴의 원형 부품, 긴 팔, 낡은 코트와 모자는 참조와 잘 맞는다. 금발과 둥근 얼굴의 어린 여자아이는 앰버의 연령대와 주요 외형에 부합하지만, 눈을 감고 울어 눈의 형태는 비교하기 어렵다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 앰버의 겉옷은 참조의 단순한 남색 상의와 차이가 있다. 읽을 수 있는 글자나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 팔은 어깨 관절에서 자연스럽게 아래로 늘어지고 상체는 앰버 쪽으로 기울어 있다. 두 인물의 하체와 발판 접촉점은 전경 금속면 뒤에 가려져 직접 확인되지 않는다. 상체만으로 무지지 부유나 불가능한 자세라고 판단할 근거는 없다. 모자는 머리에 놓여 있고 코트는 몸에서 자연스럽게 늘어진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.875
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.625
   },
   "violations": {
    "B": [
     "[gemini-pro] 이전 숏 레퍼런스의 금지된 요소(성인 여성의 체크무늬 셔츠와 갈색 앞치마)가 앰버의 캐릭터에 누출됨",
     "[gpt-high] 왼쪽 컨테이너 문에 식별 가능한 문자·숫자 표기가 있어 읽을 수 있는 글자를 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 625
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 찰리의 축 처진 어깨와 앰버를 굽어보는 앵글을 훌륭하게 구현했으며, 레퍼런스의 금지 사항들을 모두 충실히 준수했습니다."
   },
   {
    "label": "B",
    "score": 625,
    "verdict_ko": "이전 숏 레퍼런스에 등장했던 인물의 특정 의상을 앰버에게 그대로 적용하여 지시사항을 정면으로 위반했습니다.  ★위반: [gemini-pro] 이전 숏 레퍼런스의 금지된 요소(성인 여성의 체크무늬 셔츠와 갈색 앞치마)가 앰버의 캐릭터에 누출됨 / [gpt-high] 왼쪽 컨테이너 문에 식별 가능한 문자·숫자 표기가 있어 읽을 수 있는 글자를 금지한 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh7_sel.png",
    "asset_id": "2d4ed774-9a0b-4232-9e60-89a67937b101",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-6a50-7203-9d11-535d7d0ea334",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S38sh7"
  },
  "staged_characters_added": [
   "C03"
  ]
 },
 "S38sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:42:45.311909+00:00",
  "fingerprint": "2dbe0de09efc8cd6daff6fd0e93e5a8f00c95ed33b8f87b8495541ae49be7f02",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S38sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S38sh13_sel.png",
  "source_sha256": "554c9f9b2909fac820d59f4278b9ba2a9a9a62e5eaeeb2323f6d3b39d813abc4",
  "file": "S38sh13_cine.png",
  "staged_sha256": "8ccf223e8a94aa1626e2cbba3636ef52caebfc603f82ba5859d8860b7c9aa5d2",
  "latency_ms": 9803
 },
 "S39sh2::signage": {
  "fp": "1dda73e67ecab9ed",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::flood_search_water": {
  "input_fingerprint": "299a7dbb0b706ee5",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "flood_search_water",
    "tags": [
     "S39sh2",
     "S39sh6"
    ]
   },
   "context_sig": "a97b192acf54dd12"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 수몰된 난민촌의 풍경. 페드로와 구도환 등도 둥둥 떠다니고\n- 고무보트를 타고 수몰된 난민촌을 가로지르는 민병대원들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 수몰된 난민촌의 풍경. 페드로와 구도환 등도 둥둥 떠다니고\n- 고무보트를 타고 수몰된 난민촌을 가로지르는 민병대원들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_search_water_8a2584.png",
  "asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da",
  "input_asset_ids": [
   "019a6799-a901-4c35-8fb2-1592f569dde6"
  ],
  "origin_tag": "S39sh2",
  "place_text": "At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.",
  "origin_inputs": {
   "place_text": "At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.",
   "time_of_day_en": "day",
   "conti_asset_id": "019a6799-a901-4c35-8fb2-1592f569dde6"
  }
 },
 "S39sh2::bgfirst_bg": {
  "input_fingerprint": "ac709eafdaaff59f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2__bgfirst_bg.png",
  "asset_id": "22172d82-3681-4cc3-9212-7d3b5566ad92",
  "input_asset_ids": [
   "019a6799-a901-4c35-8fb2-1592f569dde6",
   "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da"
  ]
 },
 "S39sh2": {
  "input_fingerprint": "af61c1fb041d18ba",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The refugee settlement is submerged in daylight, with containers and other wreckage afloat. 페드로: He is afloat in the flooded settlement, wet from immersion. 구도환: He is afloat in the flooded settlement, wet from immersion; his condition is not yet established.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 페드로와 구도환 right now, so 페드로와 구도환's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 페드로와 구도환: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The refugee settlement is submerged in daylight, with containers and other wreckage afloat. 페드로: He is afloat in the flooded settlement, wet from immersion. 구도환: He is afloat in the flooded settlement, wet from immersion; his condition is not yet established.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 페드로와 구도환 right now, so 페드로와 구도환's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 페드로와 구도환: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The refugee settlement is submerged in daylight, with containers and other wreckage afloat. 페드로: He is afloat in the flooded settlement, wet from immersion. 구도환: He is afloat in the flooded settlement, wet from immersion; his condition is not yet established.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 페드로와 구도환 right now, so 페드로와 구도환's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 페드로와 구도환: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2__bgfirst_bg.png",
     "asset_id": "22172d82-3681-4cc3-9212-7d3b5566ad92",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S39sh2.png",
     "asset_id": "019a6799-a901-4c35-8fb2-1592f569dde6",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1023014>",
     "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_search_water_8a2584.png",
     "asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:800935>",
     "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1023014>",
     "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 사람 모두 지친 기색으로 앞쪽 수면과 허공을 향해 시선을 떨구고 있음.",
    "built_space": "좌측의 침수된 컨테이너와 우측의 수몰된 가옥, 전봇대 등 배경 요소가 위치 레퍼런스와 정확히 일치하게 배치됨.",
    "entities": "페드로의 외모와 의상은 레퍼런스와 일치하나, 구도환은 레퍼런스에 지정된 파카가 아닌 가죽 재질의 자켓을 입고 있어 의상 불일치가 있음.",
    "hard_violations": [
     "[gemini-pro] 구도환의 왼손(화면 우측)이 엄지손가락 위치가 반대인 오른손 구조로 렌더링되어 해부학적으로 물리적 불가능 상태임"
    ],
    "physics": "두 인물이 물에 떠서 널빤지에 체중을 싣고 있으며 손으로 가장자리를 꽉 쥐고 지탱하는 물리적 텐션이 잘 보이나, 구도환의 손가락 구조 오류가 있음."
   },
   {
    "label": "B",
    "direction": "두 사람 모두 피로한 표정으로 멍하니 정면 방향을 응시하고 있음.",
    "built_space": "위치 레퍼런스에 등장하는 좌측 컨테이너, 떠다니는 드럼통, 우측 가옥 및 전봇대가 완벽하게 배경으로 구현됨.",
    "entities": "페드로의 앳된 외모와 의상은 잘 맞으며 얼굴에 묻은 진흙이 상황을 잘 보여줌. 구도환은 레퍼런스 파카 안쪽의 이너웨어와 유사한 젖은 스웨터를 입고 있음.",
    "hard_violations": [
     "[gemini-pro] 페드로의 왼손(화면 좌측) 엄지손가락 위치가 반대로 렌더링되어 다른 사람의 오른손처럼 보이는 물리적/해부학적 불가 오류가 있음"
    ],
    "physics": "물에 떠 있는 부력 상태는 자연스러우나, 페드로의 왼손이 널빤지를 움켜쥐지 않고 손가락을 펴서 얹어두기만 하여 매달려 버티는 행동 묘사에 어긋나며 손의 형태도 불가능함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "구도환의 겉옷 의상이 레퍼런스와 다르고 치명적인 왼손 해부학 오류가 존재하지만, 두 인물이 널빤지를 '꽉 움켜쥔' 행동 묘사와 전체적인 프레이밍이 프롬프트에 더 잘 부합합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "페드로의 왼손에 치명적인 해부학적 오류가 있을 뿐만 아니라 널빤지를 움켜쥐지 않고 그저 평평하게 얹어만 두고 있어 핵심 행동 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 사람 모두 지친 기색으로 앞쪽 수면과 허공을 향해 시선을 떨구고 있음.",
        "built_space": "좌측의 침수된 컨테이너와 우측의 수몰된 가옥, 전봇대 등 배경 요소가 위치 레퍼런스와 정확히 일치하게 배치됨.",
        "entities": "페드로의 외모와 의상은 레퍼런스와 일치하나, 구도환은 레퍼런스에 지정된 파카가 아닌 가죽 재질의 자켓을 입고 있어 의상 불일치가 있음.",
        "hard_violations": [
         "구도환의 왼손(화면 우측)이 엄지손가락 위치가 반대인 오른손 구조로 렌더링되어 해부학적으로 물리적 불가능 상태임"
        ],
        "physics": "두 인물이 물에 떠서 널빤지에 체중을 싣고 있으며 손으로 가장자리를 꽉 쥐고 지탱하는 물리적 텐션이 잘 보이나, 구도환의 손가락 구조 오류가 있음."
       },
       {
        "label": "B",
        "direction": "두 사람 모두 피로한 표정으로 멍하니 정면 방향을 응시하고 있음.",
        "built_space": "위치 레퍼런스에 등장하는 좌측 컨테이너, 떠다니는 드럼통, 우측 가옥 및 전봇대가 완벽하게 배경으로 구현됨.",
        "entities": "페드로의 앳된 외모와 의상은 잘 맞으며 얼굴에 묻은 진흙이 상황을 잘 보여줌. 구도환은 레퍼런스 파카 안쪽의 이너웨어와 유사한 젖은 스웨터를 입고 있음.",
        "hard_violations": [
         "페드로의 왼손(화면 좌측) 엄지손가락 위치가 반대로 렌더링되어 다른 사람의 오른손처럼 보이는 물리적/해부학적 불가 오류가 있음"
        ],
        "physics": "물에 떠 있는 부력 상태는 자연스러우나, 페드로의 왼손이 널빤지를 움켜쥐지 않고 손가락을 펴서 얹어두기만 하여 매달려 버티는 행동 묘사에 어긋나며 손의 형태도 불가능함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "구도환의 겉옷 의상이 레퍼런스와 다르고 치명적인 왼손 해부학 오류가 존재하지만, 두 인물이 널빤지를 '꽉 움켜쥔' 행동 묘사와 전체적인 프레이밍이 프롬프트에 더 잘 부합합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "페드로의 왼손에 치명적인 해부학적 오류가 있을 뿐만 아니라 널빤지를 움켜쥐지 않고 그저 평평하게 얹어만 두고 있어 핵심 행동 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "두 사람 모두 지친 기색으로 앞쪽 수면과 허공을 향해 시선을 떨구고 있음.",
        "built_space": "좌측의 침수된 컨테이너와 우측의 수몰된 가옥, 전봇대 등 배경 요소가 위치 레퍼런스와 정확히 일치하게 배치됨.",
        "entities": "페드로의 외모와 의상은 레퍼런스와 일치하나, 구도환은 레퍼런스에 지정된 파카가 아닌 가죽 재질의 자켓을 입고 있어 의상 불일치가 있음.",
        "hard_violations": [
         "구도환의 왼손(화면 우측)이 엄지손가락 위치가 반대인 오른손 구조로 렌더링되어 해부학적으로 물리적 불가능 상태임"
        ],
        "physics": "두 인물이 물에 떠서 널빤지에 체중을 싣고 있으며 손으로 가장자리를 꽉 쥐고 지탱하는 물리적 텐션이 잘 보이나, 구도환의 손가락 구조 오류가 있음."
       },
       {
        "label": "B",
        "direction": "두 사람 모두 피로한 표정으로 멍하니 정면 방향을 응시하고 있음.",
        "built_space": "위치 레퍼런스에 등장하는 좌측 컨테이너, 떠다니는 드럼통, 우측 가옥 및 전봇대가 완벽하게 배경으로 구현됨.",
        "entities": "페드로의 앳된 외모와 의상은 잘 맞으며 얼굴에 묻은 진흙이 상황을 잘 보여줌. 구도환은 레퍼런스 파카 안쪽의 이너웨어와 유사한 젖은 스웨터를 입고 있음.",
        "hard_violations": [
         "페드로의 왼손(화면 좌측) 엄지손가락 위치가 반대로 렌더링되어 다른 사람의 오른손처럼 보이는 물리적/해부학적 불가 오류가 있음"
        ],
        "physics": "물에 떠 있는 부력 상태는 자연스러우나, 페드로의 왼손이 널빤지를 움켜쥐지 않고 손가락을 펴서 얹어두기만 하여 매달려 버티는 행동 묘사에 어긋나며 손의 형태도 불가능함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "흙탕물 속 지친 상체와 사선 널빤지의 중간 숏은 충실하지만, 두 사람의 한 손 파지 자세가 비슷하고 구도환의 외투·목 의상이 참조와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지친 상체 중심의 중간 숏에서 페드로의 한 손 파지와 구도환의 양손 파지를 사선 널빤지로 연결하며, 얼굴과 보이는 의상도 참조에 더 가깝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "페드로는 화면 오른쪽 앞쪽을, 구도환도 오른쪽 전방의 화면 밖을 바라본다. 특정 응시 대상은 보이지 않으며 본문에도 지정되지 않았다. 두 사람은 각각 보이는 한 손을 앞쪽의 같은 널빤지로 뻗어 가장자리를 감싸고 있다.",
        "built_space": "왼쪽 배경에 큰 침수 컨테이너 하나, 중앙에 기울어진 드럼통 하나, 오른쪽에 박공지붕 건물 하나와 전봇대 하나가 보이며 장소 참조의 배치와 부합한다. 작은 컨테이너와 잔해가 먼 수면에 흩어져 있다. 두 사람은 구조물 위가 아니라 물속에 있고, 공동으로 잡은 널빤지는 화면 하단 중앙을 비스듬히 가로지르며 화면 면적의 5분의 2보다 작다.",
        "entities": "인물은 두 남성뿐이다. 왼쪽 페드로는 짙은 젖은 머리와 어두운 후드 상의를 갖췄으나 참조의 앳된 얼굴보다 다소 나이 들어 보이고 검은 비니는 없다. 오른쪽 구도환은 검은 머리의 한국인 중년 남성이라는 설정에 부합하지만, 참조의 올리브색 외투와 두꺼운 목 감싸개 대신 갈색 계열 상의가 보인다. 불투명한 흙탕물, 젖은 목재 널빤지, 부유 잔해와 낮의 조명은 확인되며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통 아래쪽은 물에 잠겨 있고, 수중 부력과 손으로 잡은 부유 널빤지가 버티는 자세를 뒷받침한다. 각자의 보이는 손가락은 목재 가장자리를 감싸며 팔과 자연스럽게 이어진다. 널빤지와 주변 잔해는 수면에 접해 떠 있으며, 지지 없이 공중에 뜬 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "페드로는 눈을 낮춰 널빤지 너머 가까운 전방 수면을 보고, 구도환은 왼쪽 아래의 널빤지와 페드로 쪽으로 시선을 둔다. 페드로의 앞쪽 한 손과 구도환의 두 손이 동일한 널빤지 가장자리를 향해 뻗어 실제로 움켜쥐고 있다.",
        "built_space": "왼쪽 큰 침수 컨테이너 하나, 중앙의 기울어진 드럼통 하나, 오른쪽 박공지붕 건물 하나와 전봇대 하나가 장소 참조와 같은 상대적 위치에 있다. 먼 컨테이너와 목재 잔해는 작게 배치된다. 두 사람의 상체와 팔이 중간 숏을 채우고, 널빤지 윗면과 가까운 가장자리가 왼쪽 아래에서 오른쪽 중간으로 비스듬히 이어진다. 널빤지의 면적은 화면의 5분의 2 미만이다.",
        "entities": "인물은 페드로와 구도환에 해당하는 두 남성뿐이다. 페드로는 앳된 얼굴, 짙은 머리, 어두운 후드 상의가 참조에 가깝지만 비니는 없다. 구도환은 검은 머리의 한국인 중년 남성으로 보이고, 얼굴과 어두운 올리브색 외투가 참조에 가깝다. 참조의 두꺼운 목 감싸개는 뚜렷하게 보이지 않는다. 두 사람의 젖은 옷과 손, 불투명한 흙탕물, 실물 목재 질감의 널빤지, 부유 잔해와 주간 환경이 확인된다. 읽을 수 있는 문구나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "두 사람은 하체와 몸통 일부가 잠긴 상태에서 수면의 부력과 널빤지에 의지한다. 페드로는 한 손으로 가장자리를 감싸고 반대 팔은 수면 가까이 두며, 구도환은 양손으로 가장자리를 잡고 팔을 널빤지 위에 걸친다. 손과 목재의 접촉 및 팔의 연결이 자연스럽고, 널빤지는 수면에 닿아 있다. 지지 없는 공중 부양이나 불가능한 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "흙탕물 속 지친 상체와 사선 널빤지의 중간 숏은 충실하지만, 두 사람의 한 손 파지 자세가 비슷하고 구도환의 외투·목 의상이 참조와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지친 상체 중심의 중간 숏에서 페드로의 한 손 파지와 구도환의 양손 파지를 사선 널빤지로 연결하며, 얼굴과 보이는 의상도 참조에 더 가깝다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "페드로는 화면 오른쪽 앞쪽을, 구도환도 오른쪽 전방의 화면 밖을 바라본다. 특정 응시 대상은 보이지 않으며 본문에도 지정되지 않았다. 두 사람은 각각 보이는 한 손을 앞쪽의 같은 널빤지로 뻗어 가장자리를 감싸고 있다.",
        "built_space": "왼쪽 배경에 큰 침수 컨테이너 하나, 중앙에 기울어진 드럼통 하나, 오른쪽에 박공지붕 건물 하나와 전봇대 하나가 보이며 장소 참조의 배치와 부합한다. 작은 컨테이너와 잔해가 먼 수면에 흩어져 있다. 두 사람은 구조물 위가 아니라 물속에 있고, 공동으로 잡은 널빤지는 화면 하단 중앙을 비스듬히 가로지르며 화면 면적의 5분의 2보다 작다.",
        "entities": "인물은 두 남성뿐이다. 왼쪽 페드로는 짙은 젖은 머리와 어두운 후드 상의를 갖췄으나 참조의 앳된 얼굴보다 다소 나이 들어 보이고 검은 비니는 없다. 오른쪽 구도환은 검은 머리의 한국인 중년 남성이라는 설정에 부합하지만, 참조의 올리브색 외투와 두꺼운 목 감싸개 대신 갈색 계열 상의가 보인다. 불투명한 흙탕물, 젖은 목재 널빤지, 부유 잔해와 낮의 조명은 확인되며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통 아래쪽은 물에 잠겨 있고, 수중 부력과 손으로 잡은 부유 널빤지가 버티는 자세를 뒷받침한다. 각자의 보이는 손가락은 목재 가장자리를 감싸며 팔과 자연스럽게 이어진다. 널빤지와 주변 잔해는 수면에 접해 떠 있으며, 지지 없이 공중에 뜬 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "페드로는 눈을 낮춰 널빤지 너머 가까운 전방 수면을 보고, 구도환은 왼쪽 아래의 널빤지와 페드로 쪽으로 시선을 둔다. 페드로의 앞쪽 한 손과 구도환의 두 손이 동일한 널빤지 가장자리를 향해 뻗어 실제로 움켜쥐고 있다.",
        "built_space": "왼쪽 큰 침수 컨테이너 하나, 중앙의 기울어진 드럼통 하나, 오른쪽 박공지붕 건물 하나와 전봇대 하나가 장소 참조와 같은 상대적 위치에 있다. 먼 컨테이너와 목재 잔해는 작게 배치된다. 두 사람의 상체와 팔이 중간 숏을 채우고, 널빤지 윗면과 가까운 가장자리가 왼쪽 아래에서 오른쪽 중간으로 비스듬히 이어진다. 널빤지의 면적은 화면의 5분의 2 미만이다.",
        "entities": "인물은 페드로와 구도환에 해당하는 두 남성뿐이다. 페드로는 앳된 얼굴, 짙은 머리, 어두운 후드 상의가 참조에 가깝지만 비니는 없다. 구도환은 검은 머리의 한국인 중년 남성으로 보이고, 얼굴과 어두운 올리브색 외투가 참조에 가깝다. 참조의 두꺼운 목 감싸개는 뚜렷하게 보이지 않는다. 두 사람의 젖은 옷과 손, 불투명한 흙탕물, 실물 목재 질감의 널빤지, 부유 잔해와 주간 환경이 확인된다. 읽을 수 있는 문구나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "두 사람은 하체와 몸통 일부가 잠긴 상태에서 수면의 부력과 널빤지에 의지한다. 페드로는 한 손으로 가장자리를 감싸고 반대 팔은 수면 가까이 두며, 구도환은 양손으로 가장자리를 잡고 팔을 널빤지 위에 걸친다. 손과 목재의 접촉 및 팔의 연결이 자연스럽고, 널빤지는 수면에 닿아 있다. 지지 없는 공중 부양이나 불가능한 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.639
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.389
   },
   "violations": {
    "A": [
     "[gemini-pro] 구도환의 왼손(화면 우측)이 엄지손가락 위치가 반대인 오른손 구조로 렌더링되어 해부학적으로 물리적 불가능 상태임"
    ],
    "B": [
     "[gemini-pro] 페드로의 왼손(화면 좌측) 엄지손가락 위치가 반대로 렌더링되어 다른 사람의 오른손처럼 보이는 물리적/해부학적 불가 오류가 있음"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1389
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "구도환의 겉옷 의상이 레퍼런스와 다르고 치명적인 왼손 해부학 오류가 존재하지만, 두 인물이 널빤지를 '꽉 움켜쥔' 행동 묘사와 전체적인 프레이밍이 프롬프트에 더 잘 부합합니다.  ★위반: [gemini-pro] 구도환의 왼손(화면 우측)이 엄지손가락 위치가 반대인 오른손 구조로 렌더링되어 해부학적으로 물리적 불가능 상태임"
   },
   {
    "label": "B",
    "score": 1389,
    "verdict_ko": "페드로의 왼손에 치명적인 해부학적 오류가 있을 뿐만 아니라 널빤지를 움켜쥐지 않고 그저 평평하게 얹어만 두고 있어 핵심 행동 지시를 위반했습니다.  ★위반: [gemini-pro] 페드로의 왼손(화면 좌측) 엄지손가락 위치가 반대로 렌더링되어 다른 사람의 오른손처럼 보이는 물리적/해부학적 불가 오류가 있음"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_search_water_8a2584.png",
    "asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:800935>",
    "asset_id": "c6941a05-0ea7-4671-b0bf-b68d784cb7a7",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1023014>",
    "asset_id": "69242df4-28df-460f-8304-f67f36afffcc",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-6bfb-7ff8-979b-d7c902bc7be2",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2__bgfirst_bg.png",
   "bg_asset_id": "22172d82-3681-4cc3-9212-7d3b5566ad92",
   "bg_record_key": "S39sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "flood_search_water",
   "groupbg_asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S39sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:44:37.674834+00:00",
  "fingerprint": "91cae37f6ae5b6bf593c8f4a127b2f58f8a6915bd5afb65760d8290aa7fcc2ae",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S39sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S39sh2_sel.png",
  "source_sha256": "ec13be1eb9801f1cbbb996532ee30740d1c2cf34a03d7ac5f10cc1af0220d7cd",
  "file": "S39sh2_cine.png",
  "staged_sha256": "c5eaebf07b908c98fbd6a35214dc301151585eb2d1a978e6dcff21d5764b53ca",
  "latency_ms": 10449
 },
 "S39sh6::signage": {
  "fp": "aebc244dc47b0747",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S39sh6": {
  "input_fingerprint": "250239d3b4f69169",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 보트 끝에 비스듬히 서서, 긴 막대기 끝으로 수면 위의 시체 어깨를 꾹 누르며 힘을 가하고 있는 민병대원의 거친 mid-action 자세.\n\nLOCATION (lock): At the open bow of an inflatable patrol boat moving through the flooded refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubber boat (Carrying the militia member across the flooded settlement) — The end and adjacent outer side are visible from alongside; used as Provides a limited right-side support and scale reference for the braced posture; Long pole (Its tip presses into the floating body's shoulder) — Runs diagonally from the militia member's hands toward the lower-left contact point; used as Creates the visible line of force between the standing figure and the body; Floodwater (Surrounding the boat and floating body); used as Separates the corpse from the boat while keeping both within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral daytime ambient light maintains clear physical contact and restrained tonal contrast without extending the helicopter searchlight into this shot.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The corpse floats at the water's surface while a militia member presses a pole against its shoulder to inspect it. The source does not establish whether the body faces upward or downward, or how its head and limbs are arranged.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubber boats move through the submerged settlement among floating containers and wreckage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원 right now, so 민병대원's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 보트 끝에 비스듬히 서서, 긴 막대기 끝으로 수면 위의 시체 어깨를 꾹 누르며 힘을 가하고 있는 민병대원의 거친 mid-action 자세.\n\nLOCATION (lock): At the open bow of an inflatable patrol boat moving through the flooded refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubber boat (Carrying the militia member across the flooded settlement) — The end and adjacent outer side are visible from alongside; used as Provides a limited right-side support and scale reference for the braced posture; Long pole (Its tip presses into the floating body's shoulder) — Runs diagonally from the militia member's hands toward the lower-left contact point; used as Creates the visible line of force between the standing figure and the body; Floodwater (Surrounding the boat and floating body); used as Separates the corpse from the boat while keeping both within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral daytime ambient light maintains clear physical contact and restrained tonal contrast without extending the helicopter searchlight into this shot.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The corpse floats at the water's surface while a militia member presses a pole against its shoulder to inspect it. The source does not establish whether the body faces upward or downward, or how its head and limbs are arranged.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubber boats move through the submerged settlement among floating containers and wreckage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원 right now, so 민병대원's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 보트 끝에 비스듬히 서서, 긴 막대기 끝으로 수면 위의 시체 어깨를 꾹 누르며 힘을 가하고 있는 민병대원의 거친 mid-action 자세.\n\nLOCATION (lock): At the open bow of an inflatable patrol boat moving through the flooded refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubber boat (Carrying the militia member across the flooded settlement) — The end and adjacent outer side are visible from alongside; used as Provides a limited right-side support and scale reference for the braced posture; Long pole (Its tip presses into the floating body's shoulder) — Runs diagonally from the militia member's hands toward the lower-left contact point; used as Creates the visible line of force between the standing figure and the body; Floodwater (Surrounding the boat and floating body); used as Separates the corpse from the boat while keeping both within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral daytime ambient light maintains clear physical contact and restrained tonal contrast without extending the helicopter searchlight into this shot.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The corpse floats at the water's surface while a militia member presses a pole against its shoulder to inspect it. The source does not establish whether the body faces upward or downward, or how its head and limbs are arranged.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubber boats move through the submerged settlement among floating containers and wreckage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원 right now, so 민병대원's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "막대기 끝이 엎드린 시체의 등과 왼쪽 어깨뼈 부근을 정확히 향해 닿아 있음.",
    "built_space": "침수된 야외. 우측 고무보트 안에 남성이 서 있으며, 배경의 컨테이너와 잔해 배치는 레퍼런스의 환경과 일치함.",
    "entities": "전술 조끼를 입은 민병대원과 물에 뜬 시체. 두 인물 모두 레퍼런스 이미지에 등장하는 두 남성의 얼굴과 동일하게 생성됨.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)"
    ],
    "physics": "남성은 보트 바닥에 체중을 싣고 서 있고, 시체는 물에 자연스럽게 일부 잠긴 채 떠 있음. 막대기는 두 손에 쥐어져 물리적으로 시체를 누르고 있음."
   },
   {
    "label": "B",
    "direction": "막대기 끝이 프롬프트가 지시한 시체의 어깨가 아닌 가슴과 복부 쪽을 향하고 있음.",
    "built_space": "침수된 야외. 우측 고무보트에 남성이 서 있고, 배경의 구조물은 레퍼런스 환경을 잘 반영함.",
    "entities": "녹색 재킷의 민병대원과 하늘을 향해 떠 있는 시체. 민병대원의 얼굴이 레퍼런스의 중년 남성 얼굴과 완벽히 동일함.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)",
     "[gemini-pro] 시체의 다리와 발뒤꿈치가 수면 위로 완전히 노출된 물리적으로 불가능한 부유 자세",
     "[gemini-pro] 막대기가 시체의 몸통 표면을 누르는 대신 몸을 뚫고 들어간 물리적 오류"
    ],
    "physics": "시체가 마치 단단한 평지에 누운 듯 수면에 떠 있어 중력과 부력 표현이 완전히 어긋나며, 막대기 끝이 시체의 몸통과 비정상적으로 겹쳐짐."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "레퍼런스의 얼굴을 배제하라는 지시를 위반했으나, 막대기로 시체를 누르는 물리적 동작과 부력의 묘사가 비교적 자연스러움."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "레퍼런스 인물 복사 위반과 더불어, 시체가 평지에 누운 듯 수면에 떠 있고 막대기가 몸을 관통하는 물리적 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "막대기 끝이 엎드린 시체의 등과 왼쪽 어깨뼈 부근을 정확히 향해 닿아 있음.",
        "built_space": "침수된 야외. 우측 고무보트 안에 남성이 서 있으며, 배경의 컨테이너와 잔해 배치는 레퍼런스의 환경과 일치함.",
        "entities": "전술 조끼를 입은 민병대원과 물에 뜬 시체. 두 인물 모두 레퍼런스 이미지에 등장하는 두 남성의 얼굴과 동일하게 생성됨.",
        "hard_violations": [
         "레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)"
        ],
        "physics": "남성은 보트 바닥에 체중을 싣고 서 있고, 시체는 물에 자연스럽게 일부 잠긴 채 떠 있음. 막대기는 두 손에 쥐어져 물리적으로 시체를 누르고 있음."
       },
       {
        "label": "B",
        "direction": "막대기 끝이 프롬프트가 지시한 시체의 어깨가 아닌 가슴과 복부 쪽을 향하고 있음.",
        "built_space": "침수된 야외. 우측 고무보트에 남성이 서 있고, 배경의 구조물은 레퍼런스 환경을 잘 반영함.",
        "entities": "녹색 재킷의 민병대원과 하늘을 향해 떠 있는 시체. 민병대원의 얼굴이 레퍼런스의 중년 남성 얼굴과 완벽히 동일함.",
        "hard_violations": [
         "레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)",
         "시체의 다리와 발뒤꿈치가 수면 위로 완전히 노출된 물리적으로 불가능한 부유 자세",
         "막대기가 시체의 몸통 표면을 누르는 대신 몸을 뚫고 들어간 물리적 오류"
        ],
        "physics": "시체가 마치 단단한 평지에 누운 듯 수면에 떠 있어 중력과 부력 표현이 완전히 어긋나며, 막대기 끝이 시체의 몸통과 비정상적으로 겹쳐짐."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "레퍼런스의 얼굴을 배제하라는 지시를 위반했으나, 막대기로 시체를 누르는 물리적 동작과 부력의 묘사가 비교적 자연스러움."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "레퍼런스 인물 복사 위반과 더불어, 시체가 평지에 누운 듯 수면에 떠 있고 막대기가 몸을 관통하는 물리적 오류가 있음."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "막대기 끝이 엎드린 시체의 등과 왼쪽 어깨뼈 부근을 정확히 향해 닿아 있음.",
        "built_space": "침수된 야외. 우측 고무보트 안에 남성이 서 있으며, 배경의 컨테이너와 잔해 배치는 레퍼런스의 환경과 일치함.",
        "entities": "전술 조끼를 입은 민병대원과 물에 뜬 시체. 두 인물 모두 레퍼런스 이미지에 등장하는 두 남성의 얼굴과 동일하게 생성됨.",
        "hard_violations": [
         "레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)"
        ],
        "physics": "남성은 보트 바닥에 체중을 싣고 서 있고, 시체는 물에 자연스럽게 일부 잠긴 채 떠 있음. 막대기는 두 손에 쥐어져 물리적으로 시체를 누르고 있음."
       },
       {
        "label": "B",
        "direction": "막대기 끝이 프롬프트가 지시한 시체의 어깨가 아닌 가슴과 복부 쪽을 향하고 있음.",
        "built_space": "침수된 야외. 우측 고무보트에 남성이 서 있고, 배경의 구조물은 레퍼런스 환경을 잘 반영함.",
        "entities": "녹색 재킷의 민병대원과 하늘을 향해 떠 있는 시체. 민병대원의 얼굴이 레퍼런스의 중년 남성 얼굴과 완벽히 동일함.",
        "hard_violations": [
         "레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)",
         "시체의 다리와 발뒤꿈치가 수면 위로 완전히 노출된 물리적으로 불가능한 부유 자세",
         "막대기가 시체의 몸통 표면을 누르는 대신 몸을 뚫고 들어간 물리적 오류"
        ],
        "physics": "시체가 마치 단단한 평지에 누운 듯 수면에 떠 있어 중력과 부력 표현이 완전히 어긋나며, 막대기 끝이 시체의 몸통과 비정상적으로 겹쳐짐."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "오른쪽 보트에 몸을 지지한 민병대원의 거친 상체 동작과 왼쪽 아래 시체 어깨에 닿는 막대기 끝이 요구된 미디엄 숏과 힘의 방향을 충실히 구현한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "침수 장소와 보트 위 자세는 부합하지만, 막대기 끝이 어깨가 아닌 몸 옆 수면에 닿고 구도도 요구된 미디엄 숏보다 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "민병대원은 왼쪽 아래 시체 쪽을 보고 두 손으로 막대기를 잡는다. 막대기는 요구대로 왼쪽 아래로 뻗지만, 시체 몸통을 가로질러 끝이 화면 아래쪽 팔꿈치 부근의 수면에 닿는다. 어깨를 끝으로 누르는 접촉은 보이지 않는다. 시체는 얼굴을 위로 향하고 있다.",
        "built_space": "오른쪽에 고무보트 한 척의 둥근 선수와 긴 외측 튜브, 내부 가로 좌판 하나가 보인다. 민병대원은 선수 안쪽에 다리를 벌리고 서 있으며 하퇴는 튜브에 가려진다. 왼쪽의 큰 녹슨 컨테이너, 오른쪽의 침수된 박공지붕 건물 한 채와 전봇대 하나가 참고 장소의 배치를 잘 이어받는다. 다만 보트와 시체 전신까지 넓게 담아 오른쪽의 제한적인 지지물이라는 보트의 역할보다 존재감이 커졌다.",
        "entities": "민병대원 한 명과 시체 한 명만 보인다. 민병대원은 짧은 검은 머리의 중년 동아시아계 남성으로 보이며, 낡은 올리브색 군복과 검은 완장을 착용했다. 시체는 짧은 머리와 수염이 있는 성인 남성으로 보이고 어두운 젖은 옷을 입었다. 별도의 인물 참고나 구체적인 시체 인종 지정은 없어 외형 불일치는 확정할 수 없다. 긴 나무 막대기, 고무보트, 흙탕물, 컨테이너와 목재 잔해가 있으며 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "막대기는 양손으로 지지되며 민병대원은 보트 내부에 하체를 둔 채 앞으로 기울어 있다. 발은 가려져 있지만 공중에 뜬 자세는 아니다. 시체의 몸통과 벌어진 팔다리는 수면의 부력으로 지지되고, 위를 향한 발끝은 수면에 놓인 다리에서 이어진다. 뚜렷한 무지지 신체 부위는 없으나 막대기 끝이 어깨에 닿지 않아 요구된 누르는 힘은 전달되지 않는다."
       },
       {
        "label": "B",
        "direction": "민병대원은 왼쪽 아래의 접촉 부위를 내려다본다. 양손에서 뻗은 막대기의 끝은 엎드린 시체의 어깨 뒤쪽에 닿으며, 상체를 기울여 누르는 방향과 막대기의 대각선이 일치한다. 시체의 얼굴은 아래쪽 수면을 향한다.",
        "built_space": "오른쪽에 고무보트 한 척의 선수와 인접 외측 튜브가 보인다. 튜브에 밧줄과 부착 고리가 있으며 내부 좌판은 드러나지 않는다. 민병대원의 골반과 하체는 선수 안쪽에 놓이고 상체가 바깥으로 기울어, 보트가 오른쪽 지지물로 기능한다. 왼쪽 큰 컨테이너와 뒤쪽 침수 컨테이너들, 목재 잔해가 같은 침수 정착지의 재료와 환경을 유지한다. 참고의 지붕과 전봇대는 이 구도에 보이지 않고 컨테이너는 더 크게 잡혔다. 민병대원은 대체로 허리 위 중심의 미디엄 숏으로 제시된다.",
        "entities": "민병대원 한 명과 시체 한 명만 있다. 민병대원은 짧은 검은 머리의 중년 동아시아계 남성으로 보이며 낡은 회갈색 재킷, 올리브색 전술 조끼와 장갑을 착용했다. 시체는 짧은 검은 머리의 성인 동아시아계 남성으로 보이고 젖은 어두운 옷과 신발을 착용했다. 구체적인 인물 외형이나 복장을 고정하는 별도 참고는 없다. 긴 나무 막대기, 고무보트, 홍수 물과 떠 있는 잔해가 모두 식별되며, 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "민병대원은 골반과 하체를 보트 안쪽에 두고 튜브 가장자리에 기대며 상체를 내민다. 발은 가려져 있지만 보트에 지지되는 자세로 읽힌다. 두 장갑 낀 손이 막대기를 떨어져 잡아 어깨 접촉점으로 힘을 전달한다. 시체의 머리와 팔은 수면 및 인접 잔해 높이에 낮게 놓이고 몸통과 다리는 부력으로 지지되며, 스스로 들어 올린 팔다리는 없다. 보트와 시체 사이에는 물이 이어지고 보트 주변의 물결도 이동 상황과 양립한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "오른쪽 보트에 몸을 지지한 민병대원의 거친 상체 동작과 왼쪽 아래 시체 어깨에 닿는 막대기 끝이 요구된 미디엄 숏과 힘의 방향을 충실히 구현한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "침수 장소와 보트 위 자세는 부합하지만, 막대기 끝이 어깨가 아닌 몸 옆 수면에 닿고 구도도 요구된 미디엄 숏보다 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "민병대원은 왼쪽 아래 시체 쪽을 보고 두 손으로 막대기를 잡는다. 막대기는 요구대로 왼쪽 아래로 뻗지만, 시체 몸통을 가로질러 끝이 화면 아래쪽 팔꿈치 부근의 수면에 닿는다. 어깨를 끝으로 누르는 접촉은 보이지 않는다. 시체는 얼굴을 위로 향하고 있다.",
        "built_space": "오른쪽에 고무보트 한 척의 둥근 선수와 긴 외측 튜브, 내부 가로 좌판 하나가 보인다. 민병대원은 선수 안쪽에 다리를 벌리고 서 있으며 하퇴는 튜브에 가려진다. 왼쪽의 큰 녹슨 컨테이너, 오른쪽의 침수된 박공지붕 건물 한 채와 전봇대 하나가 참고 장소의 배치를 잘 이어받는다. 다만 보트와 시체 전신까지 넓게 담아 오른쪽의 제한적인 지지물이라는 보트의 역할보다 존재감이 커졌다.",
        "entities": "민병대원 한 명과 시체 한 명만 보인다. 민병대원은 짧은 검은 머리의 중년 동아시아계 남성으로 보이며, 낡은 올리브색 군복과 검은 완장을 착용했다. 시체는 짧은 머리와 수염이 있는 성인 남성으로 보이고 어두운 젖은 옷을 입었다. 별도의 인물 참고나 구체적인 시체 인종 지정은 없어 외형 불일치는 확정할 수 없다. 긴 나무 막대기, 고무보트, 흙탕물, 컨테이너와 목재 잔해가 있으며 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "막대기는 양손으로 지지되며 민병대원은 보트 내부에 하체를 둔 채 앞으로 기울어 있다. 발은 가려져 있지만 공중에 뜬 자세는 아니다. 시체의 몸통과 벌어진 팔다리는 수면의 부력으로 지지되고, 위를 향한 발끝은 수면에 놓인 다리에서 이어진다. 뚜렷한 무지지 신체 부위는 없으나 막대기 끝이 어깨에 닿지 않아 요구된 누르는 힘은 전달되지 않는다."
       },
       {
        "label": "A",
        "direction": "민병대원은 왼쪽 아래의 접촉 부위를 내려다본다. 양손에서 뻗은 막대기의 끝은 엎드린 시체의 어깨 뒤쪽에 닿으며, 상체를 기울여 누르는 방향과 막대기의 대각선이 일치한다. 시체의 얼굴은 아래쪽 수면을 향한다.",
        "built_space": "오른쪽에 고무보트 한 척의 선수와 인접 외측 튜브가 보인다. 튜브에 밧줄과 부착 고리가 있으며 내부 좌판은 드러나지 않는다. 민병대원의 골반과 하체는 선수 안쪽에 놓이고 상체가 바깥으로 기울어, 보트가 오른쪽 지지물로 기능한다. 왼쪽 큰 컨테이너와 뒤쪽 침수 컨테이너들, 목재 잔해가 같은 침수 정착지의 재료와 환경을 유지한다. 참고의 지붕과 전봇대는 이 구도에 보이지 않고 컨테이너는 더 크게 잡혔다. 민병대원은 대체로 허리 위 중심의 미디엄 숏으로 제시된다.",
        "entities": "민병대원 한 명과 시체 한 명만 있다. 민병대원은 짧은 검은 머리의 중년 동아시아계 남성으로 보이며 낡은 회갈색 재킷, 올리브색 전술 조끼와 장갑을 착용했다. 시체는 짧은 검은 머리의 성인 동아시아계 남성으로 보이고 젖은 어두운 옷과 신발을 착용했다. 구체적인 인물 외형이나 복장을 고정하는 별도 참고는 없다. 긴 나무 막대기, 고무보트, 홍수 물과 떠 있는 잔해가 모두 식별되며, 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "민병대원은 골반과 하체를 보트 안쪽에 두고 튜브 가장자리에 기대며 상체를 내민다. 발은 가려져 있지만 보트에 지지되는 자세로 읽힌다. 두 장갑 낀 손이 막대기를 떨어져 잡아 어깨 접촉점으로 힘을 전달한다. 시체의 머리와 팔은 수면 및 인접 잔해 높이에 낮게 놓이고 몸통과 다리는 부력으로 지지되며, 스스로 들어 올린 팔다리는 없다. 보트와 시체 사이에는 물이 이어지고 보트 주변의 물결도 이동 상황과 양립한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.25
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)"
    ],
    "B": [
     "[gemini-pro] 레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)",
     "[gemini-pro] 시체의 다리와 발뒤꿈치가 수면 위로 완전히 노출된 물리적으로 불가능한 부유 자세",
     "[gemini-pro] 막대기가 시체의 몸통 표면을 누르는 대신 몸을 뚫고 들어간 물리적 오류"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1000
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스의 얼굴을 배제하라는 지시를 위반했으나, 막대기로 시체를 누르는 물리적 동작과 부력의 묘사가 비교적 자연스러움.  ★위반: [gemini-pro] 레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반)"
   },
   {
    "label": "B",
    "score": 1000,
    "verdict_ko": "레퍼런스 인물 복사 위반과 더불어, 시체가 평지에 누운 듯 수면에 떠 있고 막대기가 몸을 관통하는 물리적 오류가 있음.  ★위반: [gemini-pro] 레퍼런스 이미지의 인물 얼굴을 그대로 복사하여 사용함 (금지 조건 위반) / [gemini-pro] 시체의 다리와 발뒤꿈치가 수면 위로 완전히 노출된 물리적으로 불가능한 부유 자세 / [gemini-pro] 막대기가 시체의 몸통 표면을 누르는 대신 몸을 뚫고 들어간 물리적 오류"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2_sel.png",
    "asset_id": "6999c042-aed8-416d-94a6-2d5ba0a0b1ec",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-70b7-722e-8fe8-6fb2fde5b1e7",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S39sh2"
  }
 },
 "S39sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:46:04.142226+00:00",
  "fingerprint": "a77a88f4ced814315f4a75349f66fe389defb078633d62ab18885b90c212aded",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S39sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S39sh6_sel.png",
  "source_sha256": "0b7d47ef5409a6b6a0d3635b478a7ed63ef57c680ad8c35661c753ba4324042c",
  "file": "S39sh6_cine.png",
  "staged_sha256": "f5567a7b2eacc653af28cc835cd45276c3809483c59e97120c4e3e1ea95dd6a3",
  "latency_ms": 9509
 },
 "S40sh1::signage": {
  "fp": "2ce0de0d9ed5006c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S40sh1::bgfirst_bg": {
  "input_fingerprint": "991a18db2d387744",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1__bgfirst_bg.png",
  "asset_id": "3f8ff6ab-9364-4e5b-a480-4d9ff441d727",
  "input_asset_ids": [
   "e04d942a-da09-45cf-bfc4-228283629bc9",
   "2c19fbb8-19d6-4259-a297-df4959b88a58"
  ]
 },
 "S40sh1": {
  "input_fingerprint": "79fdfc750c15b172",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is daylight on the outskirts of the flooded refugee settlement. 앰버: She has stopped from exhaustion, with tears covering her face and her attention repeatedly turning backward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is daylight on the outskirts of the flooded refugee settlement. 앰버: She has stopped from exhaustion, with tears covering her face and her attention repeatedly turning backward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is daylight on the outskirts of the flooded refugee settlement. 앰버: She has stopped from exhaustion, with tears covering her face and her attention repeatedly turning backward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1__bgfirst_bg.png",
     "asset_id": "3f8ff6ab-9364-4e5b-a480-4d9ff441d727",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S40sh1.png",
     "asset_id": "e04d942a-da09-45cf-bfc4-228283629bc9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L194B02.png",
     "asset_id": "2c19fbb8-19d6-4259-a297-df4959b88a58",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선이 화면 좌측으로 향하며 뒤를 돌아보는 동작을 명확히 보여줌.",
    "built_space": "도로가 우측으로 비스듬히 멀어지며 지시된 오른쪽 쐐기형 빈 공간을 정확히 형성함. 건물이 좌측에 위치함.",
    "entities": "앰버의 외모, 땀과 눈물, 남색 티셔츠와 멜빵바지, 머리 위의 방독면이 기준과 완벽히 일치함.",
    "hard_violations": [],
    "physics": "화면 밖 지면을 기반으로 지친 상체가 자연스럽게 멈춰 서 있음."
   },
   {
    "label": "B",
    "direction": "시선이 카메라 정면을 똑바로 응시하여 뒤를 돌아본다는 지시와 어긋남.",
    "built_space": "도로가 화면 중앙에서 일직선으로 멀어지며, 지정된 비대칭 프레이밍을 형성하지 못함.",
    "entities": "앰버의 이목구비는 일치하나 티셔츠가 회색이며 머리 위 방독면 소품이 누락됨.",
    "hard_violations": [],
    "physics": "자연스러운 움직임 없이 경직된 자세로 지면에 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지시된 비대칭적 도로 구도, 뒤를 돌아보는 시선, 캐릭터의 의상과 소품(방독면)을 모두 정확히 구현한 결과물입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "정면 응시와 대칭적 도로 구도로 연출 지시를 위반했으며, 의상 색상 오류 및 방독면 누락이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선이 화면 좌측으로 향하며 뒤를 돌아보는 동작을 명확히 보여줌.",
        "built_space": "도로가 우측으로 비스듬히 멀어지며 지시된 오른쪽 쐐기형 빈 공간을 정확히 형성함. 건물이 좌측에 위치함.",
        "entities": "앰버의 외모, 땀과 눈물, 남색 티셔츠와 멜빵바지, 머리 위의 방독면이 기준과 완벽히 일치함.",
        "hard_violations": [],
        "physics": "화면 밖 지면을 기반으로 지친 상체가 자연스럽게 멈춰 서 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 정면을 똑바로 응시하여 뒤를 돌아본다는 지시와 어긋남.",
        "built_space": "도로가 화면 중앙에서 일직선으로 멀어지며, 지정된 비대칭 프레이밍을 형성하지 못함.",
        "entities": "앰버의 이목구비는 일치하나 티셔츠가 회색이며 머리 위 방독면 소품이 누락됨.",
        "hard_violations": [],
        "physics": "자연스러운 움직임 없이 경직된 자세로 지면에 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지시된 비대칭적 도로 구도, 뒤를 돌아보는 시선, 캐릭터의 의상과 소품(방독면)을 모두 정확히 구현한 결과물입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "정면 응시와 대칭적 도로 구도로 연출 지시를 위반했으며, 의상 색상 오류 및 방독면 누락이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선이 화면 좌측으로 향하며 뒤를 돌아보는 동작을 명확히 보여줌.",
        "built_space": "도로가 우측으로 비스듬히 멀어지며 지시된 오른쪽 쐐기형 빈 공간을 정확히 형성함. 건물이 좌측에 위치함.",
        "entities": "앰버의 외모, 땀과 눈물, 남색 티셔츠와 멜빵바지, 머리 위의 방독면이 기준과 완벽히 일치함.",
        "hard_violations": [],
        "physics": "화면 밖 지면을 기반으로 지친 상체가 자연스럽게 멈춰 서 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 정면을 똑바로 응시하여 뒤를 돌아본다는 지시와 어긋남.",
        "built_space": "도로가 화면 중앙에서 일직선으로 멀어지며, 지정된 비대칭 프레이밍을 형성하지 못함.",
        "entities": "앰버의 이목구비는 일치하나 티셔츠가 회색이며 머리 위 방독면 소품이 누락됨.",
        "hard_violations": [],
        "physics": "자연스러운 움직임 없이 경직된 자세로 지면에 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "상체 미디엄 숏과 빈 낮 도로는 맞지만, 중앙의 정면 응시가 뒤를 돌아보는 상태와 오른쪽 여백 구도를 놓치며 얼굴·상의·머리 장비도 참조와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "왼쪽에 지친 상체를 두고 오른쪽으로 빈 도로를 열어 둔 미디엄 숏이며, 뒤쪽을 살피는 시선과 땀·눈물, 참조 의상 및 장소 특징을 가장 충실하게 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 몸통이 거의 정면이고 눈은 카메라 쪽을 향한다. 뒤쪽 탈출 경로를 돌아보는 시선은 보이지 않는다. 도로는 인물 뒤에서 화면 왼쪽 먼 곳으로 비스듬히 이어진다.",
        "built_space": "인물은 아스팔트 차도 중앙 부근에 있고 허리 부근에서 잘린다. 오른쪽에 연석과 인도, 낡은 콘크리트 건물 일부 및 판자로 막힌 창 한 구획이 보이며, 왼쪽 원경에 별도의 낮은 건물 한 동이 있다. 그 사이에 천막과 철망, 여러 전신주가 이어진다. 참조의 재료와 정착지 분위기는 유사하지만 상점 출입구와 창 배열은 이 화면에서 확인되지 않는다. 빈 도로는 양옆에 있으나 인물이 중앙에 있어 지정된 오른쪽 쐐기형 여백이 약하다.",
        "entities": "보이는 사람은 금발 여자아이 한 명뿐이다. 어린 연령대는 맞지만 참조보다 얼굴이 길고 눈이 작아 보이며, 지정된 혼혈 정체성은 외관만으로 확정할 수 없다. 갈색 작업 멜빵바지와 가슴 주머니 도구, 허리의 렌치는 보인다. 티셔츠는 참조의 남색 대신 회색이며 머리 위 호흡 보호 장비는 없다. 얼굴에 오염과 젖은 흔적은 있으나 눈물과 땀으로 엉망이 된 상태는 비교적 약하다. 판독 가능한 문자는 없다.",
        "hard_violations": [],
        "physics": "몸통은 수직으로 서 있고 양팔이 아래로 늘어져 있다. 발은 프레임 밖이므로 접지는 직접 확인할 수 없지만, 몸이 떠 있거나 지지 불가능한 자세는 아니다. 도구는 주머니와 허리 장구에 걸려 있다. 다만 양쪽 어깨와 팔이 거의 대칭이라 지쳐 움직임을 멈춘 순간보다 정면으로 자세를 잡은 모습에 가깝다."
       },
       {
        "label": "B",
        "direction": "몸통은 앞으로 조금 숙인 채 얼굴과 두 눈을 화면 왼쪽 바깥으로 돌리고 있다. 시선의 구체적 대상은 화면에 없지만, 카메라를 보지 않고 뒤쪽을 살피는 행동으로 읽힌다. 빈 도로는 인물 오른쪽에서 오른쪽 먼 소실점으로 비스듬히 이어진다.",
        "built_space": "인물의 상체가 화면 왼쪽을 차지하고 오른쪽에는 넓고 비어 있는 아스팔트 도로가 남는다. 왼쪽 배경에는 낡은 단층 콘크리트 상점, 유리 출입구 한 구획, 그 왼쪽의 유리창 두 구획 중 일부, 출입구 옆 단말기 한 개가 보인다. 연석과 인도 뒤로 철망, 천막, 전신주가 이어져 참조 장소의 주요 재료와 시설을 유지한다. 인물은 인도가 아닌 차도에 있으며 시설과 인물의 상대 크기도 자연스럽다.",
        "entities": "금발 여자아이 한 명만 있으며 둥근 얼굴과 큰 눈, 체격이 참조의 열 살 앰버에 더 가깝다. 혼혈 정체성 자체는 외관만으로 단정할 수 없다. 남색 티셔츠, 갈색 작업 멜빵바지, 가슴 주머니 도구와 머리 위 호흡 보호 장비가 참조와 대응한다. 보호 장비의 세부 형태는 참조와 조금 다르다. 이마와 코, 볼에 땀의 광택이 있고 볼을 따라 눈물 자국이 보인다. 주변에 다른 사람이나 차량은 없고 읽을 수 있는 문자도 없다.",
        "hard_violations": [],
        "physics": "상체를 앞으로 기울이고 어깨를 낮춘 자세가 탈진으로 멈춘 순간에 자연스럽다. 팔은 아래로 이어지며 손과 발은 프레임 밖이라 무릎 등에 손을 짚었는지는 확인할 수 없다. 보이는 몸통과 골반의 연결에는 부유나 불가능한 지지 문제가 없다. 머리 장비는 머리에 얹혀 끈으로 고정되고 도구는 가슴 주머니에 꽂혀 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "상체 미디엄 숏과 빈 낮 도로는 맞지만, 중앙의 정면 응시가 뒤를 돌아보는 상태와 오른쪽 여백 구도를 놓치며 얼굴·상의·머리 장비도 참조와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "왼쪽에 지친 상체를 두고 오른쪽으로 빈 도로를 열어 둔 미디엄 숏이며, 뒤쪽을 살피는 시선과 땀·눈물, 참조 의상 및 장소 특징을 가장 충실하게 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 몸통이 거의 정면이고 눈은 카메라 쪽을 향한다. 뒤쪽 탈출 경로를 돌아보는 시선은 보이지 않는다. 도로는 인물 뒤에서 화면 왼쪽 먼 곳으로 비스듬히 이어진다.",
        "built_space": "인물은 아스팔트 차도 중앙 부근에 있고 허리 부근에서 잘린다. 오른쪽에 연석과 인도, 낡은 콘크리트 건물 일부 및 판자로 막힌 창 한 구획이 보이며, 왼쪽 원경에 별도의 낮은 건물 한 동이 있다. 그 사이에 천막과 철망, 여러 전신주가 이어진다. 참조의 재료와 정착지 분위기는 유사하지만 상점 출입구와 창 배열은 이 화면에서 확인되지 않는다. 빈 도로는 양옆에 있으나 인물이 중앙에 있어 지정된 오른쪽 쐐기형 여백이 약하다.",
        "entities": "보이는 사람은 금발 여자아이 한 명뿐이다. 어린 연령대는 맞지만 참조보다 얼굴이 길고 눈이 작아 보이며, 지정된 혼혈 정체성은 외관만으로 확정할 수 없다. 갈색 작업 멜빵바지와 가슴 주머니 도구, 허리의 렌치는 보인다. 티셔츠는 참조의 남색 대신 회색이며 머리 위 호흡 보호 장비는 없다. 얼굴에 오염과 젖은 흔적은 있으나 눈물과 땀으로 엉망이 된 상태는 비교적 약하다. 판독 가능한 문자는 없다.",
        "hard_violations": [],
        "physics": "몸통은 수직으로 서 있고 양팔이 아래로 늘어져 있다. 발은 프레임 밖이므로 접지는 직접 확인할 수 없지만, 몸이 떠 있거나 지지 불가능한 자세는 아니다. 도구는 주머니와 허리 장구에 걸려 있다. 다만 양쪽 어깨와 팔이 거의 대칭이라 지쳐 움직임을 멈춘 순간보다 정면으로 자세를 잡은 모습에 가깝다."
       },
       {
        "label": "A",
        "direction": "몸통은 앞으로 조금 숙인 채 얼굴과 두 눈을 화면 왼쪽 바깥으로 돌리고 있다. 시선의 구체적 대상은 화면에 없지만, 카메라를 보지 않고 뒤쪽을 살피는 행동으로 읽힌다. 빈 도로는 인물 오른쪽에서 오른쪽 먼 소실점으로 비스듬히 이어진다.",
        "built_space": "인물의 상체가 화면 왼쪽을 차지하고 오른쪽에는 넓고 비어 있는 아스팔트 도로가 남는다. 왼쪽 배경에는 낡은 단층 콘크리트 상점, 유리 출입구 한 구획, 그 왼쪽의 유리창 두 구획 중 일부, 출입구 옆 단말기 한 개가 보인다. 연석과 인도 뒤로 철망, 천막, 전신주가 이어져 참조 장소의 주요 재료와 시설을 유지한다. 인물은 인도가 아닌 차도에 있으며 시설과 인물의 상대 크기도 자연스럽다.",
        "entities": "금발 여자아이 한 명만 있으며 둥근 얼굴과 큰 눈, 체격이 참조의 열 살 앰버에 더 가깝다. 혼혈 정체성 자체는 외관만으로 단정할 수 없다. 남색 티셔츠, 갈색 작업 멜빵바지, 가슴 주머니 도구와 머리 위 호흡 보호 장비가 참조와 대응한다. 보호 장비의 세부 형태는 참조와 조금 다르다. 이마와 코, 볼에 땀의 광택이 있고 볼을 따라 눈물 자국이 보인다. 주변에 다른 사람이나 차량은 없고 읽을 수 있는 문자도 없다.",
        "hard_violations": [],
        "physics": "상체를 앞으로 기울이고 어깨를 낮춘 자세가 탈진으로 멈춘 순간에 자연스럽다. 팔은 아래로 이어지며 손과 발은 프레임 밖이라 무릎 등에 손을 짚었는지는 확인할 수 없다. 보이는 몸통과 골반의 연결에는 부유나 불가능한 지지 문제가 없다. 머리 장비는 머리에 얹혀 끈으로 고정되고 도구는 가슴 주머니에 꽂혀 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.056
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.056
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1056
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 비대칭적 도로 구도, 뒤를 돌아보는 시선, 캐릭터의 의상과 소품(방독면)을 모두 정확히 구현한 결과물입니다."
   },
   {
    "label": "B",
    "score": 1056,
    "verdict_ko": "정면 응시와 대칭적 도로 구도로 연출 지시를 위반했으며, 의상 색상 오류 및 방독면 누락이 발생했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L194B02.png",
    "asset_id": "2c19fbb8-19d6-4259-a297-df4959b88a58",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-7260-79aa-a12a-49d55019c9b1",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1__bgfirst_bg.png",
   "bg_asset_id": "3f8ff6ab-9364-4e5b-a480-4d9ff441d727",
   "bg_record_key": "S40sh1::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S40sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:47:08.145104+00:00",
  "fingerprint": "1b782b0cc70bb48200cd8bf91ed260590f533f9526d9dc8c6a77fe335c8352cb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S40sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S40sh1_sel.png",
  "source_sha256": "5440ea37bcf147bd2e82c75593223e3b35ba05edb35435f5012006fa07fc70c7",
  "file": "S40sh1_cine.png",
  "staged_sha256": "bc30984546f8121d84035dd609db0dc2cc6d3cba693b7f69ac39c1850f4052e1",
  "latency_ms": 9647
 },
 "S40sh4::signage": {
  "fp": "6edb47a98acf4916",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S40sh4": {
  "input_fingerprint": "aa4222f63ec93cef",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버를 향해 미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 현우의 절박한 얼굴.\n\nLOCATION (lock): On the exposed roadside at the refugee settlement's outskirts, where the exhausted child has stopped. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Street (No police are present in the street) — Only a restricted portion remains visible behind 현우; used as Soft peripheral context keeps the shot connected to the escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established daytime ambient light and restrained contrast so urgency comes from expression rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the empty gray asphalt and the same daylight conditions at this roadside stopping point. Exclude floodwater, floating panels, and search boats from the earlier flooded district.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape continues through the daylight outskirts toward the city. 앰버: She remains exhausted and tear-streaked, reluctant to move on. 현우: He retains the facial bruises and untreated dog-bite wound on his leg. The intact contact card remains hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버를 향해 미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 현우의 절박한 얼굴.\n\nLOCATION (lock): On the exposed roadside at the refugee settlement's outskirts, where the exhausted child has stopped. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Street (No police are present in the street) — Only a restricted portion remains visible behind 현우; used as Soft peripheral context keeps the shot connected to the escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established daytime ambient light and restrained contrast so urgency comes from expression rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the empty gray asphalt and the same daylight conditions at this roadside stopping point. Exclude floodwater, floating panels, and search boats from the earlier flooded district.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape continues through the daylight outskirts toward the city. 앰버: She remains exhausted and tear-streaked, reluctant to move on. 현우: He retains the facial bruises and untreated dog-bite wound on his leg. The intact contact card remains hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버를 향해 미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 현우의 절박한 얼굴.\n\nLOCATION (lock): On the exposed roadside at the refugee settlement's outskirts, where the exhausted child has stopped. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Street (No police are present in the street) — Only a restricted portion remains visible behind 현우; used as Soft peripheral context keeps the shot connected to the escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established daytime ambient light and restrained contrast so urgency comes from expression rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the empty gray asphalt and the same daylight conditions at this roadside stopping point. Exclude floodwater, floating panels, and search boats from the earlier flooded district.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape continues through the daylight outskirts toward the city. 앰버: She remains exhausted and tear-streaked, reluctant to move on. 현우: He retains the facial bruises and untreated dog-bite wound on his leg. The intact contact card remains hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 소리치는 방향이 프레임 좌측에 있는 금발 머리의 앰버를 향하고 있습니다.",
    "built_space": "이전 샷에서 볼 수 있는 텅 빈 아스팔트 도로와 전신주가 배경으로 제한적으로 보입니다.",
    "entities": "현우는 레퍼런스와 일치하는 외모에 헝클어진 머리, 요구된 얼굴의 멍을 뚜렷하게 보여줍니다. 좌측에는 앰버의 금발 머리가 보입니다.",
    "hard_violations": [],
    "physics": "몸을 살짝 기울여 소리치는 자연스러운 자세를 취하고 있으며, 보이지 않는 지면에 의해 안정적으로 지탱되고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우가 프레임 좌측의 앰버를 향해 시선을 고정하고 입을 벌려 소리치고 있습니다.",
    "built_space": "아스팔트 도로, 차선, 전신주 및 먼 산이 배경에 올바르게 배치되어 있습니다.",
    "entities": "현우의 외모는 레퍼런스와 일치하나 상처의 흔적은 A보다 덜 뚜렷합니다. 앰버는 이전 샷에서 지정된 파란색 셔츠와 멜빵을 입은 채 뒤통수와 옆모습이 보입니다.",
    "hard_violations": [],
    "physics": "앞으로 몸을 숙이며 소리치는 역동적인 자세가 지면에 의해 잘 지탱되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "지시된 클로즈업 프레이밍을 정확히 구현했으며, 미간을 찌푸리고 절박하게 소리치는 표정과 얼굴의 상처(멍) 등 프롬프트의 디테일을 매우 사실적으로 표현했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트의 요구사항을 훌륭하게 따랐으며 이전 샷의 앰버 의상 디테일까지 잘 살렸으나, A에 비해 현우 얼굴의 상처 표현과 표정의 절박함이 미세하게 덜합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 소리치는 방향이 프레임 좌측에 있는 금발 머리의 앰버를 향하고 있습니다.",
        "built_space": "이전 샷에서 볼 수 있는 텅 빈 아스팔트 도로와 전신주가 배경으로 제한적으로 보입니다.",
        "entities": "현우는 레퍼런스와 일치하는 외모에 헝클어진 머리, 요구된 얼굴의 멍을 뚜렷하게 보여줍니다. 좌측에는 앰버의 금발 머리가 보입니다.",
        "hard_violations": [],
        "physics": "몸을 살짝 기울여 소리치는 자연스러운 자세를 취하고 있으며, 보이지 않는 지면에 의해 안정적으로 지탱되고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우가 프레임 좌측의 앰버를 향해 시선을 고정하고 입을 벌려 소리치고 있습니다.",
        "built_space": "아스팔트 도로, 차선, 전신주 및 먼 산이 배경에 올바르게 배치되어 있습니다.",
        "entities": "현우의 외모는 레퍼런스와 일치하나 상처의 흔적은 A보다 덜 뚜렷합니다. 앰버는 이전 샷에서 지정된 파란색 셔츠와 멜빵을 입은 채 뒤통수와 옆모습이 보입니다.",
        "hard_violations": [],
        "physics": "앞으로 몸을 숙이며 소리치는 역동적인 자세가 지면에 의해 잘 지탱되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "지시된 클로즈업 프레이밍을 정확히 구현했으며, 미간을 찌푸리고 절박하게 소리치는 표정과 얼굴의 상처(멍) 등 프롬프트의 디테일을 매우 사실적으로 표현했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트의 요구사항을 훌륭하게 따랐으며 이전 샷의 앰버 의상 디테일까지 잘 살렸으나, A에 비해 현우 얼굴의 상처 표현과 표정의 절박함이 미세하게 덜합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 소리치는 방향이 프레임 좌측에 있는 금발 머리의 앰버를 향하고 있습니다.",
        "built_space": "이전 샷에서 볼 수 있는 텅 빈 아스팔트 도로와 전신주가 배경으로 제한적으로 보입니다.",
        "entities": "현우는 레퍼런스와 일치하는 외모에 헝클어진 머리, 요구된 얼굴의 멍을 뚜렷하게 보여줍니다. 좌측에는 앰버의 금발 머리가 보입니다.",
        "hard_violations": [],
        "physics": "몸을 살짝 기울여 소리치는 자연스러운 자세를 취하고 있으며, 보이지 않는 지면에 의해 안정적으로 지탱되고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우가 프레임 좌측의 앰버를 향해 시선을 고정하고 입을 벌려 소리치고 있습니다.",
        "built_space": "아스팔트 도로, 차선, 전신주 및 먼 산이 배경에 올바르게 배치되어 있습니다.",
        "entities": "현우의 외모는 레퍼런스와 일치하나 상처의 흔적은 A보다 덜 뚜렷합니다. 앰버는 이전 샷에서 지정된 파란색 셔츠와 멜빵을 입은 채 뒤통수와 옆모습이 보입니다.",
        "hard_violations": [],
        "physics": "앞으로 몸을 숙이며 소리치는 역동적인 자세가 지면에 의해 잘 지탱되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "앰버를 향한 시선과 찌푸린 미간, 크게 벌린 입은 정확하지만, 앰버와 도로의 비중이 커 현우의 절박한 얼굴에 집중하는 클로즈업이 B보다 느슨하고 얼굴의 멍도 덜 분명하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 얼굴을 더 밀착해 담아 절박한 외침과 앰버를 향한 시선을 강조하며, 얼굴의 멍과 회색 셔츠, 낮의 빈 도로를 충실하게 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽 전경의 앰버 얼굴을 향해 고개와 눈을 돌리고 입을 크게 벌린다. 앰버도 현우 쪽으로 얼굴을 향하고 있어 외침의 대상이 명확하다. 무기나 방향성 있는 소품, 이동 중인 물체는 없다.",
        "built_space": "두 사람 뒤로 회색 아스팔트와 흰 차선, 맞은편 연석, 마른 풀밭과 여러 전신주가 보인다. 건물이나 실내 설비는 프레임에 없으며 경찰과 다른 행인도 없다. 이전 장면의 외곽 도로 재질과 낮빛은 이어지지만, 도로가 비교적 넓고 선명하게 남아 제한적인 주변 맥락이라는 지시에는 다소 느슨하다.",
        "entities": "현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로, 참고 인물과 대체로 유사한 얼굴과 회색 셔츠를 보인다. 미간을 깊게 찌푸리고 입을 크게 벌렸으며 얼굴에 오염과 옅은 상처 흔적이 있지만 멍은 뚜렷하지 않다. 왼쪽 앰버의 일부 옆얼굴, 금발, 남색 상의와 갈색 멜빵은 이전 장면과 부합한다. 앰버의 눈물과 전체 얼굴은 이 각도에서 확인하기 어렵다. 다리 상처와 신발 속 카드는 프레임 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 머리는 목과 어깨에 자연스럽게 연결되고 상체가 앰버 쪽으로 기울어 있다. 앰버의 머리와 어깨도 정상적으로 연결된다. 하체와 발의 지지점은 클로즈업 밖이므로 확인할 수 없지만, 공중에 뜬 신체나 지지 없이 떠 있는 물체는 없다. 입과 얼굴 근육의 움직임은 실제 외침으로 가능한 형태다."
       },
       {
        "label": "B",
        "direction": "현우의 눈과 얼굴은 화면 왼쪽에 일부 보이는 앰버를 향한다. 렌즈를 응시하는 것이 아니라 가까운 상대에게 소리치는 관계가 읽힌다. 앰버는 뒷머리 위주로 보여 정확한 눈 방향은 확인할 수 없다. 무기나 이동 물체는 없다.",
        "built_space": "현우의 얼굴 뒤로 빈 회색 도로와 차선, 오른쪽으로 물러나는 연석과 마른 풀, 흐릿한 전신주들이 보인다. 실내 설비나 반사는 없고 경찰과 추가 인물도 없다. 배경은 얕은 초점으로 주변에만 남아 도로의 맥락을 제공하며, 참고 장면의 낮빛과 외곽 도로 환경에 부합한다.",
        "entities": "현우는 참고 이미지와 유사한 젊은 동아시아계 남성의 얼굴, 헝클어진 검은 머리와 회색 셔츠를 갖고 있다. 눈썹 주변과 볼의 멍, 긁힌 흔적이 보이며 깊게 모인 미간과 크게 열린 입이 절박한 외침을 표현한다. 왼쪽 가장자리에는 앰버로 읽히는 금발 머리와 얼굴 가장자리만 있어 나이, 혼혈 외모, 의상과 눈물은 직접 확인하기 어렵다. 다리와 신발은 프레임 밖이므로 상처와 숨긴 카드도 보이지 않는다. 추가 인물이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "현우의 머리와 목, 어깨가 자연스럽게 이어지고 턱을 내린 외침의 자세도 해부학적으로 가능하다. 앰버는 머리 일부만 보이며 부유를 시사하는 배치는 없다. 발과 지면 접촉은 프레임 밖이다. 손에 든 물건이나 별도 지지가 필요한 공중 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "앰버를 향한 시선과 찌푸린 미간, 크게 벌린 입은 정확하지만, 앰버와 도로의 비중이 커 현우의 절박한 얼굴에 집중하는 클로즈업이 B보다 느슨하고 얼굴의 멍도 덜 분명하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우의 얼굴을 더 밀착해 담아 절박한 외침과 앰버를 향한 시선을 강조하며, 얼굴의 멍과 회색 셔츠, 낮의 빈 도로를 충실하게 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽 전경의 앰버 얼굴을 향해 고개와 눈을 돌리고 입을 크게 벌린다. 앰버도 현우 쪽으로 얼굴을 향하고 있어 외침의 대상이 명확하다. 무기나 방향성 있는 소품, 이동 중인 물체는 없다.",
        "built_space": "두 사람 뒤로 회색 아스팔트와 흰 차선, 맞은편 연석, 마른 풀밭과 여러 전신주가 보인다. 건물이나 실내 설비는 프레임에 없으며 경찰과 다른 행인도 없다. 이전 장면의 외곽 도로 재질과 낮빛은 이어지지만, 도로가 비교적 넓고 선명하게 남아 제한적인 주변 맥락이라는 지시에는 다소 느슨하다.",
        "entities": "현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로, 참고 인물과 대체로 유사한 얼굴과 회색 셔츠를 보인다. 미간을 깊게 찌푸리고 입을 크게 벌렸으며 얼굴에 오염과 옅은 상처 흔적이 있지만 멍은 뚜렷하지 않다. 왼쪽 앰버의 일부 옆얼굴, 금발, 남색 상의와 갈색 멜빵은 이전 장면과 부합한다. 앰버의 눈물과 전체 얼굴은 이 각도에서 확인하기 어렵다. 다리 상처와 신발 속 카드는 프레임 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 머리는 목과 어깨에 자연스럽게 연결되고 상체가 앰버 쪽으로 기울어 있다. 앰버의 머리와 어깨도 정상적으로 연결된다. 하체와 발의 지지점은 클로즈업 밖이므로 확인할 수 없지만, 공중에 뜬 신체나 지지 없이 떠 있는 물체는 없다. 입과 얼굴 근육의 움직임은 실제 외침으로 가능한 형태다."
       },
       {
        "label": "A",
        "direction": "현우의 눈과 얼굴은 화면 왼쪽에 일부 보이는 앰버를 향한다. 렌즈를 응시하는 것이 아니라 가까운 상대에게 소리치는 관계가 읽힌다. 앰버는 뒷머리 위주로 보여 정확한 눈 방향은 확인할 수 없다. 무기나 이동 물체는 없다.",
        "built_space": "현우의 얼굴 뒤로 빈 회색 도로와 차선, 오른쪽으로 물러나는 연석과 마른 풀, 흐릿한 전신주들이 보인다. 실내 설비나 반사는 없고 경찰과 추가 인물도 없다. 배경은 얕은 초점으로 주변에만 남아 도로의 맥락을 제공하며, 참고 장면의 낮빛과 외곽 도로 환경에 부합한다.",
        "entities": "현우는 참고 이미지와 유사한 젊은 동아시아계 남성의 얼굴, 헝클어진 검은 머리와 회색 셔츠를 갖고 있다. 눈썹 주변과 볼의 멍, 긁힌 흔적이 보이며 깊게 모인 미간과 크게 열린 입이 절박한 외침을 표현한다. 왼쪽 가장자리에는 앰버로 읽히는 금발 머리와 얼굴 가장자리만 있어 나이, 혼혈 외모, 의상과 눈물은 직접 확인하기 어렵다. 다리와 신발은 프레임 밖이므로 상처와 숨긴 카드도 보이지 않는다. 추가 인물이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "현우의 머리와 목, 어깨가 자연스럽게 이어지고 턱을 내린 외침의 자세도 해부학적으로 가능하다. 앰버는 머리 일부만 보이며 부유를 시사하는 배치는 없다. 발과 지면 접촉은 프레임 밖이다. 손에 든 물건이나 별도 지지가 필요한 공중 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.789
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.789
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1789
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 프레이밍을 정확히 구현했으며, 미간을 찌푸리고 절박하게 소리치는 표정과 얼굴의 상처(멍) 등 프롬프트의 디테일을 매우 사실적으로 표현했습니다."
   },
   {
    "label": "B",
    "score": 1789,
    "verdict_ko": "프롬프트의 요구사항을 훌륭하게 따랐으며 이전 샷의 앰버 의상 디테일까지 잘 살렸으나, A에 비해 현우 얼굴의 상처 표현과 표정의 절박함이 미세하게 덜합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1_sel.png",
    "asset_id": "1b766e88-1c90-4312-815c-0bace2663863",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-75aa-7d1d-a07b-0de158354cc3",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S40sh1"
  },
  "staged_characters_added": [
   "C03"
  ]
 },
 "S40sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:48:03.843377+00:00",
  "fingerprint": "49854d7df2e20348502da736ed3008888eeefca8c7afb199749dd6bc8b8291d5",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S40sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S40sh4_sel.png",
  "source_sha256": "4bcd9ce7439e5beadc0fffeda89f5a77af060d0cb3f7c1002b753a1ca1fba324",
  "file": "S40sh4_cine.png",
  "staged_sha256": "ef8b25befb8f4caf86b535d1fc5dd4fc34420a0f23cd9cab5e3406f68d0b5a5c",
  "latency_ms": 10193
 },
 "S40sh6::signage": {
  "fp": "6a851e20df9e5221",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "네온 불빛이 켜진 무인 점포"
   }
  ],
  "dropped": []
 },
 "S40sh6": {
  "input_fingerprint": "a0ae3f5f339f551b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 네온 불빛이 켜진 무인 점포를 배경으로, 한쪽 발을 바닥에서 떼고 앞으로 달려나가는 자세의 현우와 앰버, 그리고 그 뒤를 묵묵히 따라 육중한 다리를 길게 내딛은 찰리의 뒷모습 전경.\n\nLOCATION (lock): On the street approaching an unattended shop, with its illuminated exterior signage ahead of the fleeing group. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unattended shop across the street in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Unattended shop across the street (Visible across the street beyond the fleeing group) — Its street-facing exterior is seen obliquely in the upper-right background; used as Anchors the wider geography while remaining subordinate to the running figures; Street (No police are present) — The route extends away from the camera, with the shop across it; used as Provides separation between the camera, the three runners, and the opposite frontage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains dominant, with the shop's stated neon contributing only a localized accent rather than recoloring the street.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An unmanned store stands across the street along the daylight escape route. 앰버: She is moving again but remains exhausted and tear-streaked. 현우: He continues escaping with facial bruises and the untreated leg wound. The contact card remains concealed inside his shoe. 찰리: He follows along the street, still in his old coat-and-hat disguise over the worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 네온 불빛이 켜진 무인 점포를 배경으로, 한쪽 발을 바닥에서 떼고 앞으로 달려나가는 자세의 현우와 앰버, 그리고 그 뒤를 묵묵히 따라 육중한 다리를 길게 내딛은 찰리의 뒷모습 전경.\n\nLOCATION (lock): On the street approaching an unattended shop, with its illuminated exterior signage ahead of the fleeing group. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unattended shop across the street in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Unattended shop across the street (Visible across the street beyond the fleeing group) — Its street-facing exterior is seen obliquely in the upper-right background; used as Anchors the wider geography while remaining subordinate to the running figures; Street (No police are present) — The route extends away from the camera, with the shop across it; used as Provides separation between the camera, the three runners, and the opposite frontage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains dominant, with the shop's stated neon contributing only a localized accent rather than recoloring the street.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An unmanned store stands across the street along the daylight escape route. 앰버: She is moving again but remains exhausted and tear-streaked. 현우: He continues escaping with facial bruises and the untreated leg wound. The contact card remains concealed inside his shoe. 찰리: He follows along the street, still in his old coat-and-hat disguise over the worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 네온 불빛이 켜진 무인 점포를 배경으로, 한쪽 발을 바닥에서 떼고 앞으로 달려나가는 자세의 현우와 앰버, 그리고 그 뒤를 묵묵히 따라 육중한 다리를 길게 내딛은 찰리의 뒷모습 전경.\n\nLOCATION (lock): On the street approaching an unattended shop, with its illuminated exterior signage ahead of the fleeing group. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unattended shop across the street in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Unattended shop across the street (Visible across the street beyond the fleeing group) — Its street-facing exterior is seen obliquely in the upper-right background; used as Anchors the wider geography while remaining subordinate to the running figures; Street (No police are present) — The route extends away from the camera, with the shop across it; used as Provides separation between the camera, the three runners, and the opposite frontage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains dominant, with the shop's stated neon contributing only a localized accent rather than recoloring the street.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An unmanned store stands across the street along the daylight escape route. 앰버: She is moving again but remains exhausted and tear-streaked. 현우: He continues escaping with facial bruises and the untreated leg wound. The contact card remains concealed inside his shoe. 찰리: He follows along the street, still in his old coat-and-hat disguise over the worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 앰버는 카메라를 향해 앞으로 달려오고 있으나, 찰리는 카메라를 등지고 반대 방향(배경 쪽)으로 걸어가고 있어 서로 엇갈리며 '뒤를 따르는' 방향성이 성립하지 않음.",
    "built_space": "도로가 깊이감을 주며 뻗어 있고, 지시된 무인 점포가 우측 상단 배경에 위치하여 프레임 구도를 충족함.",
    "entities": "현우(타박상, 의상), 앰버(머리 위 고글 포함), 찰리(코트, 모자, 뒷모습) 모두 레퍼런스의 외형에 부합함.",
    "hard_violations": [
     "[gemini-pro] 간판 및 창문에 읽을 수 있는 문자(안작잡, OPEN)가 뚜렷하게 노출됨 (금지된 텍스트 포함)",
     "[gpt-high] 찰리를 도주하는 두 사람과 반대 방향으로 걷게 하여 같은 경로에서 뒤따른다는 필수 배치를 위반했다.",
     "[gpt-high] 큰 간판의 한글과 두 네온 표지의 영문이 읽혀 글자 금지 조건을 위반했다."
    ],
    "physics": "현우와 앰버는 한 발로 바닥을 딛고 달리는 자세이며, 찰리도 한 발을 내딛고 걷는 자세를 지면의 지지와 함께 자연스럽게 표현함."
   },
   {
    "label": "B",
    "direction": "현우와 앰버는 좌측 전경으로 달려오고, 찰리는 우측 전경에서 카메라를 등지고 그들을 마주보는 방향으로 서 있어 '뒤를 따르는' 관계가 아님.",
    "built_space": "도로가 가로로 뻗어 있고 점포가 배경 대부분을 차지하여, 점포를 우측 상단 배경에 배치하라는 구도 지시와 어긋남.",
    "entities": "앰버의 머리 위 고글이 누락되었고, 현우의 다리에는 지문에 없는 피 묻은 붕대가 그려짐. 찰리의 뒷모습은 기준에 부합함.",
    "hard_violations": [
     "[gemini-pro] 간판 및 창문에 읽을 수 있는 문자(매믹창명전소, OPEN)가 노출됨 (금지된 텍스트 포함)",
     "[gemini-pro] 현우의 다리에 치료받지 않은 상처 대신 붕대가 추가됨 (발명된 사물)",
     "[gpt-high] 찰리를 두 사람의 뒤가 아니라 진행 방향 앞에 세워 서로 마주하게 했다. 지정된 후행 배치를 반대로 만든다.",
     "[gpt-high] 판독 가능한 한글 간판과 영문 네온 문자를 노출하여 글자 금지 조건을 위반했다."
    ],
    "physics": "현우와 앰버의 달리는 자세는 바닥에 잘 지지되어 있으나, 찰리는 다리를 내딛지 않고 두 발로 멈춰 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캐릭터의 외형과 프레임 구도는 지시사항에 가깝게 구현되었으나, 텍스트 노출 금지를 어기고 일행이 서로 반대 방향으로 이동하여 '뒤를 따르는' 관계를 완전히 위반함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "텍스트 노출 금지 위반과 더불어 앰버의 고글 누락, 현우의 붕대 추가 등 객체 오류가 있으며, 인물들의 이동 방향이 서로 마주보고 있어 지시된 상황과 불일치함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 카메라를 향해 앞으로 달려오고 있으나, 찰리는 카메라를 등지고 반대 방향(배경 쪽)으로 걸어가고 있어 서로 엇갈리며 '뒤를 따르는' 방향성이 성립하지 않음.",
        "built_space": "도로가 깊이감을 주며 뻗어 있고, 지시된 무인 점포가 우측 상단 배경에 위치하여 프레임 구도를 충족함.",
        "entities": "현우(타박상, 의상), 앰버(머리 위 고글 포함), 찰리(코트, 모자, 뒷모습) 모두 레퍼런스의 외형에 부합함.",
        "hard_violations": [
         "간판 및 창문에 읽을 수 있는 문자(안작잡, OPEN)가 뚜렷하게 노출됨 (금지된 텍스트 포함)"
        ],
        "physics": "현우와 앰버는 한 발로 바닥을 딛고 달리는 자세이며, 찰리도 한 발을 내딛고 걷는 자세를 지면의 지지와 함께 자연스럽게 표현함."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 좌측 전경으로 달려오고, 찰리는 우측 전경에서 카메라를 등지고 그들을 마주보는 방향으로 서 있어 '뒤를 따르는' 관계가 아님.",
        "built_space": "도로가 가로로 뻗어 있고 점포가 배경 대부분을 차지하여, 점포를 우측 상단 배경에 배치하라는 구도 지시와 어긋남.",
        "entities": "앰버의 머리 위 고글이 누락되었고, 현우의 다리에는 지문에 없는 피 묻은 붕대가 그려짐. 찰리의 뒷모습은 기준에 부합함.",
        "hard_violations": [
         "간판 및 창문에 읽을 수 있는 문자(매믹창명전소, OPEN)가 노출됨 (금지된 텍스트 포함)",
         "현우의 다리에 치료받지 않은 상처 대신 붕대가 추가됨 (발명된 사물)"
        ],
        "physics": "현우와 앰버의 달리는 자세는 바닥에 잘 지지되어 있으나, 찰리는 다리를 내딛지 않고 두 발로 멈춰 서 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캐릭터의 외형과 프레임 구도는 지시사항에 가깝게 구현되었으나, 텍스트 노출 금지를 어기고 일행이 서로 반대 방향으로 이동하여 '뒤를 따르는' 관계를 완전히 위반함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "텍스트 노출 금지 위반과 더불어 앰버의 고글 누락, 현우의 붕대 추가 등 객체 오류가 있으며, 인물들의 이동 방향이 서로 마주보고 있어 지시된 상황과 불일치함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 카메라를 향해 앞으로 달려오고 있으나, 찰리는 카메라를 등지고 반대 방향(배경 쪽)으로 걸어가고 있어 서로 엇갈리며 '뒤를 따르는' 방향성이 성립하지 않음.",
        "built_space": "도로가 깊이감을 주며 뻗어 있고, 지시된 무인 점포가 우측 상단 배경에 위치하여 프레임 구도를 충족함.",
        "entities": "현우(타박상, 의상), 앰버(머리 위 고글 포함), 찰리(코트, 모자, 뒷모습) 모두 레퍼런스의 외형에 부합함.",
        "hard_violations": [
         "간판 및 창문에 읽을 수 있는 문자(안작잡, OPEN)가 뚜렷하게 노출됨 (금지된 텍스트 포함)"
        ],
        "physics": "현우와 앰버는 한 발로 바닥을 딛고 달리는 자세이며, 찰리도 한 발을 내딛고 걷는 자세를 지면의 지지와 함께 자연스럽게 표현함."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 좌측 전경으로 달려오고, 찰리는 우측 전경에서 카메라를 등지고 그들을 마주보는 방향으로 서 있어 '뒤를 따르는' 관계가 아님.",
        "built_space": "도로가 가로로 뻗어 있고 점포가 배경 대부분을 차지하여, 점포를 우측 상단 배경에 배치하라는 구도 지시와 어긋남.",
        "entities": "앰버의 머리 위 고글이 누락되었고, 현우의 다리에는 지문에 없는 피 묻은 붕대가 그려짐. 찰리의 뒷모습은 기준에 부합함.",
        "hard_violations": [
         "간판 및 창문에 읽을 수 있는 문자(매믹창명전소, OPEN)가 노출됨 (금지된 텍스트 포함)",
         "현우의 다리에 치료받지 않은 상처 대신 붕대가 추가됨 (발명된 사물)"
        ],
        "physics": "현우와 앰버의 달리는 자세는 바닥에 잘 지지되어 있으나, 찰리는 다리를 내딛지 않고 두 발로 멈춰 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우와 앰버가 앞을 가로막은 찰리 쪽으로 달려오는 배치여서 뒷모습의 동반 도주와 후행 관계를 뒤집었고, 판독 가능한 간판도 노출했다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "낮의 외곽 도로와 찰리의 걸음은 상대적으로 가깝지만, 두 사람은 카메라 쪽으로 달리고 찰리는 반대로 향해 후행 관계를 깨뜨렸으며 읽히는 간판도 남겼다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 얼굴과 가슴을 카메라 쪽으로 드러내며 화면 오른쪽 전경의 찰리를 향해 달린다. 두 사람의 시선도 대체로 찰리 쪽이다. 찰리는 카메라에 등을 보이고 두 사람이 오는 방향을 마주한다. 세 인물이 카메라에서 멀어지는 같은 경로로 달리고 찰리가 뒤따르는 장면이 아니다.",
        "built_space": "도로 건너 점포 한 곳이 화면 상단 중앙부터 오른쪽 가장자리까지 크게 차지하며, 유리 출입구와 여러 쇼윈도 구획, 내부 선반, 작은 네온 표지 하나가 보인다. 왼쪽에는 별도 건물 전면과 설비함이 있고 오른쪽 끝에는 파란 수거함이 있다. 점포는 비스듬히 보이지만 작은 상단 오른쪽 배경이라기보다 화면 전체의 주된 배경이다. 참고의 잡초와 전신주가 이어지는 개방된 도로보다 밀집된 상가 거리로 보인다. 불가능한 반사는 확인되지 않는다.",
        "entities": "등장인물은 세 명뿐이며 경찰은 없다. 현우는 젊은 동아시아계 남성으로 보이고 헝클어진 검은 머리, 회색 셔츠, 녹색 계열 바지와 얼굴의 상처가 맞는다. 다만 무릎에 피 묻은 흰 붕대가 있어 치료하지 않은 다리 상처와 다르다. 앰버는 금발 여자아이이고 남색 상의와 갈색 멜빵바지를 입지만 참고의 머리 위 장비는 보이지 않는다. 찰리는 모자와 낡은 코트, 베이지색 장갑판과 큰 금속 팔을 갖췄다. 신발 속 카드는 보이지 않으므로 판단 대상이 아니다. 간판의 한글과 네온의 영문 일부가 판독 가능하다.",
        "hard_violations": [
         "찰리를 두 사람의 뒤가 아니라 진행 방향 앞에 세워 서로 마주하게 했다. 지정된 후행 배치를 반대로 만든다.",
         "판독 가능한 한글 간판과 영문 네온 문자를 노출하여 글자 금지 조건을 위반했다."
        ],
        "physics": "앰버는 한쪽 부츠로 노면을 지지하고 반대쪽 다리를 뒤로 접어 달린다. 현우는 두 발이 잠깐 뜬 달리기의 비행 국면으로 보이며, 굽힌 뒷다리와 앞으로 내려오는 발이 이륙과 다음 착지를 설명한다. 근거 없는 부유는 아니다. 찰리는 양쪽 금속 발을 노면에 놓아 안정적으로 서 있지만, 육중한 다리를 길게 내딛으며 따라가는 동작은 없다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 점포를 등지고 카메라를 향해 달리며 시선도 전방인 카메라 쪽으로 향한다. 왼쪽의 찰리는 등을 보인 채 점포 방향으로 걷는다. 따라서 찰리는 두 사람과 반대 방향으로 이동하며, 세 사람 모두 점포 쪽 도주로를 따라 멀어지는 뒷모습이라는 지시와 맞지 않는다.",
        "built_space": "점포 한 곳이 길 건너 상단 중앙과 오른쪽을 차지한다. 중앙 유리 출입구와 양옆 유리창, 가로 간판 하나, 왼쪽 세로 간판 하나, 네온 표지 두 개와 창 아래 상자 묶음들이 보인다. 점포 앞 횡방향 도로와 카메라 쪽으로 이어지는 거친 포장면이 공간을 분리한다. 전신주, 전선, 마른 풀과 낮은 주변 건물은 참고의 외곽 도로 분위기에 더 가깝다. 다만 점포 전면이 거의 정면이고 크게 보여 상단 오른쪽의 비스듬한 종속 배경이라는 요구에는 못 미친다. 참고에서 점포 설비의 개수는 확인되지 않아 중복 여부를 단정할 수 없다.",
        "entities": "세 인물만 등장하고 경찰은 없다. 현우의 젊은 동아시아계 외모, 검은 머리, 회색 셔츠와 녹색 계열 바지는 참고와 대체로 맞고 얼굴의 상처도 보인다. 다리 상처는 뚜렷하지 않다. 앰버는 금발 여자아이로 남색 상의, 갈색 멜빵바지와 머리 위 검은 장비를 갖췄고 지친 표정을 보이지만 눈물 자국은 분명하지 않다. 찰리의 낡은 코트와 모자, 베이지색 금속 팔과 다리가 보이며 얼굴은 뒷모습에 가려져 있다. 신발 속 카드는 확인할 수 없다. 점포에는 읽을 수 있는 한글 음절과 영문 네온 문자가 있다.",
        "hard_violations": [
         "찰리를 도주하는 두 사람과 반대 방향으로 걷게 하여 같은 경로에서 뒤따른다는 필수 배치를 위반했다.",
         "큰 간판의 한글과 두 네온 표지의 영문이 읽혀 글자 금지 조건을 위반했다."
        ],
        "physics": "현우와 앰버는 뒤로 접힌 다리와 착지를 준비하는 앞발, 굽힌 팔로 달리기의 짧은 공중 국면을 보여준다. 발밑 그림자와 노면의 거리가 작고 다음 착지 위치도 자연스러워 설명 없는 부유로 보이지 않는다. 찰리는 한쪽 금속 발로 체중을 지지하고 다른 발을 들어 옮기는 보행 자세다. 물리적으로 가능한 걸음이지만 요구한 길고 육중한 보폭은 약하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우와 앰버가 앞을 가로막은 찰리 쪽으로 달려오는 배치여서 뒷모습의 동반 도주와 후행 관계를 뒤집었고, 판독 가능한 간판도 노출했다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "낮의 외곽 도로와 찰리의 걸음은 상대적으로 가깝지만, 두 사람은 카메라 쪽으로 달리고 찰리는 반대로 향해 후행 관계를 깨뜨렸으며 읽히는 간판도 남겼다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 앰버는 얼굴과 가슴을 카메라 쪽으로 드러내며 화면 오른쪽 전경의 찰리를 향해 달린다. 두 사람의 시선도 대체로 찰리 쪽이다. 찰리는 카메라에 등을 보이고 두 사람이 오는 방향을 마주한다. 세 인물이 카메라에서 멀어지는 같은 경로로 달리고 찰리가 뒤따르는 장면이 아니다.",
        "built_space": "도로 건너 점포 한 곳이 화면 상단 중앙부터 오른쪽 가장자리까지 크게 차지하며, 유리 출입구와 여러 쇼윈도 구획, 내부 선반, 작은 네온 표지 하나가 보인다. 왼쪽에는 별도 건물 전면과 설비함이 있고 오른쪽 끝에는 파란 수거함이 있다. 점포는 비스듬히 보이지만 작은 상단 오른쪽 배경이라기보다 화면 전체의 주된 배경이다. 참고의 잡초와 전신주가 이어지는 개방된 도로보다 밀집된 상가 거리로 보인다. 불가능한 반사는 확인되지 않는다.",
        "entities": "등장인물은 세 명뿐이며 경찰은 없다. 현우는 젊은 동아시아계 남성으로 보이고 헝클어진 검은 머리, 회색 셔츠, 녹색 계열 바지와 얼굴의 상처가 맞는다. 다만 무릎에 피 묻은 흰 붕대가 있어 치료하지 않은 다리 상처와 다르다. 앰버는 금발 여자아이이고 남색 상의와 갈색 멜빵바지를 입지만 참고의 머리 위 장비는 보이지 않는다. 찰리는 모자와 낡은 코트, 베이지색 장갑판과 큰 금속 팔을 갖췄다. 신발 속 카드는 보이지 않으므로 판단 대상이 아니다. 간판의 한글과 네온의 영문 일부가 판독 가능하다.",
        "hard_violations": [
         "찰리를 두 사람의 뒤가 아니라 진행 방향 앞에 세워 서로 마주하게 했다. 지정된 후행 배치를 반대로 만든다.",
         "판독 가능한 한글 간판과 영문 네온 문자를 노출하여 글자 금지 조건을 위반했다."
        ],
        "physics": "앰버는 한쪽 부츠로 노면을 지지하고 반대쪽 다리를 뒤로 접어 달린다. 현우는 두 발이 잠깐 뜬 달리기의 비행 국면으로 보이며, 굽힌 뒷다리와 앞으로 내려오는 발이 이륙과 다음 착지를 설명한다. 근거 없는 부유는 아니다. 찰리는 양쪽 금속 발을 노면에 놓아 안정적으로 서 있지만, 육중한 다리를 길게 내딛으며 따라가는 동작은 없다."
       },
       {
        "label": "A",
        "direction": "현우와 앰버는 점포를 등지고 카메라를 향해 달리며 시선도 전방인 카메라 쪽으로 향한다. 왼쪽의 찰리는 등을 보인 채 점포 방향으로 걷는다. 따라서 찰리는 두 사람과 반대 방향으로 이동하며, 세 사람 모두 점포 쪽 도주로를 따라 멀어지는 뒷모습이라는 지시와 맞지 않는다.",
        "built_space": "점포 한 곳이 길 건너 상단 중앙과 오른쪽을 차지한다. 중앙 유리 출입구와 양옆 유리창, 가로 간판 하나, 왼쪽 세로 간판 하나, 네온 표지 두 개와 창 아래 상자 묶음들이 보인다. 점포 앞 횡방향 도로와 카메라 쪽으로 이어지는 거친 포장면이 공간을 분리한다. 전신주, 전선, 마른 풀과 낮은 주변 건물은 참고의 외곽 도로 분위기에 더 가깝다. 다만 점포 전면이 거의 정면이고 크게 보여 상단 오른쪽의 비스듬한 종속 배경이라는 요구에는 못 미친다. 참고에서 점포 설비의 개수는 확인되지 않아 중복 여부를 단정할 수 없다.",
        "entities": "세 인물만 등장하고 경찰은 없다. 현우의 젊은 동아시아계 외모, 검은 머리, 회색 셔츠와 녹색 계열 바지는 참고와 대체로 맞고 얼굴의 상처도 보인다. 다리 상처는 뚜렷하지 않다. 앰버는 금발 여자아이로 남색 상의, 갈색 멜빵바지와 머리 위 검은 장비를 갖췄고 지친 표정을 보이지만 눈물 자국은 분명하지 않다. 찰리의 낡은 코트와 모자, 베이지색 금속 팔과 다리가 보이며 얼굴은 뒷모습에 가려져 있다. 신발 속 카드는 확인할 수 없다. 점포에는 읽을 수 있는 한글 음절과 영문 네온 문자가 있다.",
        "hard_violations": [
         "찰리를 도주하는 두 사람과 반대 방향으로 걷게 하여 같은 경로에서 뒤따른다는 필수 배치를 위반했다.",
         "큰 간판의 한글과 두 네온 표지의 영문이 읽혀 글자 금지 조건을 위반했다."
        ],
        "physics": "현우와 앰버는 뒤로 접힌 다리와 착지를 준비하는 앞발, 굽힌 팔로 달리기의 짧은 공중 국면을 보여준다. 발밑 그림자와 노면의 거리가 작고 다음 착지 위치도 자연스러워 설명 없는 부유로 보이지 않는다. 찰리는 한쪽 금속 발로 체중을 지지하고 다른 발을 들어 옮기는 보행 자세다. 물리적으로 가능한 걸음이지만 요구한 길고 육중한 보폭은 약하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.417
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "A": [
     "[gemini-pro] 간판 및 창문에 읽을 수 있는 문자(안작잡, OPEN)가 뚜렷하게 노출됨 (금지된 텍스트 포함)",
     "[gpt-high] 찰리를 도주하는 두 사람과 반대 방향으로 걷게 하여 같은 경로에서 뒤따른다는 필수 배치를 위반했다.",
     "[gpt-high] 큰 간판의 한글과 두 네온 표지의 영문이 읽혀 글자 금지 조건을 위반했다."
    ],
    "B": [
     "[gemini-pro] 간판 및 창문에 읽을 수 있는 문자(매믹창명전소, OPEN)가 노출됨 (금지된 텍스트 포함)",
     "[gemini-pro] 현우의 다리에 치료받지 않은 상처 대신 붕대가 추가됨 (발명된 사물)",
     "[gpt-high] 찰리를 두 사람의 뒤가 아니라 진행 방향 앞에 세워 서로 마주하게 했다. 지정된 후행 배치를 반대로 만든다.",
     "[gpt-high] 판독 가능한 한글 간판과 영문 네온 문자를 노출하여 글자 금지 조건을 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "캐릭터의 외형과 프레임 구도는 지시사항에 가깝게 구현되었으나, 텍스트 노출 금지를 어기고 일행이 서로 반대 방향으로 이동하여 '뒤를 따르는' 관계를 완전히 위반함.  ★위반: [gemini-pro] 간판 및 창문에 읽을 수 있는 문자(안작잡, OPEN)가 뚜렷하게 노출됨 (금지된 텍스트 포함) / [gpt-high] 찰리를 도주하는 두 사람과 반대 방향으로 걷게 하여 같은 경로에서 뒤따른다는 필수 배치를 위반했다. / [gpt-high] 큰 간판의 한글과 두 네온 표지의 영문이 읽혀 글자 금지 조건을 위반했다."
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "텍스트 노출 금지 위반과 더불어 앰버의 고글 누락, 현우의 붕대 추가 등 객체 오류가 있으며, 인물들의 이동 방향이 서로 마주보고 있어 지시된 상황과 불일치함.  ★위반: [gemini-pro] 간판 및 창문에 읽을 수 있는 문자(매믹창명전소, OPEN)가 노출됨 (금지된 텍스트 포함) / [gemini-pro] 현우의 다리에 치료받지 않은 상처 대신 붕대가 추가됨 (발명된 사물) / [gpt-high] 찰리를 두 사람의 뒤가 아니라 진행 방향 앞에 세워 서로 마주하게 했다. 지정된 후행 배치를 반대로 만든다. / [gpt-high] 판독 가능한 한글 간판과 영문 네온 문자를 노출하여 글자 금지 조건을 위반했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 앰버, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh4_sel.png",
    "asset_id": "4f61ca20-52bc-4eff-a722-1cde8eb30f6c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-7758-77f8-84f8-ff74c294a690",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S40sh4"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S40sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:49:31.300239+00:00",
  "fingerprint": "b30532f6c36e1a95d2697520b5699f69f042ac581aeac599beb63132d7a05f5c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S40sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S40sh6_sel.png",
  "source_sha256": "b2e823d4d8e6bd9261872c0fd422e0d49835c3ba6f7a2d907c20ded14882dfb1",
  "file": "S40sh6_cine.png",
  "staged_sha256": "e1bf3c06b36d1d3e6e7ec6de1d5c437a7300025412513a0e2b7446be4aa46bfa",
  "latency_ms": 10267
 },
 "S41sh14::signage": {
  "fp": "14c8720cad146e81",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S41sh14": {
  "input_fingerprint": "3779d8a018c9684c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM now has a fist-sized breach in its front. The scavenged duffel bag and backpack contain a disposable phone, a map, food, drinks and the remaining COPD medicine. 찰리: His metal fist is embedded in the ATM front. He still has the old coat and hat; the store blanket has not yet been handed over.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM now has a fist-sized breach in its front. The scavenged duffel bag and backpack contain a disposable phone, a map, food, drinks and the remaining COPD medicine. 찰리: His metal fist is embedded in the ATM front. He still has the old coat and hat; the store blanket has not yet been handed over.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM now has a fist-sized breach in its front. The scavenged duffel bag and backpack contain a disposable phone, a map, food, drinks and the remaining COPD medicine. 찰리: His metal fist is embedded in the ATM front. He still has the old coat and hat; the store blanket has not yet been handed over.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14__bgfirst_bg.png",
     "asset_id": "6398fb6a-6f79-4ad1-90ee-719e7320869a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S41sh14.png",
     "asset_id": "2b2818ac-406a-4c9d-805c-3eb8df283def",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L195B01.png",
     "asset_id": "830dae63-3f71-4c94-9045-5c77a027e48e",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 팔이 인출기 앞면을 향해 뻗어 주먹을 꽂아 넣음.",
    "built_space": "상점 내부. 우측에 인출기, 좌측 배경에 유리문이 위치하여 공간 구조와 시점이 적절함.",
    "entities": "찰리(오래된 코트 착용, 베이지색 장갑판 팔 일치), 파손된 현금인출기 일치. 금지된 텍스트 없음.",
    "hard_violations": [],
    "physics": "팔은 몸체에 연결되어 지지되며, 유리 파편은 타격 충격으로 인해 공중에 튀어오름."
   },
   {
    "label": "B",
    "direction": "찰리의 팔이 인출기 전면 화면 부위를 향해 주먹을 꽂음.",
    "built_space": "상점 내부. 좌측 벽면, 중앙 인출기, 우측 진열대가 레퍼런스의 배치와 정확히 일치함.",
    "entities": "찰리(지시와 달리 줄무늬 담요 착용 불일치, 베이지색 장갑판 팔 일치), 현금인출기 일치. 읽을 수 있는 숫자(32,000) 존재.",
    "hard_violations": [
     "[gemini-pro] 공간에 하나뿐인 인출기 파손 구멍이 두 개로 중복 생성됨 (레퍼런스의 기존 구멍과 주먹이 만든 새 구멍이 동시 존재)",
     "[gpt-high] 오른쪽 진열대 가격표의 가격 숫자가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반합니다."
    ],
    "physics": "팔은 몸에 지지되며, 파편들은 기기 틈새와 바닥에 놓임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 의상(오래된 코트)과 텍스트 제약을 잘 지켰으나, 팔이 2D 그래픽처럼 묘사되어 질감 현실성 지침을 위반함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "실사 렌더링은 우수하나, 의상 지침을 무시하고 담요를 입혔으며, 인출기의 파손 구멍이 중복 생성되고 금지된 텍스트가 포함됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔이 인출기 앞면을 향해 뻗어 주먹을 꽂아 넣음.",
        "built_space": "상점 내부. 우측에 인출기, 좌측 배경에 유리문이 위치하여 공간 구조와 시점이 적절함.",
        "entities": "찰리(오래된 코트 착용, 베이지색 장갑판 팔 일치), 파손된 현금인출기 일치. 금지된 텍스트 없음.",
        "hard_violations": [],
        "physics": "팔은 몸체에 연결되어 지지되며, 유리 파편은 타격 충격으로 인해 공중에 튀어오름."
       },
       {
        "label": "B",
        "direction": "찰리의 팔이 인출기 전면 화면 부위를 향해 주먹을 꽂음.",
        "built_space": "상점 내부. 좌측 벽면, 중앙 인출기, 우측 진열대가 레퍼런스의 배치와 정확히 일치함.",
        "entities": "찰리(지시와 달리 줄무늬 담요 착용 불일치, 베이지색 장갑판 팔 일치), 현금인출기 일치. 읽을 수 있는 숫자(32,000) 존재.",
        "hard_violations": [
         "공간에 하나뿐인 인출기 파손 구멍이 두 개로 중복 생성됨 (레퍼런스의 기존 구멍과 주먹이 만든 새 구멍이 동시 존재)"
        ],
        "physics": "팔은 몸에 지지되며, 파편들은 기기 틈새와 바닥에 놓임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 의상(오래된 코트)과 텍스트 제약을 잘 지켰으나, 팔이 2D 그래픽처럼 묘사되어 질감 현실성 지침을 위반함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "실사 렌더링은 우수하나, 의상 지침을 무시하고 담요를 입혔으며, 인출기의 파손 구멍이 중복 생성되고 금지된 텍스트가 포함됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔이 인출기 앞면을 향해 뻗어 주먹을 꽂아 넣음.",
        "built_space": "상점 내부. 우측에 인출기, 좌측 배경에 유리문이 위치하여 공간 구조와 시점이 적절함.",
        "entities": "찰리(오래된 코트 착용, 베이지색 장갑판 팔 일치), 파손된 현금인출기 일치. 금지된 텍스트 없음.",
        "hard_violations": [],
        "physics": "팔은 몸체에 연결되어 지지되며, 유리 파편은 타격 충격으로 인해 공중에 튀어오름."
       },
       {
        "label": "B",
        "direction": "찰리의 팔이 인출기 전면 화면 부위를 향해 주먹을 꽂음.",
        "built_space": "상점 내부. 좌측 벽면, 중앙 인출기, 우측 진열대가 레퍼런스의 배치와 정확히 일치함.",
        "entities": "찰리(지시와 달리 줄무늬 담요 착용 불일치, 베이지색 장갑판 팔 일치), 현금인출기 일치. 읽을 수 있는 숫자(32,000) 존재.",
        "hard_violations": [
         "공간에 하나뿐인 인출기 파손 구멍이 두 개로 중복 생성됨 (레퍼런스의 기존 구멍과 주먹이 만든 새 구멍이 동시 존재)"
        ],
        "physics": "팔은 몸에 지지되며, 파편들은 기기 틈새와 바닥에 놓임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "타격부를 가깝게 잡았지만 읽히는 가격 숫자가 금지 조건을 위반하고, 주먹 옆의 별도 구멍과 이미 두른 줄무늬 담요도 지정된 파손·소지 상태와 어긋납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "사선으로 보이는 현금인출기 앞면에 금속 주먹이 박힌 순간과 낡은 외투를 잘 구현했으나, 몸통과 매장 바닥의 비중이 커 요구된 밀착 구도보다 다소 넓습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 왼쪽의 팔이 오른쪽 현금인출기 조작부를 향해 뻗어 있고, 주먹은 찢어진 화면 부근에 걸쳐 들어가 있습니다. 그 오른쪽에는 주먹이 들어가지 않은 별도의 구멍이 보입니다. 얼굴과 시선은 프레임 밖입니다.",
        "built_space": "주된 회색 현금인출기 한 대의 앞면과 오른쪽 측면이 보이고, 왼쪽 뒤에는 참조에도 있는 파란 장치 일부가 보입니다. 오른쪽 가장자리에는 담요 진열대 하나, 왼쪽 배경에는 유리 출입문이 있습니다. 장소의 주요 재료와 배치는 참조에 가깝지만, 화면 부근 파손과 오른쪽 외장 구멍이 따로 있어 단일 주먹 크기 관통부보다 파손 범위가 큽니다.",
        "entities": "한 인물의 몸통 일부와 굵은 금속 팔만 보입니다. 마모된 샌드 베이지 장갑판과 검은 관절은 찰리 참조와 잘 맞습니다. 얼굴과 모자는 보이지 않아 판단 대상이 아닙니다. 왼쪽 몸통에 줄무늬 담요가 이미 둘러져 있어 아직 담요를 받지 않았다는 조건과 다릅니다. 현금인출기와 상품 진열대는 식별되며, 오른쪽 가격표에는 읽을 수 있는 숫자가 남아 있습니다.",
        "hard_violations": [
         "오른쪽 진열대 가격표의 가격 숫자가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반합니다."
        ],
        "physics": "주먹은 손목·전완·팔꿈치·상완으로 이어져 왼쪽 몸통에 연결되어 있습니다. 찢어진 외장판은 기기에 붙어 있고 작은 파편은 아래 조작 선반에 놓여 있어 지지가 확인됩니다. 주먹의 진입 자체는 가능하지만, 옆으로 떨어진 별도 구멍까지 이 주먹의 현재 타격으로 생겼다는 관계는 명확하지 않습니다."
       },
       {
        "label": "B",
        "direction": "왼쪽 어깨에서 뻗은 팔이 오른쪽 현금인출기 앞면으로 향하고, 손목 너머의 주먹이 파손 구멍 안에 박혀 있습니다. 타격 방향과 관통 지점이 일치하며 아직 주먹을 빼지 않은 상태입니다. 얼굴과 시선은 보이지 않습니다.",
        "built_space": "회색 현금인출기 한 대의 앞면과 인접 오른쪽 측면을 비스듬히 보여줍니다. 왼쪽 뒤에는 파란 장치 일부, 배경에는 유리 출입문과 높은 창, 상품 선반 일부가 있고 오른쪽 끝에는 담요 진열대가 있습니다. 참조 장소의 구성과 맞지만, 바닥과 상체가 차지하는 면적이 커 타격부 중심의 클로즈업으로는 다소 넓습니다.",
        "entities": "한 인물의 몸통 일부와 육중한 금속 팔이 보이며, 샌드 베이지 장갑판과 노출된 기계 관절이 찰리의 외형에 부합합니다. 낡은 갈색 외투와 상단의 모자 챙 일부가 보이고 줄무늬 매장 담요는 두르지 않았습니다. 얼굴·다리·가방은 프레임 밖이므로 평가하지 않습니다. 현금인출기 앞면의 관통부와 상품 진열대가 확인되며 명확하게 읽히는 문구는 없습니다.",
        "hard_violations": [],
        "physics": "팔은 외투 안의 어깨에서 손목과 주먹까지 연속적으로 연결되고, 주먹은 기기 내부에 들어가 있습니다. 구멍 가장자리의 판재는 외장에 붙어 있습니다. 공중의 파편은 바로 옆 관통부에서 타격으로 튀어나온 순간으로 설명되므로 지지 없는 정지 물체로 볼 근거는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "타격부를 가깝게 잡았지만 읽히는 가격 숫자가 금지 조건을 위반하고, 주먹 옆의 별도 구멍과 이미 두른 줄무늬 담요도 지정된 파손·소지 상태와 어긋납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "사선으로 보이는 현금인출기 앞면에 금속 주먹이 박힌 순간과 낡은 외투를 잘 구현했으나, 몸통과 매장 바닥의 비중이 커 요구된 밀착 구도보다 다소 넓습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 왼쪽의 팔이 오른쪽 현금인출기 조작부를 향해 뻗어 있고, 주먹은 찢어진 화면 부근에 걸쳐 들어가 있습니다. 그 오른쪽에는 주먹이 들어가지 않은 별도의 구멍이 보입니다. 얼굴과 시선은 프레임 밖입니다.",
        "built_space": "주된 회색 현금인출기 한 대의 앞면과 오른쪽 측면이 보이고, 왼쪽 뒤에는 참조에도 있는 파란 장치 일부가 보입니다. 오른쪽 가장자리에는 담요 진열대 하나, 왼쪽 배경에는 유리 출입문이 있습니다. 장소의 주요 재료와 배치는 참조에 가깝지만, 화면 부근 파손과 오른쪽 외장 구멍이 따로 있어 단일 주먹 크기 관통부보다 파손 범위가 큽니다.",
        "entities": "한 인물의 몸통 일부와 굵은 금속 팔만 보입니다. 마모된 샌드 베이지 장갑판과 검은 관절은 찰리 참조와 잘 맞습니다. 얼굴과 모자는 보이지 않아 판단 대상이 아닙니다. 왼쪽 몸통에 줄무늬 담요가 이미 둘러져 있어 아직 담요를 받지 않았다는 조건과 다릅니다. 현금인출기와 상품 진열대는 식별되며, 오른쪽 가격표에는 읽을 수 있는 숫자가 남아 있습니다.",
        "hard_violations": [
         "오른쪽 진열대 가격표의 가격 숫자가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반합니다."
        ],
        "physics": "주먹은 손목·전완·팔꿈치·상완으로 이어져 왼쪽 몸통에 연결되어 있습니다. 찢어진 외장판은 기기에 붙어 있고 작은 파편은 아래 조작 선반에 놓여 있어 지지가 확인됩니다. 주먹의 진입 자체는 가능하지만, 옆으로 떨어진 별도 구멍까지 이 주먹의 현재 타격으로 생겼다는 관계는 명확하지 않습니다."
       },
       {
        "label": "A",
        "direction": "왼쪽 어깨에서 뻗은 팔이 오른쪽 현금인출기 앞면으로 향하고, 손목 너머의 주먹이 파손 구멍 안에 박혀 있습니다. 타격 방향과 관통 지점이 일치하며 아직 주먹을 빼지 않은 상태입니다. 얼굴과 시선은 보이지 않습니다.",
        "built_space": "회색 현금인출기 한 대의 앞면과 인접 오른쪽 측면을 비스듬히 보여줍니다. 왼쪽 뒤에는 파란 장치 일부, 배경에는 유리 출입문과 높은 창, 상품 선반 일부가 있고 오른쪽 끝에는 담요 진열대가 있습니다. 참조 장소의 구성과 맞지만, 바닥과 상체가 차지하는 면적이 커 타격부 중심의 클로즈업으로는 다소 넓습니다.",
        "entities": "한 인물의 몸통 일부와 육중한 금속 팔이 보이며, 샌드 베이지 장갑판과 노출된 기계 관절이 찰리의 외형에 부합합니다. 낡은 갈색 외투와 상단의 모자 챙 일부가 보이고 줄무늬 매장 담요는 두르지 않았습니다. 얼굴·다리·가방은 프레임 밖이므로 평가하지 않습니다. 현금인출기 앞면의 관통부와 상품 진열대가 확인되며 명확하게 읽히는 문구는 없습니다.",
        "hard_violations": [],
        "physics": "팔은 외투 안의 어깨에서 손목과 주먹까지 연속적으로 연결되고, 주먹은 기기 내부에 들어가 있습니다. 구멍 가장자리의 판재는 외장에 붙어 있습니다. 공중의 파편은 바로 옆 관통부에서 타격으로 튀어나온 순간으로 설명되므로 지지 없는 정지 물체로 볼 근거는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.875
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.625
   },
   "violations": {
    "B": [
     "[gemini-pro] 공간에 하나뿐인 인출기 파손 구멍이 두 개로 중복 생성됨 (레퍼런스의 기존 구멍과 주먹이 만든 새 구멍이 동시 존재)",
     "[gpt-high] 오른쪽 진열대 가격표의 가격 숫자가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 625
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 의상(오래된 코트)과 텍스트 제약을 잘 지켰으나, 팔이 2D 그래픽처럼 묘사되어 질감 현실성 지침을 위반함."
   },
   {
    "label": "B",
    "score": 625,
    "verdict_ko": "실사 렌더링은 우수하나, 의상 지침을 무시하고 담요를 입혔으며, 인출기의 파손 구멍이 중복 생성되고 금지된 텍스트가 포함됨.  ★위반: [gemini-pro] 공간에 하나뿐인 인출기 파손 구멍이 두 개로 중복 생성됨 (레퍼런스의 기존 구멍과 주먹이 만든 새 구멍이 동시 존재) / [gpt-high] 오른쪽 진열대 가격표의 가격 숫자가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L195B01.png",
    "asset_id": "830dae63-3f71-4c94-9045-5c77a027e48e",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-7913-7cab-bf5c-4782da1b2412",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14__bgfirst_bg.png",
   "bg_asset_id": "6398fb6a-6f79-4ad1-90ee-719e7320869a",
   "bg_record_key": "S41sh14::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S41sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:51:58.161835+00:00",
  "fingerprint": "1d76538035ca42824b2ba94c7226f214be8feff49aa4771f824a39b2e5131adf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S41sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S41sh14_sel.png",
  "source_sha256": "32e328ebbbaccc0871de7827786e28771aa880535efe37efe4e7ff1fb6bc3c56",
  "file": "S41sh14_cine.png",
  "staged_sha256": "2fbd18c3806e98afeba6e37d7cac2745179771bc67292dad2713d237d6b31213",
  "latency_ms": 16832
 },
 "S41sh22::signage": {
  "fp": "39e75a7ed2eaf5aa",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S41sh22": {
  "input_fingerprint": "41c33679a9dbed20",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 진열대 뒤에 웅크려 몸을 숨긴 채 바깥쪽을 긴장된 눈빛으로 엿보는 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Behind a merchandise shelving unit inside the unattended shop, looking toward the entrance in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Display end leading toward the off-screen entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 몸을 숨긴 진열대 (Separates the crouching group from the entrance approach) — Its concealed side faces the camera; its end at screen right marks the route toward the entrance; used as Creates a visible boundary between safety and exposure without obscuring the group's full bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the store's established ambient illumination and restrained contrast, allowing concealment to come from the display rather than an invented lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM remains breached, with cash spilled from the opening and some already collected in a bag. The store entrance has been opened, and the scavenged supplies remain packed. 현우: He crouches behind a display shelf with the bag containing supplies and collected cash. His facial bruises and leg wound remain, and the contact card is still hidden in his shoe. 앰버: She crouches behind a display shelf, still wearing the scavenged bag and retaining the replacement shoes. 찰리: He is concealed behind a display shelf and now has the store blanket for covering his body, in addition to the earlier coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 진열대 뒤에 웅크려 몸을 숨긴 채 바깥쪽을 긴장된 눈빛으로 엿보는 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Behind a merchandise shelving unit inside the unattended shop, looking toward the entrance in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Display end leading toward the off-screen entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 몸을 숨긴 진열대 (Separates the crouching group from the entrance approach) — Its concealed side faces the camera; its end at screen right marks the route toward the entrance; used as Creates a visible boundary between safety and exposure without obscuring the group's full bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the store's established ambient illumination and restrained contrast, allowing concealment to come from the display rather than an invented lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM remains breached, with cash spilled from the opening and some already collected in a bag. The store entrance has been opened, and the scavenged supplies remain packed. 현우: He crouches behind a display shelf with the bag containing supplies and collected cash. His facial bruises and leg wound remain, and the contact card is still hidden in his shoe. 앰버: She crouches behind a display shelf, still wearing the scavenged bag and retaining the replacement shoes. 찰리: He is concealed behind a display shelf and now has the store blanket for covering his body, in addition to the earlier coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 진열대 뒤에 웅크려 몸을 숨긴 채 바깥쪽을 긴장된 눈빛으로 엿보는 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Behind a merchandise shelving unit inside the unattended shop, looking toward the entrance in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Display end leading toward the off-screen entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 몸을 숨긴 진열대 (Separates the crouching group from the entrance approach) — Its concealed side faces the camera; its end at screen right marks the route toward the entrance; used as Creates a visible boundary between safety and exposure without obscuring the group's full bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the store's established ambient illumination and restrained contrast, allowing concealment to come from the display rather than an invented lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM remains breached, with cash spilled from the opening and some already collected in a bag. The store entrance has been opened, and the scavenged supplies remain packed. 현우: He crouches behind a display shelf with the bag containing supplies and collected cash. His facial bruises and leg wound remain, and the contact card is still hidden in his shoe. 앰버: She crouches behind a display shelf, still wearing the scavenged bag and retaining the replacement shoes. 찰리: He is concealed behind a display shelf and now has the store blanket for covering his body, in addition to the earlier coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우, 앰버, 찰리 세 명 모두 프레임 오른쪽의 출입구 방향을 주시하고 있습니다.",
    "built_space": "상점 내부로, 왼쪽에 진열대가 있고 오른쪽에 파손된 ATM이 배치되어 있으며 인물들은 진열대 옆에 위치해 있습니다.",
    "entities": "얼굴에 상처가 있고 자루를 든 현우, 금발의 앰버, 모자와 담요를 두른 고릴라 형태의 로봇 찰리가 모두 제시된 외형과 일치합니다.",
    "hard_violations": [
     "[gpt-high] 전경 상품 포장에 판독 가능한 한글 문자가 노출되어 문자 금지 조건을 위반한다."
    ],
    "physics": "세 인물 모두 바닥에 발과 무릎을 대고 안정적으로 지탱하며 웅크린 자세를 유지하고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우와 앰버, 우측의 찰리는 오른쪽을 주시하지만, 좌측에 나타난 또 다른 찰리는 화면 왼쪽을 향하고 있습니다.",
    "built_space": "상점 중앙에 철망 형태의 진열대가 있고 오른쪽에 부서진 ATM과 출입구가 위치합니다.",
    "entities": "현우와 앰버의 외형은 일치하나, 로봇 찰리가 좌우에 각각 한 명씩 등장하여 지시와 다릅니다.",
    "hard_violations": [
     "[gemini-pro] 중복된 캐릭터 (찰리가 두 명 등장함)",
     "[gpt-high] 찰리의 몸이 두 개로 복제되어 허용되지 않은 네 번째 인물이 등장한다.",
     "[gpt-high] 현금자동입출금기 상단에 읽을 수 있는 ‘ATM’ 문자가 노출되어 있다."
    ],
    "physics": "인물들은 바닥에 발을 딛고 무릎을 굽힌 채 웅크려 체중을 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 인물 구성, 진열대 뒤에 웅크린 자세, 시선 방향을 명확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "특정 캐릭터(찰리)가 양옆에 두 명으로 중복 생성되는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우, 앰버, 찰리 세 명 모두 프레임 오른쪽의 출입구 방향을 주시하고 있습니다.",
        "built_space": "상점 내부로, 왼쪽에 진열대가 있고 오른쪽에 파손된 ATM이 배치되어 있으며 인물들은 진열대 옆에 위치해 있습니다.",
        "entities": "얼굴에 상처가 있고 자루를 든 현우, 금발의 앰버, 모자와 담요를 두른 고릴라 형태의 로봇 찰리가 모두 제시된 외형과 일치합니다.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 발과 무릎을 대고 안정적으로 지탱하며 웅크린 자세를 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버, 우측의 찰리는 오른쪽을 주시하지만, 좌측에 나타난 또 다른 찰리는 화면 왼쪽을 향하고 있습니다.",
        "built_space": "상점 중앙에 철망 형태의 진열대가 있고 오른쪽에 부서진 ATM과 출입구가 위치합니다.",
        "entities": "현우와 앰버의 외형은 일치하나, 로봇 찰리가 좌우에 각각 한 명씩 등장하여 지시와 다릅니다.",
        "hard_violations": [
         "중복된 캐릭터 (찰리가 두 명 등장함)"
        ],
        "physics": "인물들은 바닥에 발을 딛고 무릎을 굽힌 채 웅크려 체중을 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 인물 구성, 진열대 뒤에 웅크린 자세, 시선 방향을 명확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "특정 캐릭터(찰리)가 양옆에 두 명으로 중복 생성되는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우, 앰버, 찰리 세 명 모두 프레임 오른쪽의 출입구 방향을 주시하고 있습니다.",
        "built_space": "상점 내부로, 왼쪽에 진열대가 있고 오른쪽에 파손된 ATM이 배치되어 있으며 인물들은 진열대 옆에 위치해 있습니다.",
        "entities": "얼굴에 상처가 있고 자루를 든 현우, 금발의 앰버, 모자와 담요를 두른 고릴라 형태의 로봇 찰리가 모두 제시된 외형과 일치합니다.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 발과 무릎을 대고 안정적으로 지탱하며 웅크린 자세를 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버, 우측의 찰리는 오른쪽을 주시하지만, 좌측에 나타난 또 다른 찰리는 화면 왼쪽을 향하고 있습니다.",
        "built_space": "상점 중앙에 철망 형태의 진열대가 있고 오른쪽에 부서진 ATM과 출입구가 위치합니다.",
        "entities": "현우와 앰버의 외형은 일치하나, 로봇 찰리가 좌우에 각각 한 명씩 등장하여 지시와 다릅니다.",
        "hard_violations": [
         "중복된 캐릭터 (찰리가 두 명 등장함)"
        ],
        "physics": "인물들은 바닥에 발을 딛고 무릎을 굽힌 채 웅크려 체중을 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "찰리가 두 명으로 복제된 치명적 오류가 있으며, 일행의 시선도 오른쪽 뒤 출입구가 아닌 왼쪽을 향한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "세 인물과 열린 출입문, 찰리의 변장 및 앰버의 가방은 더 충실하지만, 출입구 반대쪽을 엿보며 진열대가 전신을 가리는 구도는 지시와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 화면 왼쪽 바깥을 바라본다. 왼쪽 찰리는 정면 아래쪽, 오른쪽 찰리는 일행 앞쪽으로 얼굴을 향한다. 실제 출입구는 화면 오른쪽 뒤에 있어 누구의 경계 방향도 그 출입구와 맞지 않는다. 무기나 겨냥하는 도구는 없다.",
        "built_space": "중앙에 철망형 상품 진열대 한 조, 왼쪽 전경에 별도 선반, 왼쪽 뒤에 계산대, 오른쪽에 파손된 현금자동입출금기 한 대와 냉동 진열함, 오른쪽 뒤에 유리 출입문이 보인다. 회색 타일과 높은 창, 주간 조명은 참고 장소와 유사하다. 일행의 몸 전체는 보이지만 오른쪽 찰리는 진열대 끝을 벗어나 출입구 쪽 통로에 노출된다. 출입문을 화면 밖에 두라는 구도와 달리 문 전체가 보이며, 열린 상태도 명확하지 않다.",
        "entities": "요구된 세 인물이 아니라 현우, 앰버와 찰리 두 명으로 총 네 몸이 있다. 현우는 젊은 동아시아계 남성으로 검은 머리, 회색 셔츠, 녹색 바지, 얼굴 멍과 다리 붕대가 보인다. 금발 여자아이 앰버는 남색 상의와 갈색 작업복을 입었지만 착용한 가방은 확인되지 않는다. 두 찰리 모두 흰 기계 얼굴과 베이지 장갑판, 덮개를 갖췄으며 모자는 왼쪽 개체만 썼다. 현우의 투명 가방에는 물품과 지폐가 보이나 지폐의 국가와 액면은 확정할 수 없다. 기계 상단의 ‘ATM’은 읽을 수 있어 문자 금지에 어긋난다.",
        "hard_violations": [
         "찰리의 몸이 두 개로 복제되어 허용되지 않은 네 번째 인물이 등장한다.",
         "현금자동입출금기 상단에 읽을 수 있는 ‘ATM’ 문자가 노출되어 있다."
        ],
        "physics": "현우는 바닥에 디딘 부츠와 굽힌 다리로 몸을 지탱하고, 앰버는 무릎을 바닥에 대고 있다. 두 찰리도 발 또는 접힌 하체가 바닥에 닿아 있다. 현우가 가방 입구를 잡고 있으며 가방 바닥은 타일에 놓여 있다. 덮개는 어깨에 걸쳐 아래로 처진다. 지지 없이 떠 있는 몸이나 물건은 확인되지 않는다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 왼쪽 전경 진열대 끝 너머, 카메라 왼쪽 방향을 긴장된 표정으로 바라본다. 찰리의 얼굴도 대체로 카메라 쪽을 향한다. 열린 출입문은 세 인물의 오른쪽 뒤에 있으므로 출입구 접근 방향을 엿보는 시선 관계가 성립하지 않는다. 겨냥하는 무기나 도구는 없다.",
        "built_space": "왼쪽 전경의 긴 상품 선반 한 조, 그 끝에 모인 세 인물, 오른쪽의 파손된 현금자동입출금기 한 대, 맨 오른쪽 상품 선반, 뒤쪽의 열린 유리 출입문이 보인다. 높은 창과 문틀, 기계의 재질 및 회색 바닥은 이전 장소에 가깝다. 그러나 카메라는 진열대의 은폐면보다 상품 진열면을 크게 보여 주고, 선반이 현우의 몸 일부를 가린다. 찰리의 하체도 다른 인물에게 가려져 세 사람의 전신을 명료하게 보여 주지 못한다. 선반은 뒤쪽 출입구와 일행 사이를 차단하지 않으며, 출입구도 지시와 달리 화면 안에 크게 보인다.",
        "entities": "현우, 앰버, 찰리 세 명만 보인다. 현우의 검은 머리, 앳된 동아시아계 얼굴, 회색 셔츠와 녹색 바지 및 얼굴 멍은 대체로 맞는다. 다리 상처는 확인되지 않는다. 앰버는 금발의 어린 여자아이로 남색 상의와 갈색 작업복, 등에 착용한 가방이 보인다. 찰리는 흰 기계 얼굴과 베이지 장갑판, 모자 및 몸을 덮는 천을 갖췄다. 현우가 꾸러미를 안고 있으나 내부의 현금과 물품은 확인할 수 없다. 기계는 파손되어 있지만 쏟아진 현금은 명확하지 않다. 전경 상품 포장에 읽을 수 있는 한글 인쇄가 남아 있다.",
        "hard_violations": [
         "전경 상품 포장에 판독 가능한 한글 문자가 노출되어 문자 금지 조건을 위반한다."
        ],
        "physics": "현우는 부츠를 바닥에 디딘 채 다리를 굽히고 양손과 팔로 꾸러미를 받친다. 앰버는 무릎과 신발로 바닥을 지지하며 현우의 팔에 손을 댄다. 앰버의 가방은 어깨끈으로 몸에 연결된다. 찰리의 발은 일행 뒤 바닥에 닿고 덮개는 어깨에서 늘어진다. 공중에 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "찰리가 두 명으로 복제된 치명적 오류가 있으며, 일행의 시선도 오른쪽 뒤 출입구가 아닌 왼쪽을 향한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "세 인물과 열린 출입문, 찰리의 변장 및 앰버의 가방은 더 충실하지만, 출입구 반대쪽을 엿보며 진열대가 전신을 가리는 구도는 지시와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 앰버는 화면 왼쪽 바깥을 바라본다. 왼쪽 찰리는 정면 아래쪽, 오른쪽 찰리는 일행 앞쪽으로 얼굴을 향한다. 실제 출입구는 화면 오른쪽 뒤에 있어 누구의 경계 방향도 그 출입구와 맞지 않는다. 무기나 겨냥하는 도구는 없다.",
        "built_space": "중앙에 철망형 상품 진열대 한 조, 왼쪽 전경에 별도 선반, 왼쪽 뒤에 계산대, 오른쪽에 파손된 현금자동입출금기 한 대와 냉동 진열함, 오른쪽 뒤에 유리 출입문이 보인다. 회색 타일과 높은 창, 주간 조명은 참고 장소와 유사하다. 일행의 몸 전체는 보이지만 오른쪽 찰리는 진열대 끝을 벗어나 출입구 쪽 통로에 노출된다. 출입문을 화면 밖에 두라는 구도와 달리 문 전체가 보이며, 열린 상태도 명확하지 않다.",
        "entities": "요구된 세 인물이 아니라 현우, 앰버와 찰리 두 명으로 총 네 몸이 있다. 현우는 젊은 동아시아계 남성으로 검은 머리, 회색 셔츠, 녹색 바지, 얼굴 멍과 다리 붕대가 보인다. 금발 여자아이 앰버는 남색 상의와 갈색 작업복을 입었지만 착용한 가방은 확인되지 않는다. 두 찰리 모두 흰 기계 얼굴과 베이지 장갑판, 덮개를 갖췄으며 모자는 왼쪽 개체만 썼다. 현우의 투명 가방에는 물품과 지폐가 보이나 지폐의 국가와 액면은 확정할 수 없다. 기계 상단의 ‘ATM’은 읽을 수 있어 문자 금지에 어긋난다.",
        "hard_violations": [
         "찰리의 몸이 두 개로 복제되어 허용되지 않은 네 번째 인물이 등장한다.",
         "현금자동입출금기 상단에 읽을 수 있는 ‘ATM’ 문자가 노출되어 있다."
        ],
        "physics": "현우는 바닥에 디딘 부츠와 굽힌 다리로 몸을 지탱하고, 앰버는 무릎을 바닥에 대고 있다. 두 찰리도 발 또는 접힌 하체가 바닥에 닿아 있다. 현우가 가방 입구를 잡고 있으며 가방 바닥은 타일에 놓여 있다. 덮개는 어깨에 걸쳐 아래로 처진다. 지지 없이 떠 있는 몸이나 물건은 확인되지 않는다."
       },
       {
        "label": "A",
        "direction": "현우와 앰버는 왼쪽 전경 진열대 끝 너머, 카메라 왼쪽 방향을 긴장된 표정으로 바라본다. 찰리의 얼굴도 대체로 카메라 쪽을 향한다. 열린 출입문은 세 인물의 오른쪽 뒤에 있으므로 출입구 접근 방향을 엿보는 시선 관계가 성립하지 않는다. 겨냥하는 무기나 도구는 없다.",
        "built_space": "왼쪽 전경의 긴 상품 선반 한 조, 그 끝에 모인 세 인물, 오른쪽의 파손된 현금자동입출금기 한 대, 맨 오른쪽 상품 선반, 뒤쪽의 열린 유리 출입문이 보인다. 높은 창과 문틀, 기계의 재질 및 회색 바닥은 이전 장소에 가깝다. 그러나 카메라는 진열대의 은폐면보다 상품 진열면을 크게 보여 주고, 선반이 현우의 몸 일부를 가린다. 찰리의 하체도 다른 인물에게 가려져 세 사람의 전신을 명료하게 보여 주지 못한다. 선반은 뒤쪽 출입구와 일행 사이를 차단하지 않으며, 출입구도 지시와 달리 화면 안에 크게 보인다.",
        "entities": "현우, 앰버, 찰리 세 명만 보인다. 현우의 검은 머리, 앳된 동아시아계 얼굴, 회색 셔츠와 녹색 바지 및 얼굴 멍은 대체로 맞는다. 다리 상처는 확인되지 않는다. 앰버는 금발의 어린 여자아이로 남색 상의와 갈색 작업복, 등에 착용한 가방이 보인다. 찰리는 흰 기계 얼굴과 베이지 장갑판, 모자 및 몸을 덮는 천을 갖췄다. 현우가 꾸러미를 안고 있으나 내부의 현금과 물품은 확인할 수 없다. 기계는 파손되어 있지만 쏟아진 현금은 명확하지 않다. 전경 상품 포장에 읽을 수 있는 한글 인쇄가 남아 있다.",
        "hard_violations": [
         "전경 상품 포장에 판독 가능한 한글 문자가 노출되어 문자 금지 조건을 위반한다."
        ],
        "physics": "현우는 부츠를 바닥에 디딘 채 다리를 굽히고 양손과 팔로 꾸러미를 받친다. 앰버는 무릎과 신발로 바닥을 지지하며 현우의 팔에 손을 댄다. 앰버의 가방은 어깨끈으로 몸에 연결된다. 찰리의 발은 일행 뒤 바닥에 닿고 덮개는 어깨에서 늘어진다. 공중에 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 중복된 캐릭터 (찰리가 두 명 등장함)",
     "[gpt-high] 찰리의 몸이 두 개로 복제되어 허용되지 않은 네 번째 인물이 등장한다.",
     "[gpt-high] 현금자동입출금기 상단에 읽을 수 있는 ‘ATM’ 문자가 노출되어 있다."
    ],
    "A": [
     "[gpt-high] 전경 상품 포장에 판독 가능한 한글 문자가 노출되어 문자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트가 요구한 인물 구성, 진열대 뒤에 웅크린 자세, 시선 방향을 명확하게 구현했습니다.  ★위반: [gpt-high] 전경 상품 포장에 판독 가능한 한글 문자가 노출되어 문자 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "특정 캐릭터(찰리)가 양옆에 두 명으로 중복 생성되는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 중복된 캐릭터 (찰리가 두 명 등장함) / [gpt-high] 찰리의 몸이 두 개로 복제되어 허용되지 않은 네 번째 인물이 등장한다. / [gpt-high] 현금자동입출금기 상단에 읽을 수 있는 ‘ATM’ 문자가 노출되어 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14_sel.png",
    "asset_id": "2519c1df-a835-4bc8-9341-9eed3309116d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-7c5b-79e8-b5bf-e2b08f580845",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S41sh14"
  }
 },
 "S41sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:53:25.045878+00:00",
  "fingerprint": "7f735466fc99dddbb050b883be40f9658b7d9a34e3758143baa97662f91784b1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S41sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S41sh22_sel.png",
  "source_sha256": "78cd6f88893b24019e520592ffc2b764b2b5e79ed630bad200bb80d757a2bdcc",
  "file": "S41sh22_cine.png",
  "staged_sha256": "e77cd7e848136b0bef2bb9da7292b6c2c3b46a921540ab904df57c9eebd0ebe2",
  "latency_ms": 10909
 },
 "S41sh26::signage": {
  "fp": "27d5fb0abb4823ed",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S41sh26": {
  "input_fingerprint": "31bd110278d5d159",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 꽉 껴안은 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): In the merchandise aisle inside the unattended shop, near the shelving used as cover, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 점포 진열대 (Remains behind the pair after 앰버 has left concealment) — An oblique end section remains peripheral behind 앰버; used as A subdued continuity marker connecting the embrace to the preceding hiding place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the established store illumination unchanged, rendering both smiling faces with gentle tonal separation rather than introducing a warmer light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The breached ATM and spilled cash remain unchanged. Food, drinks, the disposable phone, map and COPD medicine remain packed in the scavenged bags. 앰버: She is out of hiding and visibly delighted, still wearing the scavenged bag and retaining the replacement shoes. 라울: He stands inside the store, delighted by the reunion, with his ponytail unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 꽉 껴안은 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): In the merchandise aisle inside the unattended shop, near the shelving used as cover, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 점포 진열대 (Remains behind the pair after 앰버 has left concealment) — An oblique end section remains peripheral behind 앰버; used as A subdued continuity marker connecting the embrace to the preceding hiding place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the established store illumination unchanged, rendering both smiling faces with gentle tonal separation rather than introducing a warmer light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The breached ATM and spilled cash remain unchanged. Food, drinks, the disposable phone, map and COPD medicine remain packed in the scavenged bags. 앰버: She is out of hiding and visibly delighted, still wearing the scavenged bag and retaining the replacement shoes. 라울: He stands inside the store, delighted by the reunion, with his ponytail unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 꽉 껴안은 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): In the merchandise aisle inside the unattended shop, near the shelving used as cover, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 점포 진열대 (Remains behind the pair after 앰버 has left concealment) — An oblique end section remains peripheral behind 앰버; used as A subdued continuity marker connecting the embrace to the preceding hiding place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the established store illumination unchanged, rendering both smiling faces with gentle tonal separation rather than introducing a warmer light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The breached ATM and spilled cash remain unchanged. Food, drinks, the disposable phone, map and COPD medicine remain packed in the scavenged bags. 앰버: She is out of hiding and visibly delighted, still wearing the scavenged bag and retaining the replacement shoes. 라울: He stands inside the store, delighted by the reunion, with his ponytail unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "두 아이는 몸통을 서로 향하고 양팔로 상대의 어깨와 등을 감싼다. 얼굴은 나란히 카메라 왼쪽 바깥을 향하며 서로 눈을 맞추지는 않는다. 둘 다 이를 드러내며 활짝 웃어 재회의 기쁨은 분명하다. 무기나 조작 중인 물건은 없다.",
    "built_space": "왼쪽 전경의 큰 금속 진열대 한 줄에 약 다섯 층의 상품이 보이고, 두 사람 뒤에는 비스듬한 진열대와 벽 쪽 상품 선반이 보인다. 상부 창들이 있고 오른쪽 가장자리에는 기계 외함 일부가 있다. 두 사람은 진열대 옆 통로에 서 있어 공간적 충돌은 없다. 다만 앰버 뒤에 작고 억제된 연결 표지로 남아야 할 진열대가 화면 왼쪽 전경을 크게 점유한다. 회색 금속, 닳은 바닥과 중성적인 낮빛은 이전 장면과 유사하다.",
    "entities": "보이는 사람은 어린 여자아이 앰버와 어린 남자아이 라울 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 반소매와 갈색 작업복은 참고 이미지에 대체로 맞는다. 라울의 갈색 피부, 어린 얼굴, 뒤로 묶은 곱슬머리와 낡은 회색 티셔츠도 부합한다. 앰버의 가방은 포옹과 몸 뒤 가림 때문에 착용 여부를 확정하기 어렵다. 신발, 흩어진 현금과 포장된 물품은 프레임 밖이거나 가려져 평가할 수 없다. 왼쪽 상품에는 읽을 수 있는 한글과 영문 상표가 여러 곳 남아 있다.",
    "hard_violations": [
     "왼쪽 진열 상품의 한글 표기와 영문 상표가 읽을 수 있을 만큼 선명하여, 어떤 읽을 수 있는 글자나 로고도 없어야 한다는 조건을 위반한다."
    ],
    "physics": "두 사람의 몸통은 서 있는 자세로 화면 아래까지 이어지며 공중에 뜬 모습은 아니다. 발의 접지는 프레임 밖이라 직접 확인되지 않는다. 팔은 상대 몸을 감싸고 손은 어깨와 등에 닿아 있어 포옹으로 가능한 자세다. 상품은 선반 위에 놓여 있고 지지 없이 떠 있는 물체는 보이지 않는다."
   },
   {
    "label": "A",
    "direction": "두 아이가 가슴과 옆얼굴을 가까이 붙이고 서로의 등과 어깨를 꽉 감싼다. 앰버는 화면 왼쪽을, 라울도 대체로 카메라 왼쪽을 보며 크게 웃는다. 서로를 바라보는 눈맞춤은 없지만 요청한 웃는 포옹 동작은 명확하다. 방향을 따로 검증할 무기나 사용 중인 도구는 없다.",
    "built_space": "앰버 뒤 왼쪽에 끝 기둥이 비스듬히 보이는 독립 진열대 한 줄이 있고, 더 왼쪽 배경에 벽을 따라 놓인 상품 선반 한 줄이 보인다. 오른쪽 배경에는 열린 유리 출입구 한 곳, 상부 창들, 오른쪽 가장자리에는 현금인출기 일부가 있다. 두 사람은 진열대에서 나와 출입구 앞 통로에 서 있다. 진열대 끝이 앰버 뒤에 놓인 관계와 출입구의 낮빛은 이전 장소의 연속성을 잘 전달한다. 다만 진열대가 여전히 화면 왼쪽의 상당 부분을 차지해 작고 억제된 배경이라는 요구에는 부족하다.",
    "entities": "인물은 앰버와 라울로 보이는 어린아이 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 티셔츠, 갈색 작업복과 허리 도구가 참고에 부합하며, 낡은 가방도 어깨끈과 함께 보인다. 라울은 어린 얼굴과 갈색 피부, 뒤로 묶은 머리, 낡은 회색 티셔츠를 유지한다. 민족적 배경은 외관만으로 확정할 수 없지만 두 인물의 참고 외형과 대체로 일치한다. 신발과 현금, 가방 속 식품·전화·지도·약은 보이지 않아 상태를 판단할 수 없다. 상품 포장과 상자에 읽을 수 있는 영문 상표와 인쇄 문구가 남아 있다.",
    "hard_violations": [
     "진열대의 과자 포장과 상품 상자에 식별 가능한 영문 상표 및 인쇄 문구가 있어, 읽을 수 있는 글자와 로고를 전면 금지한 조건을 위반한다."
    ],
    "physics": "양쪽 몸통이 수직으로 화면 아래까지 이어져 서서 껴안는 자세로 읽히며, 발은 프레임 밖이다. 손과 팔은 상대의 어깨·등에 실제로 접촉하고 있어 밀착한 포옹의 지지가 자연스럽다. 앰버의 가방은 어깨끈으로 몸에 걸려 있고 허리 도구는 벨트에 매달려 있다. 상품은 선반에 지지되며, 지지 없이 떠 있는 몸이나 물체는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "포옹과 환한 웃음은 맞지만, 진열대가 주변 배경이 아니라 큰 전경을 차지하고 상품의 읽을 수 있는 글자와 로고가 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "앰버 뒤의 비스듬한 진열대 끝, 밀착한 포옹, 가방과 낮 조명의 연속성이 더 잘 맞지만, 선명한 상품 글자와 로고 때문에 완전한 적합작은 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 아이는 몸통을 서로 향하고 양팔로 상대의 어깨와 등을 감싼다. 얼굴은 나란히 카메라 왼쪽 바깥을 향하며 서로 눈을 맞추지는 않는다. 둘 다 이를 드러내며 활짝 웃어 재회의 기쁨은 분명하다. 무기나 조작 중인 물건은 없다.",
        "built_space": "왼쪽 전경의 큰 금속 진열대 한 줄에 약 다섯 층의 상품이 보이고, 두 사람 뒤에는 비스듬한 진열대와 벽 쪽 상품 선반이 보인다. 상부 창들이 있고 오른쪽 가장자리에는 기계 외함 일부가 있다. 두 사람은 진열대 옆 통로에 서 있어 공간적 충돌은 없다. 다만 앰버 뒤에 작고 억제된 연결 표지로 남아야 할 진열대가 화면 왼쪽 전경을 크게 점유한다. 회색 금속, 닳은 바닥과 중성적인 낮빛은 이전 장면과 유사하다.",
        "entities": "보이는 사람은 어린 여자아이 앰버와 어린 남자아이 라울 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 반소매와 갈색 작업복은 참고 이미지에 대체로 맞는다. 라울의 갈색 피부, 어린 얼굴, 뒤로 묶은 곱슬머리와 낡은 회색 티셔츠도 부합한다. 앰버의 가방은 포옹과 몸 뒤 가림 때문에 착용 여부를 확정하기 어렵다. 신발, 흩어진 현금과 포장된 물품은 프레임 밖이거나 가려져 평가할 수 없다. 왼쪽 상품에는 읽을 수 있는 한글과 영문 상표가 여러 곳 남아 있다.",
        "hard_violations": [
         "왼쪽 진열 상품의 한글 표기와 영문 상표가 읽을 수 있을 만큼 선명하여, 어떤 읽을 수 있는 글자나 로고도 없어야 한다는 조건을 위반한다."
        ],
        "physics": "두 사람의 몸통은 서 있는 자세로 화면 아래까지 이어지며 공중에 뜬 모습은 아니다. 발의 접지는 프레임 밖이라 직접 확인되지 않는다. 팔은 상대 몸을 감싸고 손은 어깨와 등에 닿아 있어 포옹으로 가능한 자세다. 상품은 선반 위에 놓여 있고 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "두 아이가 가슴과 옆얼굴을 가까이 붙이고 서로의 등과 어깨를 꽉 감싼다. 앰버는 화면 왼쪽을, 라울도 대체로 카메라 왼쪽을 보며 크게 웃는다. 서로를 바라보는 눈맞춤은 없지만 요청한 웃는 포옹 동작은 명확하다. 방향을 따로 검증할 무기나 사용 중인 도구는 없다.",
        "built_space": "앰버 뒤 왼쪽에 끝 기둥이 비스듬히 보이는 독립 진열대 한 줄이 있고, 더 왼쪽 배경에 벽을 따라 놓인 상품 선반 한 줄이 보인다. 오른쪽 배경에는 열린 유리 출입구 한 곳, 상부 창들, 오른쪽 가장자리에는 현금인출기 일부가 있다. 두 사람은 진열대에서 나와 출입구 앞 통로에 서 있다. 진열대 끝이 앰버 뒤에 놓인 관계와 출입구의 낮빛은 이전 장소의 연속성을 잘 전달한다. 다만 진열대가 여전히 화면 왼쪽의 상당 부분을 차지해 작고 억제된 배경이라는 요구에는 부족하다.",
        "entities": "인물은 앰버와 라울로 보이는 어린아이 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 티셔츠, 갈색 작업복과 허리 도구가 참고에 부합하며, 낡은 가방도 어깨끈과 함께 보인다. 라울은 어린 얼굴과 갈색 피부, 뒤로 묶은 머리, 낡은 회색 티셔츠를 유지한다. 민족적 배경은 외관만으로 확정할 수 없지만 두 인물의 참고 외형과 대체로 일치한다. 신발과 현금, 가방 속 식품·전화·지도·약은 보이지 않아 상태를 판단할 수 없다. 상품 포장과 상자에 읽을 수 있는 영문 상표와 인쇄 문구가 남아 있다.",
        "hard_violations": [
         "진열대의 과자 포장과 상품 상자에 식별 가능한 영문 상표 및 인쇄 문구가 있어, 읽을 수 있는 글자와 로고를 전면 금지한 조건을 위반한다."
        ],
        "physics": "양쪽 몸통이 수직으로 화면 아래까지 이어져 서서 껴안는 자세로 읽히며, 발은 프레임 밖이다. 손과 팔은 상대의 어깨·등에 실제로 접촉하고 있어 밀착한 포옹의 지지가 자연스럽다. 앰버의 가방은 어깨끈으로 몸에 걸려 있고 허리 도구는 벨트에 매달려 있다. 상품은 선반에 지지되며, 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "포옹과 환한 웃음은 맞지만, 진열대가 주변 배경이 아니라 큰 전경을 차지하고 상품의 읽을 수 있는 글자와 로고가 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "앰버 뒤의 비스듬한 진열대 끝, 밀착한 포옹, 가방과 낮 조명의 연속성이 더 잘 맞지만, 선명한 상품 글자와 로고 때문에 완전한 적합작은 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "두 아이는 몸통을 서로 향하고 양팔로 상대의 어깨와 등을 감싼다. 얼굴은 나란히 카메라 왼쪽 바깥을 향하며 서로 눈을 맞추지는 않는다. 둘 다 이를 드러내며 활짝 웃어 재회의 기쁨은 분명하다. 무기나 조작 중인 물건은 없다.",
        "built_space": "왼쪽 전경의 큰 금속 진열대 한 줄에 약 다섯 층의 상품이 보이고, 두 사람 뒤에는 비스듬한 진열대와 벽 쪽 상품 선반이 보인다. 상부 창들이 있고 오른쪽 가장자리에는 기계 외함 일부가 있다. 두 사람은 진열대 옆 통로에 서 있어 공간적 충돌은 없다. 다만 앰버 뒤에 작고 억제된 연결 표지로 남아야 할 진열대가 화면 왼쪽 전경을 크게 점유한다. 회색 금속, 닳은 바닥과 중성적인 낮빛은 이전 장면과 유사하다.",
        "entities": "보이는 사람은 어린 여자아이 앰버와 어린 남자아이 라울 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 반소매와 갈색 작업복은 참고 이미지에 대체로 맞는다. 라울의 갈색 피부, 어린 얼굴, 뒤로 묶은 곱슬머리와 낡은 회색 티셔츠도 부합한다. 앰버의 가방은 포옹과 몸 뒤 가림 때문에 착용 여부를 확정하기 어렵다. 신발, 흩어진 현금과 포장된 물품은 프레임 밖이거나 가려져 평가할 수 없다. 왼쪽 상품에는 읽을 수 있는 한글과 영문 상표가 여러 곳 남아 있다.",
        "hard_violations": [
         "왼쪽 진열 상품의 한글 표기와 영문 상표가 읽을 수 있을 만큼 선명하여, 어떤 읽을 수 있는 글자나 로고도 없어야 한다는 조건을 위반한다."
        ],
        "physics": "두 사람의 몸통은 서 있는 자세로 화면 아래까지 이어지며 공중에 뜬 모습은 아니다. 발의 접지는 프레임 밖이라 직접 확인되지 않는다. 팔은 상대 몸을 감싸고 손은 어깨와 등에 닿아 있어 포옹으로 가능한 자세다. 상품은 선반 위에 놓여 있고 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "두 아이가 가슴과 옆얼굴을 가까이 붙이고 서로의 등과 어깨를 꽉 감싼다. 앰버는 화면 왼쪽을, 라울도 대체로 카메라 왼쪽을 보며 크게 웃는다. 서로를 바라보는 눈맞춤은 없지만 요청한 웃는 포옹 동작은 명확하다. 방향을 따로 검증할 무기나 사용 중인 도구는 없다.",
        "built_space": "앰버 뒤 왼쪽에 끝 기둥이 비스듬히 보이는 독립 진열대 한 줄이 있고, 더 왼쪽 배경에 벽을 따라 놓인 상품 선반 한 줄이 보인다. 오른쪽 배경에는 열린 유리 출입구 한 곳, 상부 창들, 오른쪽 가장자리에는 현금인출기 일부가 있다. 두 사람은 진열대에서 나와 출입구 앞 통로에 서 있다. 진열대 끝이 앰버 뒤에 놓인 관계와 출입구의 낮빛은 이전 장소의 연속성을 잘 전달한다. 다만 진열대가 여전히 화면 왼쪽의 상당 부분을 차지해 작고 억제된 배경이라는 요구에는 부족하다.",
        "entities": "인물은 앰버와 라울로 보이는 어린아이 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 티셔츠, 갈색 작업복과 허리 도구가 참고에 부합하며, 낡은 가방도 어깨끈과 함께 보인다. 라울은 어린 얼굴과 갈색 피부, 뒤로 묶은 머리, 낡은 회색 티셔츠를 유지한다. 민족적 배경은 외관만으로 확정할 수 없지만 두 인물의 참고 외형과 대체로 일치한다. 신발과 현금, 가방 속 식품·전화·지도·약은 보이지 않아 상태를 판단할 수 없다. 상품 포장과 상자에 읽을 수 있는 영문 상표와 인쇄 문구가 남아 있다.",
        "hard_violations": [
         "진열대의 과자 포장과 상품 상자에 식별 가능한 영문 상표 및 인쇄 문구가 있어, 읽을 수 있는 글자와 로고를 전면 금지한 조건을 위반한다."
        ],
        "physics": "양쪽 몸통이 수직으로 화면 아래까지 이어져 서서 껴안는 자세로 읽히며, 발은 프레임 밖이다. 손과 팔은 상대의 어깨·등에 실제로 접촉하고 있어 밀착한 포옹의 지지가 자연스럽다. 앰버의 가방은 어깨끈으로 몸에 걸려 있고 허리 도구는 벨트에 매달려 있다. 상품은 선반에 지지되며, 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 4,
   "A": 5
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 4,
    "verdict_ko": "포옹과 환한 웃음은 맞지만, 진열대가 주변 배경이 아니라 큰 전경을 차지하고 상품의 읽을 수 있는 글자와 로고가 명시적 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 5,
    "verdict_ko": "앰버 뒤의 비스듬한 진열대 끝, 밀착한 포옹, 가방과 낮 조명의 연속성이 더 잘 맞지만, 선명한 상품 글자와 로고 때문에 완전한 적합작은 아니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh22_sel.png",
    "asset_id": "3110d238-841f-4543-91bc-d482d9e1756d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-7e11-7afc-a7c7-44e5f949c36d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S41sh22"
  }
 },
 "S41sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:54:40.467019+00:00",
  "fingerprint": "80268963895318ee23a3f904a7aee672cddff9d8f521b948ab55e30f6a427c20",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S41sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S41sh26_sel.png",
  "source_sha256": "3d0344364acc5a6b86084045df306cbbe8ec7042a30acea9f278fd49abad0f99",
  "file": "S41sh26_cine.png",
  "staged_sha256": "eaec0c39b324644d02345b86e9f4dac5f21e16f2b53e0703bacc140d3a83220c",
  "latency_ms": 9204
 },
 "S42sh2::signage": {
  "fp": "bdbf3f72217f9441",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S42sh2": {
  "input_fingerprint": "869639f8bbefdadb",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Gu Dohwan's dead body is enclosed in a body bag with the zipper opened to expose his face for inspection from above. His torso and limbs remain concealed, and the source does not specify his head's turn or the surface supporting the bag.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The body bag's zipper is open at the face end. It is night at the refugee-settlement entrance. 구도환: He is dead and is enclosed in a body bag with the zipper opened to expose his face for inspection from above.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Gu Dohwan's dead body is enclosed in a body bag with the zipper opened to expose his face for inspection from above. His torso and limbs remain concealed, and the source does not specify his head's turn or the surface supporting the bag.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The body bag's zipper is open at the face end. It is night at the refugee-settlement entrance. 구도환: He is dead and is enclosed in a body bag with the zipper opened to expose his face for inspection from above.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Gu Dohwan's dead body is enclosed in a body bag with the zipper opened to expose his face for inspection from above. His torso and limbs remain concealed, and the source does not specify his head's turn or the surface supporting the bag.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The body bag's zipper is open at the face end. It is night at the refugee-settlement entrance. 구도환: He is dead and is enclosed in a body bag with the zipper opened to expose his face for inspection from above.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2__bgfirst_bg.png",
     "asset_id": "0428d63b-de1f-42c5-ad4e-e5abb4f95632",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S42sh2.png",
     "asset_id": "ed925a67-f2a3-41f5-9413-65abf6516887",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1357020>",
     "asset_id": "b269127c-b439-46b5-8f0c-91ba4c73ecea",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_gate_b39179.png",
     "asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1357020>",
     "asset_id": "b269127c-b439-46b5-8f0c-91ba4c73ecea",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "두 눈을 감고 있음. 카메라는 지평선을 향해 수평으로 놓임.",
    "built_space": "야간 수용소 입구. 원경의 게이트가 정면 수평 구도로 배치됨.",
    "entities": "구도환(얼굴, 50대 남성 일치), 보디백(열린 채 바닥에 놓임).",
    "hard_violations": [
     "[gemini-pro] 내려다보는 시점(seen from above) 지시 위반",
     "[gemini-pro] 바닥의 얼굴을 내려다보면서 원경 지평선이 정면으로 보이는 물리적으로 불가능한 스테이징"
    ],
    "physics": "얼굴이 바닥에 위치하나, 수평 시점 탓에 거인처럼 보이거나 배경과 원근감이 어긋남."
   },
   {
    "label": "B",
    "direction": "두 눈을 감고 있음. 카메라는 정면 수평을 향함.",
    "built_space": "야간 수용소 입구. 배경이 정면 수평 구도로 렌더링됨.",
    "entities": "구도환(얼굴 일치), 보디백(더플백처럼 세워진 형태).",
    "hard_violations": [
     "[gemini-pro] 내려다보는 시점(seen from above) 지시 위반",
     "[gemini-pro] 사망한 시신과 가방이 중력을 거스르고 수직으로 서 있는 물리적 위반"
    ],
    "physics": "시신과 보디백이 지지대 없이 수직으로 기립해 있어 중력 및 부동 상태 지시와 모순됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "명시된 '내려다보는 시점'을 완전히 무시하고 수평 앵글을 적용하여, 바닥에 누운 시신과 정면 원경이 동시에 보이는 물리적으로 불가능한 구도를 만들었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "시점 지시 위반은 물론, 바닥에 늘어져야 할 사망한 시신과 보디백이 중력을 무시하고 수직으로 서 있는 치명적인 물리 법칙 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 감고 있음. 카메라는 지평선을 향해 수평으로 놓임.",
        "built_space": "야간 수용소 입구. 원경의 게이트가 정면 수평 구도로 배치됨.",
        "entities": "구도환(얼굴, 50대 남성 일치), 보디백(열린 채 바닥에 놓임).",
        "hard_violations": [
         "내려다보는 시점(seen from above) 지시 위반",
         "바닥의 얼굴을 내려다보면서 원경 지평선이 정면으로 보이는 물리적으로 불가능한 스테이징"
        ],
        "physics": "얼굴이 바닥에 위치하나, 수평 시점 탓에 거인처럼 보이거나 배경과 원근감이 어긋남."
       },
       {
        "label": "B",
        "direction": "두 눈을 감고 있음. 카메라는 정면 수평을 향함.",
        "built_space": "야간 수용소 입구. 배경이 정면 수평 구도로 렌더링됨.",
        "entities": "구도환(얼굴 일치), 보디백(더플백처럼 세워진 형태).",
        "hard_violations": [
         "내려다보는 시점(seen from above) 지시 위반",
         "사망한 시신과 가방이 중력을 거스르고 수직으로 서 있는 물리적 위반"
        ],
        "physics": "시신과 보디백이 지지대 없이 수직으로 기립해 있어 중력 및 부동 상태 지시와 모순됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "명시된 '내려다보는 시점'을 완전히 무시하고 수평 앵글을 적용하여, 바닥에 누운 시신과 정면 원경이 동시에 보이는 물리적으로 불가능한 구도를 만들었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "시점 지시 위반은 물론, 바닥에 늘어져야 할 사망한 시신과 보디백이 중력을 무시하고 수직으로 서 있는 치명적인 물리 법칙 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 감고 있음. 카메라는 지평선을 향해 수평으로 놓임.",
        "built_space": "야간 수용소 입구. 원경의 게이트가 정면 수평 구도로 배치됨.",
        "entities": "구도환(얼굴, 50대 남성 일치), 보디백(열린 채 바닥에 놓임).",
        "hard_violations": [
         "내려다보는 시점(seen from above) 지시 위반",
         "바닥의 얼굴을 내려다보면서 원경 지평선이 정면으로 보이는 물리적으로 불가능한 스테이징"
        ],
        "physics": "얼굴이 바닥에 위치하나, 수평 시점 탓에 거인처럼 보이거나 배경과 원근감이 어긋남."
       },
       {
        "label": "B",
        "direction": "두 눈을 감고 있음. 카메라는 정면 수평을 향함.",
        "built_space": "야간 수용소 입구. 배경이 정면 수평 구도로 렌더링됨.",
        "entities": "구도환(얼굴 일치), 보디백(더플백처럼 세워진 형태).",
        "hard_violations": [
         "내려다보는 시점(seen from above) 지시 위반",
         "사망한 시신과 가방이 중력을 거스르고 수직으로 서 있는 물리적 위반"
        ],
        "physics": "시신과 보디백이 지지대 없이 수직으로 기립해 있어 중력 및 부동 상태 지시와 모순됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "야간 장소와 인물은 대체로 맞지만, 얼굴이 작고 가슴까지 드러나며 위에서 보디백 안을 내려다보는 얼굴 클로즈업이 아니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "얼굴 중심의 클로즈업과 몸통 은폐는 더 충실하지만, 얼굴 위에서 내려다보지 않고 발치 쪽 낮은 위치에서 바라보는 시점이 결정적으로 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자는 눈을 감고 얼굴을 위로 향하고 있어 응시 대상이 없다. 카메라는 얼굴을 발치 쪽에서 비스듬히 바라보며, 얼굴 너머 출입문과 도로가 넓게 보인다. 얼굴 바로 위에서 보디백 내부를 내려다보는 시점으로 읽히지는 않는다.",
        "built_space": "중앙 출입문 한 곳과 이를 둘러싼 기둥 두 개, 양옆 경비 초소, 왼쪽 콘크리트 벽, 오른쪽 철망과 천막, 젖은 도로가 보인다. 장소의 주요 배치는 참고와 대체로 일치한다. 다만 배경 건축과 도로에 상당한 화면을 할애해 얼굴 클로즈업의 집중도가 낮다.",
        "entities": "보이는 사람은 검은 머리의 중년 동아시아계 남성 한 명뿐이며, 얼굴 윤곽과 머리 모양은 구도환 참고와 대체로 맞는다. 감긴 눈과 혈색 없는 피부가 죽은 상태를 표현한다. 검은 목도리, 갈색 상의, 털 안감 외투가 보이며 참고 의상과 부합하지만, 보디백이 가슴까지 열려 몸통을 숨기라는 조건을 벗어난다. 검은 보디백 한 개와 열린 지퍼가 있고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리 뒤쪽은 보디백 안쪽 천에 닿아 있고 목은 두꺼운 목도리에 둘러싸여 있어 명백히 공중에 떠 있는 신체 부위는 보이지 않는다. 몸은 도로 위 보디백 안에 놓인 것으로 읽힌다. 머리 위 보디백 가장자리가 높은 뾰족한 형태로 서 있어 다소 부자연스럽지만, 접힌 두꺼운 원단의 강성으로 유지될 가능성이 있어 무지지 부유로 단정할 수는 없다."
       },
       {
        "label": "B",
        "direction": "남자는 눈을 감고 얼굴을 위로 향한다. 카메라는 턱과 콧구멍이 강조되는 발치 쪽 낮은 방향에서 얼굴을 바라본다. 얼굴을 내려다보는 관찰자의 시점이 아니라 얼굴 너머 출입구를 함께 보는 낮은 시점이다.",
        "built_space": "중앙 출입문 한 곳, 문틀 기둥 두 개, 왼쪽 초소, 왼쪽 콘크리트 벽, 오른쪽 철망과 천막, 확성기 기둥 두 개가 보인다. 오른쪽 초소 부근은 인물에 가려져 있다. 젖은 도로와 야간 시설 조명은 참고 장소와 부합한다. 얼굴은 크게 잡혔으나, 출입구와 하늘까지 보이는 배경은 요구된 하향 시점과 맞지 않는다.",
        "entities": "검은 머리의 중년 동아시아계 남성 한 명이며, 참고의 얼굴과 머리 모양에 대체로 부합한다. 창백한 피부, 감긴 눈, 이완된 입이 죽은 상태를 표현한다. 얼굴과 목 주변의 검은 니트 및 약간의 외투 안감만 보이고 몸통과 팔다리는 보디백에 가려져 있다. 검은 보디백 한 개의 지퍼가 얼굴 주변에서 열려 있으며, 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리 뒤로 보디백 안감이 이어지고 목 아래에는 두꺼운 니트가 닿아 있어 머리가 받쳐진 상태로 읽힌다. 보디백은 도로에 놓여 있고 양쪽 가장자리와 앞쪽 덮개는 접혀 늘어진다. 들어 올린 팔다리나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "야간 장소와 인물은 대체로 맞지만, 얼굴이 작고 가슴까지 드러나며 위에서 보디백 안을 내려다보는 얼굴 클로즈업이 아니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "얼굴 중심의 클로즈업과 몸통 은폐는 더 충실하지만, 얼굴 위에서 내려다보지 않고 발치 쪽 낮은 위치에서 바라보는 시점이 결정적으로 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "남자는 눈을 감고 얼굴을 위로 향하고 있어 응시 대상이 없다. 카메라는 얼굴을 발치 쪽에서 비스듬히 바라보며, 얼굴 너머 출입문과 도로가 넓게 보인다. 얼굴 바로 위에서 보디백 내부를 내려다보는 시점으로 읽히지는 않는다.",
        "built_space": "중앙 출입문 한 곳과 이를 둘러싼 기둥 두 개, 양옆 경비 초소, 왼쪽 콘크리트 벽, 오른쪽 철망과 천막, 젖은 도로가 보인다. 장소의 주요 배치는 참고와 대체로 일치한다. 다만 배경 건축과 도로에 상당한 화면을 할애해 얼굴 클로즈업의 집중도가 낮다.",
        "entities": "보이는 사람은 검은 머리의 중년 동아시아계 남성 한 명뿐이며, 얼굴 윤곽과 머리 모양은 구도환 참고와 대체로 맞는다. 감긴 눈과 혈색 없는 피부가 죽은 상태를 표현한다. 검은 목도리, 갈색 상의, 털 안감 외투가 보이며 참고 의상과 부합하지만, 보디백이 가슴까지 열려 몸통을 숨기라는 조건을 벗어난다. 검은 보디백 한 개와 열린 지퍼가 있고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리 뒤쪽은 보디백 안쪽 천에 닿아 있고 목은 두꺼운 목도리에 둘러싸여 있어 명백히 공중에 떠 있는 신체 부위는 보이지 않는다. 몸은 도로 위 보디백 안에 놓인 것으로 읽힌다. 머리 위 보디백 가장자리가 높은 뾰족한 형태로 서 있어 다소 부자연스럽지만, 접힌 두꺼운 원단의 강성으로 유지될 가능성이 있어 무지지 부유로 단정할 수는 없다."
       },
       {
        "label": "A",
        "direction": "남자는 눈을 감고 얼굴을 위로 향한다. 카메라는 턱과 콧구멍이 강조되는 발치 쪽 낮은 방향에서 얼굴을 바라본다. 얼굴을 내려다보는 관찰자의 시점이 아니라 얼굴 너머 출입구를 함께 보는 낮은 시점이다.",
        "built_space": "중앙 출입문 한 곳, 문틀 기둥 두 개, 왼쪽 초소, 왼쪽 콘크리트 벽, 오른쪽 철망과 천막, 확성기 기둥 두 개가 보인다. 오른쪽 초소 부근은 인물에 가려져 있다. 젖은 도로와 야간 시설 조명은 참고 장소와 부합한다. 얼굴은 크게 잡혔으나, 출입구와 하늘까지 보이는 배경은 요구된 하향 시점과 맞지 않는다.",
        "entities": "검은 머리의 중년 동아시아계 남성 한 명이며, 참고의 얼굴과 머리 모양에 대체로 부합한다. 창백한 피부, 감긴 눈, 이완된 입이 죽은 상태를 표현한다. 얼굴과 목 주변의 검은 니트 및 약간의 외투 안감만 보이고 몸통과 팔다리는 보디백에 가려져 있다. 검은 보디백 한 개의 지퍼가 얼굴 주변에서 열려 있으며, 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리 뒤로 보디백 안감이 이어지고 목 아래에는 두꺼운 니트가 닿아 있어 머리가 받쳐진 상태로 읽힌다. 보디백은 도로에 놓여 있고 양쪽 가장자리와 앞쪽 덮개는 접혀 늘어진다. 들어 올린 팔다리나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.417
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "A": [
     "[gemini-pro] 내려다보는 시점(seen from above) 지시 위반",
     "[gemini-pro] 바닥의 얼굴을 내려다보면서 원경 지평선이 정면으로 보이는 물리적으로 불가능한 스테이징"
    ],
    "B": [
     "[gemini-pro] 내려다보는 시점(seen from above) 지시 위반",
     "[gemini-pro] 사망한 시신과 가방이 중력을 거스르고 수직으로 서 있는 물리적 위반"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "명시된 '내려다보는 시점'을 완전히 무시하고 수평 앵글을 적용하여, 바닥에 누운 시신과 정면 원경이 동시에 보이는 물리적으로 불가능한 구도를 만들었습니다.  ★위반: [gemini-pro] 내려다보는 시점(seen from above) 지시 위반 / [gemini-pro] 바닥의 얼굴을 내려다보면서 원경 지평선이 정면으로 보이는 물리적으로 불가능한 스테이징"
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "시점 지시 위반은 물론, 바닥에 늘어져야 할 사망한 시신과 보디백이 중력을 무시하고 수직으로 서 있는 치명적인 물리 법칙 오류를 범했습니다.  ★위반: [gemini-pro] 내려다보는 시점(seen from above) 지시 위반 / [gemini-pro] 사망한 시신과 가방이 중력을 거스르고 수직으로 서 있는 물리적 위반"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_gate_b39179.png",
    "asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1357020>",
    "asset_id": "b269127c-b439-46b5-8f0c-91ba4c73ecea",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-7fd4-739d-91cd-8ebfa8e3d54b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2__bgfirst_bg.png",
   "bg_asset_id": "0428d63b-de1f-42c5-ad4e-e5abb4f95632",
   "bg_record_key": "S42sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "camp_gate",
   "groupbg_asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S42sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:55:52.063356+00:00",
  "fingerprint": "2b67cb457351c97e393a05b16fd080bde0768ec44e7d8ea95caa0d8ef52036b4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S42sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S42sh2_sel.png",
  "source_sha256": "6b626b197ccb3abd1d882f6fdae604d1c19074c2b35dda71119319b873dc23f6",
  "file": "S42sh2_cine.png",
  "staged_sha256": "e41a5f4c9643c0929ef562286d4948ed05d9e83c03a266d661279996cb174c53",
  "latency_ms": 9544
 },
 "S42sh11::signage": {
  "fp": "5bb1fa41ba640e00",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S42sh11": {
  "input_fingerprint": "d2aee962e2cd6950",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문을 한 손으로 꽉 틀어쥔 채 단호하게 막아선 박철진의 팽팽한 자세.\n\nLOCATION (lock): Outside the open door of a sedan stopped on the muddy approach to the refugee settlement entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 고급 세단의 열린 차 문 (Open and held by 박철진, preventing departure) — The edge and an oblique portion of the inner face are visible at the left side; used as Connects the gripping hand to the blocked doorway while occupying less than two-fifths of the frame; 세단 출입구 (Open, with 박철진 blocking its use) — Viewed diagonally from outside; the occupant remains beyond the crop; used as Provides the spatial evidence for his obstructing posture and inward-directed gaze.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient illumination, preserving readable tension in the gripping hand and face without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan is stopped with its door open. The body bag remains open at the face end nearby. 박철진: He stands at the sedan with one hand holding the open door.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문을 한 손으로 꽉 틀어쥔 채 단호하게 막아선 박철진의 팽팽한 자세.\n\nLOCATION (lock): Outside the open door of a sedan stopped on the muddy approach to the refugee settlement entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 고급 세단의 열린 차 문 (Open and held by 박철진, preventing departure) — The edge and an oblique portion of the inner face are visible at the left side; used as Connects the gripping hand to the blocked doorway while occupying less than two-fifths of the frame; 세단 출입구 (Open, with 박철진 blocking its use) — Viewed diagonally from outside; the occupant remains beyond the crop; used as Provides the spatial evidence for his obstructing posture and inward-directed gaze.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient illumination, preserving readable tension in the gripping hand and face without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan is stopped with its door open. The body bag remains open at the face end nearby. 박철진: He stands at the sedan with one hand holding the open door.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문을 한 손으로 꽉 틀어쥔 채 단호하게 막아선 박철진의 팽팽한 자세.\n\nLOCATION (lock): Outside the open door of a sedan stopped on the muddy approach to the refugee settlement entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 고급 세단의 열린 차 문 (Open and held by 박철진, preventing departure) — The edge and an oblique portion of the inner face are visible at the left side; used as Connects the gripping hand to the blocked doorway while occupying less than two-fifths of the frame; 세단 출입구 (Open, with 박철진 blocking its use) — Viewed diagonally from outside; the occupant remains beyond the crop; used as Provides the spatial evidence for his obstructing posture and inward-directed gaze.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient illumination, preserving readable tension in the gripping hand and face without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan is stopped with its door open. The body bag remains open at the face end nearby. 박철진: He stands at the sedan with one hand holding the open door.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "박철진의 시선이 카메라 렌즈 쪽을 향하고 있어, 지시된 내부를 향한 시선(inward-directed gaze)과 어긋납니다.",
    "built_space": "차량이 우측에 있고 열린 문이 중앙을 가로막아, 프롬프트가 지시한 '화면 왼쪽 가장자리의 차 문' 프레이밍을 위반했습니다. 배경의 수용소 벽면과 진흙탕은 이전 샷과 일치합니다.",
    "entities": "박철진이 등장하나 레퍼런스의 붉은 완장이 누락되었습니다. 또한 차량 유리창 너머로 프롬프트에 명시되지 않은 추가 인물의 얼굴이 보입니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 추가 인물(창문 안의 얼굴) 등장",
     "[gpt-high] 닫힌 측면 창에 박철진의 실제 머리 기울기와 창에 대한 몸 방향으로 성립하지 않는 얼굴 반사를 생성했다."
    ],
    "physics": "오른손으로 차 문을 잡고 땅에 서 있는 자세는 지탱면과 접촉점이 자연스럽습니다."
   },
   {
    "label": "B",
    "direction": "박철진이 차량 출입구를 막아선 채 화면 우측 전경의 대상을 향해 단호한 시선을 보내고 있습니다.",
    "built_space": "차량이 왼쪽에 위치하고 열린 문이 화면 좌측 일부를 차지하며, 배경의 수용소 입구 형태와 위치가 이전 샷과 정확히 일치하게 배치되었습니다.",
    "entities": "박철진의 붉은 완장과 의상이 정확히 묘사되었고 배경에 바디백도 명시된 대로 열린 채 놓여 있으나, 화면 우측 전경에 프롬프트에 존재하지 않는 거대한 인물의 뒷모습이 포함되었습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 추가 인물(화면 우측 전경의 뒷모습) 등장",
     "[gpt-high] 박철진만 보여야 하는 숏에 오른쪽 전경의 별도 인물 몸통을 추가했다.",
     "[gpt-high] 이전 숏의 인물을 이번 숏에 등장시키지 말라는 지시와 달리 시신 가방 속 얼굴을 다시 노출했다."
    ],
    "physics": "왼손으로 열린 차 문 모서리를 단단히 잡고 진흙 바닥에 안정적으로 체중을 싣고 서 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "차 문을 오른쪽에 배치해 샷 프레이밍 지시를 완전히 어겼으며, 필수 의상인 붉은 완장 누락 및 창문에 정체불명의 인물이 나타나 실격입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지시된 좌측 차 문 프레이밍, 배경, 의상(완장) 및 바디백 배치를 훌륭하게 구현했으나, 우측 전경에 프롬프트에 없는 인물이 크게 등장하여 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선이 카메라 렌즈 쪽을 향하고 있어, 지시된 내부를 향한 시선(inward-directed gaze)과 어긋납니다.",
        "built_space": "차량이 우측에 있고 열린 문이 중앙을 가로막아, 프롬프트가 지시한 '화면 왼쪽 가장자리의 차 문' 프레이밍을 위반했습니다. 배경의 수용소 벽면과 진흙탕은 이전 샷과 일치합니다.",
        "entities": "박철진이 등장하나 레퍼런스의 붉은 완장이 누락되었습니다. 또한 차량 유리창 너머로 프롬프트에 명시되지 않은 추가 인물의 얼굴이 보입니다.",
        "hard_violations": [
         "프롬프트에 없는 추가 인물(창문 안의 얼굴) 등장"
        ],
        "physics": "오른손으로 차 문을 잡고 땅에 서 있는 자세는 지탱면과 접촉점이 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "박철진이 차량 출입구를 막아선 채 화면 우측 전경의 대상을 향해 단호한 시선을 보내고 있습니다.",
        "built_space": "차량이 왼쪽에 위치하고 열린 문이 화면 좌측 일부를 차지하며, 배경의 수용소 입구 형태와 위치가 이전 샷과 정확히 일치하게 배치되었습니다.",
        "entities": "박철진의 붉은 완장과 의상이 정확히 묘사되었고 배경에 바디백도 명시된 대로 열린 채 놓여 있으나, 화면 우측 전경에 프롬프트에 존재하지 않는 거대한 인물의 뒷모습이 포함되었습니다.",
        "hard_violations": [
         "프롬프트에 없는 추가 인물(화면 우측 전경의 뒷모습) 등장"
        ],
        "physics": "왼손으로 열린 차 문 모서리를 단단히 잡고 진흙 바닥에 안정적으로 체중을 싣고 서 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "차 문을 오른쪽에 배치해 샷 프레이밍 지시를 완전히 어겼으며, 필수 의상인 붉은 완장 누락 및 창문에 정체불명의 인물이 나타나 실격입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지시된 좌측 차 문 프레이밍, 배경, 의상(완장) 및 바디백 배치를 훌륭하게 구현했으나, 우측 전경에 프롬프트에 없는 인물이 크게 등장하여 실격입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선이 카메라 렌즈 쪽을 향하고 있어, 지시된 내부를 향한 시선(inward-directed gaze)과 어긋납니다.",
        "built_space": "차량이 우측에 있고 열린 문이 중앙을 가로막아, 프롬프트가 지시한 '화면 왼쪽 가장자리의 차 문' 프레이밍을 위반했습니다. 배경의 수용소 벽면과 진흙탕은 이전 샷과 일치합니다.",
        "entities": "박철진이 등장하나 레퍼런스의 붉은 완장이 누락되었습니다. 또한 차량 유리창 너머로 프롬프트에 명시되지 않은 추가 인물의 얼굴이 보입니다.",
        "hard_violations": [
         "프롬프트에 없는 추가 인물(창문 안의 얼굴) 등장"
        ],
        "physics": "오른손으로 차 문을 잡고 땅에 서 있는 자세는 지탱면과 접촉점이 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "박철진이 차량 출입구를 막아선 채 화면 우측 전경의 대상을 향해 단호한 시선을 보내고 있습니다.",
        "built_space": "차량이 왼쪽에 위치하고 열린 문이 화면 좌측 일부를 차지하며, 배경의 수용소 입구 형태와 위치가 이전 샷과 정확히 일치하게 배치되었습니다.",
        "entities": "박철진의 붉은 완장과 의상이 정확히 묘사되었고 배경에 바디백도 명시된 대로 열린 채 놓여 있으나, 화면 우측 전경에 프롬프트에 존재하지 않는 거대한 인물의 뒷모습이 포함되었습니다.",
        "hard_violations": [
         "프롬프트에 없는 추가 인물(화면 우측 전경의 뒷모습) 등장"
        ],
        "physics": "왼손으로 열린 차 문 모서리를 단단히 잡고 진흙 바닥에 안정적으로 체중을 싣고 서 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "금지된 추가 인물과 이전 숏의 시신이 등장하고, 박철진이 차 안이 아닌 바깥 인물을 바라보며 미디엄 숏보다 넓게 잡혀 핵심 연출을 어긴다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "문을 잡고 출입구를 막는 배치는 A보다 가깝지만, 시선이 차 밖을 향하고 창문의 얼굴 반사가 실제 자세와 맞지 않아 사용 가능한 충실도에는 못 미친다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 왼쪽으로 팔을 뻗어 차 문 가장자리를 잡지만, 얼굴과 시선은 화면 오른쪽 전경의 다른 사람을 향한다. 요구된 차 안의 화면 밖 탑승자를 향한 시선이 아니다.",
        "built_space": "열린 세단 문 한 개와 출입구 한 개가 왼쪽에 있으며, 좌석 등받이와 머리받침이 보인다. 문 안쪽 면은 비스듬한 일부보다 거의 전체가 정면으로 드러나고 화면 중앙까지 차지한다. 박철진은 출입구 자체보다 문의 자유단 바깥에 서 있다. 배경에는 중앙 출입문 구조물 한 개, 왼쪽 초소 한 개, 확성기 기둥 두 개, 콘크리트 벽과 철조망이 보여 장소의 주요 재료와 야간 분위기는 이어진다.",
        "entities": "박철진의 중년 한국인 남성 외형, 짧은 검은 머리, 검은 작업복, 붉은 완장과 허리 파우치는 인물 참조와 대체로 맞는다. 그러나 오른쪽 전경에 별도 인물의 어깨와 몸통이 크게 등장하고, 뒤쪽 열린 시신 가방에는 이전 숏 인물로 읽히는 얼굴까지 보인다. 세단과 열린 문은 존재하며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "박철진만 보여야 하는 숏에 오른쪽 전경의 별도 인물 몸통을 추가했다.",
         "이전 숏의 인물을 이번 숏에 등장시키지 말라는 지시와 달리 시신 가방 속 얼굴을 다시 노출했다."
        ],
        "physics": "박철진의 손가락은 문의 세로 프레임을 감싸며 팔과 손목의 연결도 자연스럽다. 문은 차량 경첩에 연결되어 있고 차량과 시신 가방은 지면에 놓여 있다. 박철진의 발 접점은 하단에 잘리지만 하체는 지면으로 이어지는 서 있는 자세이며 공중에 뜬 징후는 없다."
       },
       {
        "label": "B",
        "direction": "박철진은 한 손을 왼쪽 문 가장자리에 대고 몸으로 출입구를 가로막는다. 다만 고개를 차 밖 카메라 쪽으로 돌려 바라보므로 차 안 탑승자를 향한 시선 조건은 충족하지 않는다.",
        "built_space": "열린 앞문 한 개와 그 뒤 닫힌 문 한 개가 보이고, 박철진은 열린 문과 차체 사이 출입구에 서 있다. 열린 문의 안쪽 면이 비스듬히 보여 A보다 요구된 공간 관계에 가깝지만, 문은 화면 왼쪽 끝보다는 중앙 왼쪽에 놓이고 인물도 허벅지까지 잡힌다. 배경에는 중앙 출입문 구조물 한 개, 왼쪽 초소 한 개, 확성기 기둥 두 개와 양쪽 방벽·울타리가 이어진다. 오른쪽 닫힌 창에는 실제 인물과 다른 각도로 기울어진 정면 얼굴이 나타나며, 창에 등을 비스듬히 둔 현재 자세의 반사로는 맞지 않는다.",
        "entities": "짧은 검은 머리와 중년 한국인 남성 얼굴, 어두운 작업복 및 허리 파우치는 박철진 참조에 대체로 부합한다. 붉은 완장은 현재 보이는 팔에서 확인되지 않으나 반대편 팔은 가려져 있다. 검은 세단과 열린 문은 명확하고, 시신 가방은 프레임에서 확인되지 않아 이를 누락으로 판단하지 않는다. 창 속 얼굴 형상 외에는 별도 인물이 보이지 않으며 읽히는 문구도 없다.",
        "hard_violations": [
         "닫힌 측면 창에 박철진의 실제 머리 기울기와 창에 대한 몸 방향으로 성립하지 않는 얼굴 반사를 생성했다."
        ],
        "physics": "문 가장자리에 손바닥과 손가락이 접촉하고 팔은 몸으로 자연스럽게 이어진다. 꽉 움켜쥔 힘은 A보다 약하게 읽히지만 불가능한 손 자세는 아니다. 문은 경첩으로 차체에 지지되고 허리 파우치는 벨트에 매달려 있다. 발은 화면 밖이지만 몸통과 다리는 정상적인 직립 자세로 이어져 부유 문제는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "금지된 추가 인물과 이전 숏의 시신이 등장하고, 박철진이 차 안이 아닌 바깥 인물을 바라보며 미디엄 숏보다 넓게 잡혀 핵심 연출을 어긴다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "문을 잡고 출입구를 막는 배치는 A보다 가깝지만, 시선이 차 밖을 향하고 창문의 얼굴 반사가 실제 자세와 맞지 않아 사용 가능한 충실도에는 못 미친다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 왼쪽으로 팔을 뻗어 차 문 가장자리를 잡지만, 얼굴과 시선은 화면 오른쪽 전경의 다른 사람을 향한다. 요구된 차 안의 화면 밖 탑승자를 향한 시선이 아니다.",
        "built_space": "열린 세단 문 한 개와 출입구 한 개가 왼쪽에 있으며, 좌석 등받이와 머리받침이 보인다. 문 안쪽 면은 비스듬한 일부보다 거의 전체가 정면으로 드러나고 화면 중앙까지 차지한다. 박철진은 출입구 자체보다 문의 자유단 바깥에 서 있다. 배경에는 중앙 출입문 구조물 한 개, 왼쪽 초소 한 개, 확성기 기둥 두 개, 콘크리트 벽과 철조망이 보여 장소의 주요 재료와 야간 분위기는 이어진다.",
        "entities": "박철진의 중년 한국인 남성 외형, 짧은 검은 머리, 검은 작업복, 붉은 완장과 허리 파우치는 인물 참조와 대체로 맞는다. 그러나 오른쪽 전경에 별도 인물의 어깨와 몸통이 크게 등장하고, 뒤쪽 열린 시신 가방에는 이전 숏 인물로 읽히는 얼굴까지 보인다. 세단과 열린 문은 존재하며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "박철진만 보여야 하는 숏에 오른쪽 전경의 별도 인물 몸통을 추가했다.",
         "이전 숏의 인물을 이번 숏에 등장시키지 말라는 지시와 달리 시신 가방 속 얼굴을 다시 노출했다."
        ],
        "physics": "박철진의 손가락은 문의 세로 프레임을 감싸며 팔과 손목의 연결도 자연스럽다. 문은 차량 경첩에 연결되어 있고 차량과 시신 가방은 지면에 놓여 있다. 박철진의 발 접점은 하단에 잘리지만 하체는 지면으로 이어지는 서 있는 자세이며 공중에 뜬 징후는 없다."
       },
       {
        "label": "A",
        "direction": "박철진은 한 손을 왼쪽 문 가장자리에 대고 몸으로 출입구를 가로막는다. 다만 고개를 차 밖 카메라 쪽으로 돌려 바라보므로 차 안 탑승자를 향한 시선 조건은 충족하지 않는다.",
        "built_space": "열린 앞문 한 개와 그 뒤 닫힌 문 한 개가 보이고, 박철진은 열린 문과 차체 사이 출입구에 서 있다. 열린 문의 안쪽 면이 비스듬히 보여 A보다 요구된 공간 관계에 가깝지만, 문은 화면 왼쪽 끝보다는 중앙 왼쪽에 놓이고 인물도 허벅지까지 잡힌다. 배경에는 중앙 출입문 구조물 한 개, 왼쪽 초소 한 개, 확성기 기둥 두 개와 양쪽 방벽·울타리가 이어진다. 오른쪽 닫힌 창에는 실제 인물과 다른 각도로 기울어진 정면 얼굴이 나타나며, 창에 등을 비스듬히 둔 현재 자세의 반사로는 맞지 않는다.",
        "entities": "짧은 검은 머리와 중년 한국인 남성 얼굴, 어두운 작업복 및 허리 파우치는 박철진 참조에 대체로 부합한다. 붉은 완장은 현재 보이는 팔에서 확인되지 않으나 반대편 팔은 가려져 있다. 검은 세단과 열린 문은 명확하고, 시신 가방은 프레임에서 확인되지 않아 이를 누락으로 판단하지 않는다. 창 속 얼굴 형상 외에는 별도 인물이 보이지 않으며 읽히는 문구도 없다.",
        "hard_violations": [
         "닫힌 측면 창에 박철진의 실제 머리 기울기와 창에 대한 몸 방향으로 성립하지 않는 얼굴 반사를 생성했다."
        ],
        "physics": "문 가장자리에 손바닥과 손가락이 접촉하고 팔은 몸으로 자연스럽게 이어진다. 꽉 움켜쥔 힘은 A보다 약하게 읽히지만 불가능한 손 자세는 아니다. 문은 경첩으로 차체에 지지되고 허리 파우치는 벨트에 매달려 있다. 발은 화면 밖이지만 몸통과 다리는 정상적인 직립 자세로 이어져 부유 문제는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.083
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 없는 추가 인물(창문 안의 얼굴) 등장",
     "[gpt-high] 닫힌 측면 창에 박철진의 실제 머리 기울기와 창에 대한 몸 방향으로 성립하지 않는 얼굴 반사를 생성했다."
    ],
    "B": [
     "[gemini-pro] 프롬프트에 없는 추가 인물(화면 우측 전경의 뒷모습) 등장",
     "[gpt-high] 박철진만 보여야 하는 숏에 오른쪽 전경의 별도 인물 몸통을 추가했다.",
     "[gpt-high] 이전 숏의 인물을 이번 숏에 등장시키지 말라는 지시와 달리 시신 가방 속 얼굴을 다시 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1083
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "차 문을 오른쪽에 배치해 샷 프레이밍 지시를 완전히 어겼으며, 필수 의상인 붉은 완장 누락 및 창문에 정체불명의 인물이 나타나 실격입니다.  ★위반: [gemini-pro] 프롬프트에 없는 추가 인물(창문 안의 얼굴) 등장 / [gpt-high] 닫힌 측면 창에 박철진의 실제 머리 기울기와 창에 대한 몸 방향으로 성립하지 않는 얼굴 반사를 생성했다."
   },
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "지시된 좌측 차 문 프레이밍, 배경, 의상(완장) 및 바디백 배치를 훌륭하게 구현했으나, 우측 전경에 프롬프트에 없는 인물이 크게 등장하여 실격입니다.  ★위반: [gemini-pro] 프롬프트에 없는 추가 인물(화면 우측 전경의 뒷모습) 등장 / [gpt-high] 박철진만 보여야 하는 숏에 오른쪽 전경의 별도 인물 몸통을 추가했다. / [gpt-high] 이전 숏의 인물을 이번 숏에 등장시키지 말라는 지시와 달리 시신 가방 속 얼굴을 다시 노출했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2_sel.png",
    "asset_id": "c9dd30b1-0c5a-4a75-ac73-12892c6e4433",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-8478-7a19-ac12-1bcaf476cd46",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S42sh2"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S42sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:57:11.115850+00:00",
  "fingerprint": "d099cc3674e21e3b328de18090768b4feb874c217c15653c720e742acf32dff1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S42sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S42sh11_sel.png",
  "source_sha256": "6ee8665ac82b3b8ef6d39591df4d7afd51c559174a8f42ac72ca712c8f731b28",
  "file": "S42sh11_cine.png",
  "staged_sha256": "6b2c316af83d1bb2fc9eab8ffac5f67fa2223d0d7b170f947a23af7b231baa70",
  "latency_ms": 9673
 },
 "S42sh19::signage": {
  "fp": "8f3b3442a49ea966",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S42sh19": {
  "input_fingerprint": "60683ad3f34a7551",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 주시한 채 충격을 받은 듯 눈이 동그랗게 커진 박철진의 얼어붙은 얼굴 클로즈업.\n\nLOCATION (lock): On the muddy roadside at the refugee settlement entrance at night, after the sedan has departed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established nighttime ambient treatment, with controlled facial contrast that makes the sudden loss of composure legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the muddy entrance ground, nearby structures, and nighttime lighting. Exclude the luxury sedan and its open door, since the vehicle has departed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan has left with its door closed. The opened body bag remains at the nighttime settlement entrance. 박철진: He remains at the entrance, his smile gone and his face now tense.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 주시한 채 충격을 받은 듯 눈이 동그랗게 커진 박철진의 얼어붙은 얼굴 클로즈업.\n\nLOCATION (lock): On the muddy roadside at the refugee settlement entrance at night, after the sedan has departed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established nighttime ambient treatment, with controlled facial contrast that makes the sudden loss of composure legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the muddy entrance ground, nearby structures, and nighttime lighting. Exclude the luxury sedan and its open door, since the vehicle has departed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan has left with its door closed. The opened body bag remains at the nighttime settlement entrance. 박철진: He remains at the entrance, his smile gone and his face now tense.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 주시한 채 충격을 받은 듯 눈이 동그랗게 커진 박철진의 얼어붙은 얼굴 클로즈업.\n\nLOCATION (lock): On the muddy roadside at the refugee settlement entrance at night, after the sedan has departed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established nighttime ambient treatment, with controlled facial contrast that makes the sudden loss of composure legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the muddy entrance ground, nearby structures, and nighttime lighting. Exclude the luxury sedan and its open door, since the vehicle has departed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan has left with its door closed. The opened body bag remains at the nighttime settlement entrance. 박철진: He remains at the entrance, his smile gone and his face now tense.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 손에 들고 있는 핸드폰의 디스플레이를 향하고 있음.",
    "built_space": "진흙 바닥, 좌측 콘크리트 장벽, 정면의 철문 등 지정된 배경 구조가 그대로 유지되었으며, 지시대로 세단은 제거됨.",
    "entities": "박철진의 외모와 복장이 레퍼런스와 일치하며, 뒷면에 아무런 텍스트가 없는 핸드폰을 쥐고 있음.",
    "hard_violations": [],
    "physics": "한 손으로 핸드폰을 자연스럽고 안정적으로 파지하고 있음."
   },
   {
    "label": "B",
    "direction": "크게 떠진 눈의 시선이 들고 있는 핸드폰을 향하고 있음.",
    "built_space": "진흙 바닥은 유지되었으나, 이전 샷에 고정된 좌측의 콘크리트 벽이 철조망 펜스로 임의 변경됨.",
    "entities": "박철진의 형태는 일치하나, 들고 있는 핸드폰 뒷면에 읽을 수 있는 브랜드 로고가 위치함.",
    "hard_violations": [
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트(핸드폰 뒷면 SAMSUNG 로고) 노출",
     "[gemini-pro] 이전 샷과 동일해야 할 좌측 배경 구조물(콘크리트 벽)이 철조망으로 변경됨"
    ],
    "physics": "양손으로 핸드폰을 쥐고 화면을 바라보고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 얼굴 클로즈업과 충격받은 표정을 잘 연출했으며, 이전 샷의 배경 구조를 유지하고 읽을 수 있는 텍스트 없이 지시사항을 정확히 준수함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "핸드폰 뒷면에 금지된 텍스트 로고가 선명하게 노출되었고, 고정되어야 할 좌측 배경 구조물을 임의로 변경하여 핵심 지침을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 손에 들고 있는 핸드폰의 디스플레이를 향하고 있음.",
        "built_space": "진흙 바닥, 좌측 콘크리트 장벽, 정면의 철문 등 지정된 배경 구조가 그대로 유지되었으며, 지시대로 세단은 제거됨.",
        "entities": "박철진의 외모와 복장이 레퍼런스와 일치하며, 뒷면에 아무런 텍스트가 없는 핸드폰을 쥐고 있음.",
        "hard_violations": [],
        "physics": "한 손으로 핸드폰을 자연스럽고 안정적으로 파지하고 있음."
       },
       {
        "label": "B",
        "direction": "크게 떠진 눈의 시선이 들고 있는 핸드폰을 향하고 있음.",
        "built_space": "진흙 바닥은 유지되었으나, 이전 샷에 고정된 좌측의 콘크리트 벽이 철조망 펜스로 임의 변경됨.",
        "entities": "박철진의 형태는 일치하나, 들고 있는 핸드폰 뒷면에 읽을 수 있는 브랜드 로고가 위치함.",
        "hard_violations": [
         "화면 내 읽을 수 있는 텍스트(핸드폰 뒷면 SAMSUNG 로고) 노출",
         "이전 샷과 동일해야 할 좌측 배경 구조물(콘크리트 벽)이 철조망으로 변경됨"
        ],
        "physics": "양손으로 핸드폰을 쥐고 화면을 바라보고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 얼굴 클로즈업과 충격받은 표정을 잘 연출했으며, 이전 샷의 배경 구조를 유지하고 읽을 수 있는 텍스트 없이 지시사항을 정확히 준수함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "핸드폰 뒷면에 금지된 텍스트 로고가 선명하게 노출되었고, 고정되어야 할 좌측 배경 구조물을 임의로 변경하여 핵심 지침을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 손에 들고 있는 핸드폰의 디스플레이를 향하고 있음.",
        "built_space": "진흙 바닥, 좌측 콘크리트 장벽, 정면의 철문 등 지정된 배경 구조가 그대로 유지되었으며, 지시대로 세단은 제거됨.",
        "entities": "박철진의 외모와 복장이 레퍼런스와 일치하며, 뒷면에 아무런 텍스트가 없는 핸드폰을 쥐고 있음.",
        "hard_violations": [],
        "physics": "한 손으로 핸드폰을 자연스럽고 안정적으로 파지하고 있음."
       },
       {
        "label": "B",
        "direction": "크게 떠진 눈의 시선이 들고 있는 핸드폰을 향하고 있음.",
        "built_space": "진흙 바닥은 유지되었으나, 이전 샷에 고정된 좌측의 콘크리트 벽이 철조망 펜스로 임의 변경됨.",
        "entities": "박철진의 형태는 일치하나, 들고 있는 핸드폰 뒷면에 읽을 수 있는 브랜드 로고가 위치함.",
        "hard_violations": [
         "화면 내 읽을 수 있는 텍스트(핸드폰 뒷면 SAMSUNG 로고) 노출",
         "이전 샷과 동일해야 할 좌측 배경 구조물(콘크리트 벽)이 철조망으로 변경됨"
        ],
        "physics": "양손으로 핸드폰을 쥐고 화면을 바라보고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "휴대폰을 응시하는 긴장된 표정과 인물 외모, 기존 입구의 건축·야간 조명을 더 충실히 유지하지만, 얼굴 클로즈업보다 구도가 넓다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "휴대폰을 보고 놀라는 행동은 분명하지만, 얼굴 클로즈업보다 넓으며 기존 장소의 배치·조명과 인물의 피부 상태가 크게 달라졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈은 얼굴 아래쪽의 휴대폰을 향한다. 휴대폰의 후면 카메라가 관객에게 보이고 화면은 인물 쪽을 향하므로 읽는 방향은 맞는다. 눈꺼풀을 크게 벌리고 입을 조금 연 모습으로 충격을 표현한다.",
        "built_space": "왼쪽에 철망 울타리 한 줄, 오른쪽에 콘크리트 담장 한 줄, 그 사이에 젖은 도로와 여러 조명이 보인다. 이전 사진의 중앙 출입문과 콘크리트 문틀은 뚜렷하게 확인되지 않는다. 다른 시점일 가능성은 있지만 동일 장소를 입증하는 구조적 단서가 약하며, 푸른 조명이 이전 사진보다 강하다. 인물은 화면 오른쪽에 있고 가슴과 양손까지 포함되어 얼굴 클로즈업보다 넓다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명과 휴대폰 한 대가 보인다. 어두운 작업복과 어깨 장식은 참고와 유사하지만 얼굴과 손에 참고보다 훨씬 짙은 오염이 추가되었다. 세단과 다른 사람은 없다. 오른쪽 도로의 검은 주머니 모양 물체가 열린 시신 가방인지는 확정하기 어렵다. 왼쪽에는 잡동사니가 보인다. 휴대폰 뒷면의 작은 표식은 명확히 판독되지 않는다.",
        "hard_violations": [],
        "physics": "휴대폰은 손가락과 손바닥 사이에 잡혀 있고 다른 손이 아래쪽을 받치는 모습이다. 손목과 팔은 몸통으로 이어져 지지가 성립한다. 하체는 화면 밖이므로 발의 접지는 판단할 수 없지만 몸이 떠 있다는 증거는 없다. 도로의 주머니와 잡동사니는 지면에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "시선은 화면 왼쪽 전경에 든 휴대폰을 향한다. 관객에게는 휴대폰 뒷면과 카메라가 보이며 화면은 인물의 눈 쪽을 향한다. 눈을 넓게 뜨고 입술을 살짝 벌린 채 굳은 표정으로, 미소가 사라진 충격 상태가 읽힌다.",
        "built_space": "왼쪽 콘크리트 담장, 중앙의 콘크리트 출입문 틀 한 개와 금속 출입문 한 곳, 오른쪽 철망과 낮은 시설물이 이전 사진의 배치를 유지한다. 상단에는 방송 설비가 달린 기둥 두 개가 일부 보이며, 도로의 물웅덩이에 야간 조명이 반사된다. 세단은 없다. 인물은 오른쪽에 배치되지만 얼굴뿐 아니라 가슴과 긴 팔까지 보여 요청한 얼굴 클로즈업보다 넓다.",
        "entities": "짧은 검은 머리, 중년의 얼굴 윤곽을 가진 동아시아계 남성 한 명으로 참고 속 박철진과 더 가깝다. 어두운 작업복의 칼라와 어깨 장식도 이전 사진에 잘 연결된다. 성인 남성의 손이 휴대폰 한 대를 쥐고 있으며 소매도 해당 의상과 이어진다. 추가 인물이나 세단, 읽을 수 있는 글자는 없다. 열린 시신 가방은 확인되지 않지만 화면 밖에 있을 수 있어 그 부재만으로 연속성 위반을 확정할 수 없다.",
        "hard_violations": [],
        "physics": "휴대폰은 엄지와 나머지 손가락 사이에 안정적으로 잡혀 있다. 손목과 소매, 팔이 몸통으로 자연스럽게 이어지고, 휴대폰을 들어 읽다가 멈춘 자세로 성립한다. 하체는 잘려 있으므로 발의 지지는 보이지 않지만 공중 부양이나 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "휴대폰을 응시하는 긴장된 표정과 인물 외모, 기존 입구의 건축·야간 조명을 더 충실히 유지하지만, 얼굴 클로즈업보다 구도가 넓다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "휴대폰을 보고 놀라는 행동은 분명하지만, 얼굴 클로즈업보다 넓으며 기존 장소의 배치·조명과 인물의 피부 상태가 크게 달라졌다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "두 눈은 얼굴 아래쪽의 휴대폰을 향한다. 휴대폰의 후면 카메라가 관객에게 보이고 화면은 인물 쪽을 향하므로 읽는 방향은 맞는다. 눈꺼풀을 크게 벌리고 입을 조금 연 모습으로 충격을 표현한다.",
        "built_space": "왼쪽에 철망 울타리 한 줄, 오른쪽에 콘크리트 담장 한 줄, 그 사이에 젖은 도로와 여러 조명이 보인다. 이전 사진의 중앙 출입문과 콘크리트 문틀은 뚜렷하게 확인되지 않는다. 다른 시점일 가능성은 있지만 동일 장소를 입증하는 구조적 단서가 약하며, 푸른 조명이 이전 사진보다 강하다. 인물은 화면 오른쪽에 있고 가슴과 양손까지 포함되어 얼굴 클로즈업보다 넓다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명과 휴대폰 한 대가 보인다. 어두운 작업복과 어깨 장식은 참고와 유사하지만 얼굴과 손에 참고보다 훨씬 짙은 오염이 추가되었다. 세단과 다른 사람은 없다. 오른쪽 도로의 검은 주머니 모양 물체가 열린 시신 가방인지는 확정하기 어렵다. 왼쪽에는 잡동사니가 보인다. 휴대폰 뒷면의 작은 표식은 명확히 판독되지 않는다.",
        "hard_violations": [],
        "physics": "휴대폰은 손가락과 손바닥 사이에 잡혀 있고 다른 손이 아래쪽을 받치는 모습이다. 손목과 팔은 몸통으로 이어져 지지가 성립한다. 하체는 화면 밖이므로 발의 접지는 판단할 수 없지만 몸이 떠 있다는 증거는 없다. 도로의 주머니와 잡동사니는 지면에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "시선은 화면 왼쪽 전경에 든 휴대폰을 향한다. 관객에게는 휴대폰 뒷면과 카메라가 보이며 화면은 인물의 눈 쪽을 향한다. 눈을 넓게 뜨고 입술을 살짝 벌린 채 굳은 표정으로, 미소가 사라진 충격 상태가 읽힌다.",
        "built_space": "왼쪽 콘크리트 담장, 중앙의 콘크리트 출입문 틀 한 개와 금속 출입문 한 곳, 오른쪽 철망과 낮은 시설물이 이전 사진의 배치를 유지한다. 상단에는 방송 설비가 달린 기둥 두 개가 일부 보이며, 도로의 물웅덩이에 야간 조명이 반사된다. 세단은 없다. 인물은 오른쪽에 배치되지만 얼굴뿐 아니라 가슴과 긴 팔까지 보여 요청한 얼굴 클로즈업보다 넓다.",
        "entities": "짧은 검은 머리, 중년의 얼굴 윤곽을 가진 동아시아계 남성 한 명으로 참고 속 박철진과 더 가깝다. 어두운 작업복의 칼라와 어깨 장식도 이전 사진에 잘 연결된다. 성인 남성의 손이 휴대폰 한 대를 쥐고 있으며 소매도 해당 의상과 이어진다. 추가 인물이나 세단, 읽을 수 있는 글자는 없다. 열린 시신 가방은 확인되지 않지만 화면 밖에 있을 수 있어 그 부재만으로 연속성 위반을 확정할 수 없다.",
        "hard_violations": [],
        "physics": "휴대폰은 엄지와 나머지 손가락 사이에 안정적으로 잡혀 있다. 손목과 소매, 팔이 몸통으로 자연스럽게 이어지고, 휴대폰을 들어 읽다가 멈춘 자세로 성립한다. 하체는 잘려 있으므로 발의 지지는 보이지 않지만 공중 부양이나 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트(핸드폰 뒷면 SAMSUNG 로고) 노출",
     "[gemini-pro] 이전 샷과 동일해야 할 좌측 배경 구조물(콘크리트 벽)이 철조망으로 변경됨"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 얼굴 클로즈업과 충격받은 표정을 잘 연출했으며, 이전 샷의 배경 구조를 유지하고 읽을 수 있는 텍스트 없이 지시사항을 정확히 준수함."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "핸드폰 뒷면에 금지된 텍스트 로고가 선명하게 노출되었고, 고정되어야 할 좌측 배경 구조물을 임의로 변경하여 핵심 지침을 위반함.  ★위반: [gemini-pro] 화면 내 읽을 수 있는 텍스트(핸드폰 뒷면 SAMSUNG 로고) 노출 / [gemini-pro] 이전 샷과 동일해야 할 좌측 배경 구조물(콘크리트 벽)이 철조망으로 변경됨"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh11_sel.png",
    "asset_id": "5ed70ede-665b-494b-a51a-00198175fffb",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-8623-767e-aa8f-84ec9af0112c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S42sh11"
  }
 },
 "S42sh19::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:58:05.582608+00:00",
  "fingerprint": "fb5076e42015a78e91299b421ae9d5c902fc38f2b55b786a303c38d8e33ac50b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S42sh19_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S42sh19_sel.png",
  "source_sha256": "b8984dfa95ecb94239998e65e13ee4ad30758db1e556fe16b4f6de47cf5294c6",
  "file": "S42sh19_cine.png",
  "staged_sha256": "b959ccdb9e69014f8d3bb578ff0111138db97e46a8c283c2620582ea48e11220",
  "latency_ms": 9681
 },
 "S43sh2::signage": {
  "fp": "1d602c0a44fb091c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5c5e686b8cecba11": {
  "subjects": [],
  "subject_text": "국방장관 집무실\n집무용 책상과 의자, 전화가 놓인 사무 공간. 책상 앞쪽에는 대면할 수 있는 좌석 공간이 마련돼 있다.",
  "identity": "canonical",
  "scope_id": "L196",
  "scope_role": "location_interior",
  "scope_sha": "eb6522f45277b7b9"
 },
 "S43sh2::bgfirst_bg": {
  "input_fingerprint": "158b48a517dc33f0",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2__bgfirst_bg.png",
  "asset_id": "a14ea04b-9a30-4065-8253-7e970e47fb07",
  "input_asset_ids": [
   "b8079413-7782-40e9-9e87-17589712c08c",
   "345f0672-c175-4e82-98e7-7fba2f349457"
  ]
 },
 "S43sh2": {
  "input_fingerprint": "db8be656261ab674",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife previously thrown into the militia-office wall remains embedded there. 국방장관: He is in his own office, engaged in the same telephone call.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife previously thrown into the militia-office wall remains embedded there. 국방장관: He is in his own office, engaged in the same telephone call.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife previously thrown into the militia-office wall remains embedded there. 국방장관: He is in his own office, engaged in the same telephone call.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2__bgfirst_bg.png",
     "asset_id": "a14ea04b-9a30-4065-8253-7e970e47fb07",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S43sh2.png",
     "asset_id": "b8079413-7782-40e9-9e87-17589712c08c",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:835328>",
     "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L196B01.png",
     "asset_id": "345f0672-c175-4e82-98e7-7fba2f349457",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:835328>",
     "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "화면 밖 왼쪽 허공을 응시하고 있음.",
    "built_space": "레퍼런스와 일치하는 넓은 집무실(책상, 깃발, 지도, 암체어 등 배치). 창밖으로 야경이 보임.",
    "entities": "국방장관의 얼굴과 헤어는 일치하나 넥타이를 매지 않았음. 검은색 스마트폰을 들고 있음.",
    "hard_violations": [],
    "physics": "소파에 기대어 다리를 꼬고 앉아 있으며, 왼손으로 전화를 귀에 대고 오른팔은 소파 등받이에 걸치고 있음."
   },
   {
    "label": "B",
    "direction": "화면 밖 오른쪽 아래 허공을 서늘하게 응시하고 있음.",
    "built_space": "가죽 소파와 창문 밖 야경, 하단 우드 패널이 보이는 집무실 창가 위치.",
    "entities": "국방장관의 얼굴, 헤어스타일, 정장 및 넥타이가 레퍼런스와 일치함(조끼는 임의 추가됨). 투명 케이스 스마트폰을 들고 있음.",
    "hard_violations": [
     "[gpt-high] 지정된 소파 자리를 산악 사진 아래의 측벽에서 창문·수납장 바로 앞으로 옮겨, 정확한 촬영 장소의 공간 배치를 변경했다."
    ],
    "physics": "소파에 등을 깊이 기대고 앉아 있으며, 오른손으로 전화기를 귀에 안정적으로 대고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지시된 미디엄 샷 프레이밍과 상체 위주의 구도를 잘 따랐으며, 인물의 의상(넥타이 포함)도 레퍼런스에 더 부합합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "집무실 배경은 훌륭하게 구현되었으나, 미디엄 샷 지시를 어기고 앵글이 넓어져 하반신과 공간이 과도하게 노출되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "화면 밖 오른쪽 아래 허공을 서늘하게 응시하고 있음.",
        "built_space": "가죽 소파와 창문 밖 야경, 하단 우드 패널이 보이는 집무실 창가 위치.",
        "entities": "국방장관의 얼굴, 헤어스타일, 정장 및 넥타이가 레퍼런스와 일치함(조끼는 임의 추가됨). 투명 케이스 스마트폰을 들고 있음.",
        "hard_violations": [],
        "physics": "소파에 등을 깊이 기대고 앉아 있으며, 오른손으로 전화기를 귀에 안정적으로 대고 있음."
       },
       {
        "label": "A",
        "direction": "화면 밖 왼쪽 허공을 응시하고 있음.",
        "built_space": "레퍼런스와 일치하는 넓은 집무실(책상, 깃발, 지도, 암체어 등 배치). 창밖으로 야경이 보임.",
        "entities": "국방장관의 얼굴과 헤어는 일치하나 넥타이를 매지 않았음. 검은색 스마트폰을 들고 있음.",
        "hard_violations": [],
        "physics": "소파에 기대어 다리를 꼬고 앉아 있으며, 왼손으로 전화를 귀에 대고 오른팔은 소파 등받이에 걸치고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지시된 미디엄 샷 프레이밍과 상체 위주의 구도를 잘 따랐으며, 인물의 의상(넥타이 포함)도 레퍼런스에 더 부합합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "집무실 배경은 훌륭하게 구현되었으나, 미디엄 샷 지시를 어기고 앵글이 넓어져 하반신과 공간이 과도하게 노출되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 밖 오른쪽 아래 허공을 서늘하게 응시하고 있음.",
        "built_space": "가죽 소파와 창문 밖 야경, 하단 우드 패널이 보이는 집무실 창가 위치.",
        "entities": "국방장관의 얼굴, 헤어스타일, 정장 및 넥타이가 레퍼런스와 일치함(조끼는 임의 추가됨). 투명 케이스 스마트폰을 들고 있음.",
        "hard_violations": [],
        "physics": "소파에 등을 깊이 기대고 앉아 있으며, 오른손으로 전화기를 귀에 안정적으로 대고 있음."
       },
       {
        "label": "A",
        "direction": "화면 밖 왼쪽 허공을 응시하고 있음.",
        "built_space": "레퍼런스와 일치하는 넓은 집무실(책상, 깃발, 지도, 암체어 등 배치). 창밖으로 야경이 보임.",
        "entities": "국방장관의 얼굴과 헤어는 일치하나 넥타이를 매지 않았음. 검은색 스마트폰을 들고 있음.",
        "hard_violations": [],
        "physics": "소파에 기대어 다리를 꼬고 앉아 있으며, 왼손으로 전화를 귀에 대고 오른팔은 소파 등받이에 걸치고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "여유롭게 기댄 상체 중심의 구도는 더 가깝지만, 소파를 지정된 벽면이 아닌 창문 바로 앞으로 옮겨 장소의 공간 배치를 위반한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 집무실 소파 위치와 통화·기댄 자세는 충실하지만, 상체 중심 미디엄 샷보다 넓고 참고 의상의 넥타이가 빠져 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남성은 카메라가 아닌 화면 오른쪽의 허공을 차분하고 서늘하게 바라본다. 휴대전화는 손으로 귀에 대고 있으며 통화에 맞는 방향이다. 겨누는 물체나 이동하는 인물은 없다.",
        "built_space": "검은 소파 한 개의 좌판과 등받이, 왼쪽 팔걸이가 비스듬히 보인다. 소파 바로 뒤에는 창문들이 이어지고 그 아래 목재 수납장이 있다. 그러나 장소 참고에서 해당 소파는 산악 사진이 걸린 왼쪽 벽에 등을 대고 있으며, 창문과 수납장은 소파 뒤가 아니라 옆쪽에 있다. 이는 작은 배경 이동이 아니라 착석 장소의 벽면 관계를 바꾼 배치다. 불가능한 반사나 중복 인물은 보이지 않는다.",
        "entities": "성인 한국인 남성으로 설정에 부합하는 인물 한 명이며 짧고 단정한 검은 머리와 얼굴은 참고 인물에 대체로 가깝다. 짙은 정장, 흰 셔츠, 무늬 넥타이는 맞지만 참고에 없는 조끼가 보인다. 휴대전화 한 대를 사용하고 있으며 창밖은 밤이다. 읽을 수 있는 글자는 보이지 않는다. 다른 사무실 벽에 남아 있어야 하는 칼은 이 구도에 나타나지 않는다.",
        "hard_violations": [
         "지정된 소파 자리를 산악 사진 아래의 측벽에서 창문·수납장 바로 앞으로 옮겨, 정확한 촬영 장소의 공간 배치를 변경했다."
        ],
        "physics": "골반과 허벅지는 좌판에 놓이고 등은 등받이에 기대어 뒤로 실린 체중을 지탱한다. 전화기를 든 팔꿈치는 팔걸이에 얹혀 있으며 손가락이 전화기를 잡고 있다. 포갠 다리도 서로와 좌판의 지지를 받는다. 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "남성은 화면 오른쪽 위쪽의 빈 공간을 바라보며 렌즈와 눈을 맞추지 않는다. 휴대전화는 손에 잡혀 귀에 닿아 있고, 기기의 기울기도 통화 동작으로 자연스럽다. 다른 사람을 응시하거나 물체를 겨누는 장면은 아니다.",
        "built_space": "왼쪽 벽의 소파 한 개에 인물이 앉아 있고, 앞에는 낮은 탁자 한 개, 맞은편에는 안락의자 한 개가 있다. 오른쪽에는 집무 책상 한 개와 집무 의자 한 개, 깃발 두 개, 책장 한 개가 보인다. 뒤쪽 창 두 구역 사이의 한반도 지도 한 점, 왼쪽 산악 사진 한 점과 스탠드 한 개도 참고의 배치에 부합한다. 다만 방과 다리를 넓게 포함해 요청한 상체 중심 미디엄 샷보다 환경 설명의 비중이 크다.",
        "entities": "성인 한국인 남성으로 설정에 부합하는 인물 한 명이다. 짧은 검은 머리, 얼굴과 체격은 참고에 대체로 가깝다. 짙은 정장과 흰 셔츠는 맞지만 셔츠 깃을 풀었고 참고의 넥타이는 없다. 휴대전화 한 대로 통화 중이며, 검은 가죽 소파와 목재 가구는 장소 참고와 잘 맞는다. 창밖 야경으로 밤이 분명하고 읽을 수 있는 문구는 없다. 다른 장소에 있는 칼을 추가하지 않았다.",
        "hard_violations": [],
        "physics": "골반은 소파 좌판에 놓이고 뒤로 기운 몸통은 등받이가 받친다. 한쪽 팔은 소파 상단에 걸쳐 지지되고 다른 손은 전화기를 귀에 붙여 잡는다. 포갠 다리는 좌판과 반대쪽 다리가 받친다. 탁자의 책과 화분도 표면에 놓여 있으며 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "여유롭게 기댄 상체 중심의 구도는 더 가깝지만, 소파를 지정된 벽면이 아닌 창문 바로 앞으로 옮겨 장소의 공간 배치를 위반한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 집무실 소파 위치와 통화·기댄 자세는 충실하지만, 상체 중심 미디엄 샷보다 넓고 참고 의상의 넥타이가 빠져 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "남성은 카메라가 아닌 화면 오른쪽의 허공을 차분하고 서늘하게 바라본다. 휴대전화는 손으로 귀에 대고 있으며 통화에 맞는 방향이다. 겨누는 물체나 이동하는 인물은 없다.",
        "built_space": "검은 소파 한 개의 좌판과 등받이, 왼쪽 팔걸이가 비스듬히 보인다. 소파 바로 뒤에는 창문들이 이어지고 그 아래 목재 수납장이 있다. 그러나 장소 참고에서 해당 소파는 산악 사진이 걸린 왼쪽 벽에 등을 대고 있으며, 창문과 수납장은 소파 뒤가 아니라 옆쪽에 있다. 이는 작은 배경 이동이 아니라 착석 장소의 벽면 관계를 바꾼 배치다. 불가능한 반사나 중복 인물은 보이지 않는다.",
        "entities": "성인 한국인 남성으로 설정에 부합하는 인물 한 명이며 짧고 단정한 검은 머리와 얼굴은 참고 인물에 대체로 가깝다. 짙은 정장, 흰 셔츠, 무늬 넥타이는 맞지만 참고에 없는 조끼가 보인다. 휴대전화 한 대를 사용하고 있으며 창밖은 밤이다. 읽을 수 있는 글자는 보이지 않는다. 다른 사무실 벽에 남아 있어야 하는 칼은 이 구도에 나타나지 않는다.",
        "hard_violations": [
         "지정된 소파 자리를 산악 사진 아래의 측벽에서 창문·수납장 바로 앞으로 옮겨, 정확한 촬영 장소의 공간 배치를 변경했다."
        ],
        "physics": "골반과 허벅지는 좌판에 놓이고 등은 등받이에 기대어 뒤로 실린 체중을 지탱한다. 전화기를 든 팔꿈치는 팔걸이에 얹혀 있으며 손가락이 전화기를 잡고 있다. 포갠 다리도 서로와 좌판의 지지를 받는다. 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "남성은 화면 오른쪽 위쪽의 빈 공간을 바라보며 렌즈와 눈을 맞추지 않는다. 휴대전화는 손에 잡혀 귀에 닿아 있고, 기기의 기울기도 통화 동작으로 자연스럽다. 다른 사람을 응시하거나 물체를 겨누는 장면은 아니다.",
        "built_space": "왼쪽 벽의 소파 한 개에 인물이 앉아 있고, 앞에는 낮은 탁자 한 개, 맞은편에는 안락의자 한 개가 있다. 오른쪽에는 집무 책상 한 개와 집무 의자 한 개, 깃발 두 개, 책장 한 개가 보인다. 뒤쪽 창 두 구역 사이의 한반도 지도 한 점, 왼쪽 산악 사진 한 점과 스탠드 한 개도 참고의 배치에 부합한다. 다만 방과 다리를 넓게 포함해 요청한 상체 중심 미디엄 샷보다 환경 설명의 비중이 크다.",
        "entities": "성인 한국인 남성으로 설정에 부합하는 인물 한 명이다. 짧은 검은 머리, 얼굴과 체격은 참고에 대체로 가깝다. 짙은 정장과 흰 셔츠는 맞지만 셔츠 깃을 풀었고 참고의 넥타이는 없다. 휴대전화 한 대로 통화 중이며, 검은 가죽 소파와 목재 가구는 장소 참고와 잘 맞는다. 창밖 야경으로 밤이 분명하고 읽을 수 있는 문구는 없다. 다른 장소에 있는 칼을 추가하지 않았다.",
        "hard_violations": [],
        "physics": "골반은 소파 좌판에 놓이고 뒤로 기운 몸통은 등받이가 받친다. 한쪽 팔은 소파 상단에 걸쳐 지지되고 다른 손은 전화기를 귀에 붙여 잡는다. 포갠 다리는 좌판과 반대쪽 다리가 받친다. 탁자의 책과 화분도 표면에 놓여 있으며 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.667,
    "B": 1.321
   },
   "violations": {
    "B": [
     "[gpt-high] 지정된 소파 자리를 산악 사진 아래의 측벽에서 창문·수납장 바로 앞으로 옮겨, 정확한 촬영 장소의 공간 배치를 변경했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1321,
   "A": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "지시된 미디엄 샷 프레이밍과 상체 위주의 구도를 잘 따랐으며, 인물의 의상(넥타이 포함)도 레퍼런스에 더 부합합니다.  ★위반: [gpt-high] 지정된 소파 자리를 산악 사진 아래의 측벽에서 창문·수납장 바로 앞으로 옮겨, 정확한 촬영 장소의 공간 배치를 변경했다."
   },
   {
    "label": "A",
    "score": 1667,
    "verdict_ko": "집무실 배경은 훌륭하게 구현되었으나, 미디엄 샷 지시를 어기고 앵글이 넓어져 하반신과 공간이 과도하게 노출되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L196B01.png",
    "asset_id": "345f0672-c175-4e82-98e7-7fba2f349457",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:835328>",
    "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-87ce-738a-a469-47c2a2804360",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2__bgfirst_bg.png",
   "bg_asset_id": "a14ea04b-9a30-4065-8253-7e970e47fb07",
   "bg_record_key": "S43sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S43sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T17:59:18.548265+00:00",
  "fingerprint": "084d6a67ba84ad0df2b49d0f25da52e56685c943ef478e27328126928970babe",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S43sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S43sh2_sel.png",
  "source_sha256": "76720cc8dad8fe6c79aa300f60ed06981edeea9191057f4481b1aefd7a029ea2",
  "file": "S43sh2_cine.png",
  "staged_sha256": "7a2fd8f73f7a5e35aa9c6237e198acfe4b2fe070ef505ea22861992be748f745",
  "latency_ms": 10765
 },
 "S43sh4::signage": {
  "fp": "f657f571eefa4ecc",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S43sh4": {
  "input_fingerprint": "6e5c5b40e94deb5d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 귀에 댄 채 허리를 90도로 굽히고 비굴하게 고개를 숙인 박철진의 전신.\n\nLOCATION (lock): Inside the militia commander's own office, in the open standing area under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 핸드폰 (Held to 박철진's ear during the call) — Its edge and back are visible beside his lowered head; no screen content is presented; used as Links the otherwise unseen authority to his full-body submission.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambient illumination with restrained contrast, expressing submission through framing rather than a fabricated lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He remains on the telephone, bowing submissively with a trembling hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 귀에 댄 채 허리를 90도로 굽히고 비굴하게 고개를 숙인 박철진의 전신.\n\nLOCATION (lock): Inside the militia commander's own office, in the open standing area under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 핸드폰 (Held to 박철진's ear during the call) — Its edge and back are visible beside his lowered head; no screen content is presented; used as Links the otherwise unseen authority to his full-body submission.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambient illumination with restrained contrast, expressing submission through framing rather than a fabricated lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He remains on the telephone, bowing submissively with a trembling hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 귀에 댄 채 허리를 90도로 굽히고 비굴하게 고개를 숙인 박철진의 전신.\n\nLOCATION (lock): Inside the militia commander's own office, in the open standing area under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 핸드폰 (Held to 박철진's ear during the call) — Its edge and back are visible beside his lowered head; no screen content is presented; used as Links the otherwise unseen authority to his full-body submission.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambient illumination with restrained contrast, expressing submission through framing rather than a fabricated lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He remains on the telephone, bowing submissively with a trembling hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물은 카메라 방향을 향해 서서 바닥으로 깊게 고개를 숙이고 있음.",
    "built_space": "사무실 내부. 좌측에 책상 1개, 우측 벽에 칼 1개, 후면에 화이트보드 1개. 인물의 위치와 공간 구조가 레퍼런스와 정확히 일치함.",
    "entities": "박철진(검은 머리, 검은 제복). 오른손으로 핸드폰을 귀에 대고 있음. 붉은 완장이 오른팔에 위치함.",
    "hard_violations": [],
    "physics": "두 발로 바닥을 안정적으로 지지하며 허리를 90도로 굽힌 자세가 자연스러움. 오른손이 핸드폰을 단단히 쥐고 있으며, 칼은 벽에 올바르게 박혀 있음."
   },
   {
    "label": "B",
    "direction": "인물은 우측 벽을 향해 서서 벽면을 바라보며 시선을 두고 있음.",
    "built_space": "사무실 내부. 좌측에 책상 1개, 우측 벽에 칼 1개, 후면에 화이트보드 1개. 가구와 공간 배치는 적절함.",
    "entities": "박철진. 왼손으로 핸드폰을 귀에 대고 있음. 붉은 완장이 왼팔에 위치함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학: 인물의 허리 뒤쪽에 팔과 연결되지 않은 기형적인 작은 손가락들이 떠 있는 형태로 렌더링됨."
    ],
    "physics": "두 발로 서 있으나 허리를 지정된 90도로 굽히지 않고 애매한 각도로 서 있음. 허리 뒤의 손은 신체와 정상적으로 연결되지 않아 물리적 구조가 깨짐. 칼은 벽에 고정됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "90도로 허리를 굽히고 비굴하게 통화하는 핵심 행동을 완벽히 연출했으며, 공간과 소품의 물리적 일관성도 매우 훌륭합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "허리를 90도로 굽혀야 하는 지시를 따르지 않았으며, 벽을 향해 어색하게 서 있고 신체 렌더링 오류(허리 뒤쪽 손)가 심각합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 카메라 방향을 향해 서서 바닥으로 깊게 고개를 숙이고 있음.",
        "built_space": "사무실 내부. 좌측에 책상 1개, 우측 벽에 칼 1개, 후면에 화이트보드 1개. 인물의 위치와 공간 구조가 레퍼런스와 정확히 일치함.",
        "entities": "박철진(검은 머리, 검은 제복). 오른손으로 핸드폰을 귀에 대고 있음. 붉은 완장이 오른팔에 위치함.",
        "hard_violations": [],
        "physics": "두 발로 바닥을 안정적으로 지지하며 허리를 90도로 굽힌 자세가 자연스러움. 오른손이 핸드폰을 단단히 쥐고 있으며, 칼은 벽에 올바르게 박혀 있음."
       },
       {
        "label": "B",
        "direction": "인물은 우측 벽을 향해 서서 벽면을 바라보며 시선을 두고 있음.",
        "built_space": "사무실 내부. 좌측에 책상 1개, 우측 벽에 칼 1개, 후면에 화이트보드 1개. 가구와 공간 배치는 적절함.",
        "entities": "박철진. 왼손으로 핸드폰을 귀에 대고 있음. 붉은 완장이 왼팔에 위치함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학: 인물의 허리 뒤쪽에 팔과 연결되지 않은 기형적인 작은 손가락들이 떠 있는 형태로 렌더링됨."
        ],
        "physics": "두 발로 서 있으나 허리를 지정된 90도로 굽히지 않고 애매한 각도로 서 있음. 허리 뒤의 손은 신체와 정상적으로 연결되지 않아 물리적 구조가 깨짐. 칼은 벽에 고정됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "90도로 허리를 굽히고 비굴하게 통화하는 핵심 행동을 완벽히 연출했으며, 공간과 소품의 물리적 일관성도 매우 훌륭합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "허리를 90도로 굽혀야 하는 지시를 따르지 않았으며, 벽을 향해 어색하게 서 있고 신체 렌더링 오류(허리 뒤쪽 손)가 심각합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 카메라 방향을 향해 서서 바닥으로 깊게 고개를 숙이고 있음.",
        "built_space": "사무실 내부. 좌측에 책상 1개, 우측 벽에 칼 1개, 후면에 화이트보드 1개. 인물의 위치와 공간 구조가 레퍼런스와 정확히 일치함.",
        "entities": "박철진(검은 머리, 검은 제복). 오른손으로 핸드폰을 귀에 대고 있음. 붉은 완장이 오른팔에 위치함.",
        "hard_violations": [],
        "physics": "두 발로 바닥을 안정적으로 지지하며 허리를 90도로 굽힌 자세가 자연스러움. 오른손이 핸드폰을 단단히 쥐고 있으며, 칼은 벽에 올바르게 박혀 있음."
       },
       {
        "label": "B",
        "direction": "인물은 우측 벽을 향해 서서 벽면을 바라보며 시선을 두고 있음.",
        "built_space": "사무실 내부. 좌측에 책상 1개, 우측 벽에 칼 1개, 후면에 화이트보드 1개. 가구와 공간 배치는 적절함.",
        "entities": "박철진. 왼손으로 핸드폰을 귀에 대고 있음. 붉은 완장이 왼팔에 위치함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학: 인물의 허리 뒤쪽에 팔과 연결되지 않은 기형적인 작은 손가락들이 떠 있는 형태로 렌더링됨."
        ],
        "physics": "두 발로 서 있으나 허리를 지정된 90도로 굽히지 않고 애매한 각도로 서 있음. 허리 뒤의 손은 신체와 정상적으로 연결되지 않아 물리적 구조가 깨짐. 칼은 벽에 고정됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "측면 전신 와이드숏으로 약 90도의 허리 굽힘과 숙인 고개, 귀에 댄 휴대전화의 뒷면을 가장 명확하게 보여준다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "전신 통화와 깊은 인사는 충실하지만, 정면 단축으로 허리의 90도 굽힘이 덜 명확하고 휴대전화도 주로 얇은 모서리만 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "몸통은 화면 오른쪽으로 거의 수평하게 숙여져 있고 얼굴과 시선은 발 앞쪽 바닥을 향한다. 오른손의 휴대전화는 귀에 밀착되어 있으며 화면은 머리 쪽, 뒷면과 카메라 부분은 관객 쪽이다. 칼은 손잡이가 왼쪽 위로 돌출되고 칼끝이 오른쪽 아래 벽 안으로 들어가 있어 사람을 겨누는 상태가 아니다.",
        "built_space": "뒤쪽 목재 책상 1개, 그 아래 일부 보이는 바퀴 의자 1개, 오른쪽 게시판 1개, 왼쪽의 높은 금속 수납장 1개와 뒤쪽에 일부 드러난 수납장 1개가 보인다. 오른쪽 돌출 벽에는 칼 1개가 박혀 있다. 인물은 책상 앞의 빈 바닥에 서 있으며 가구와 겹쳐 들어가지 않는다. 흰색·회색 투톤 벽, 균열, 검은 걸레받이와 회색 바닥이 이전 장면과 잘 이어진다. 반사는 없다.",
        "entities": "인물은 한 명뿐이며 짧은 검은 머리의 중년 동아시아계 남성으로 박철진의 외형과 부합한다. 숙인 얼굴 때문에 세부 얼굴 일치 여부는 제한적으로 보인다. 바랜 검정 전투복, 붉은 완장과 검정 표식, 허리 파우치, 검정 전투화가 참조와 맞는다. 검정 휴대전화와 회색 손잡이의 벽에 박힌 칼도 확인된다. 읽을 수 있는 글자는 없다. 중립적인 사무실 조명이며 밤이라는 외부 단서는 확실하지 않다.",
        "hard_violations": [],
        "physics": "두 전투화가 바닥에 닿아 체중을 지탱하고, 약간 굽힌 무릎과 뒤로 둔 골반이 앞으로 숙인 상체를 받친다. 전화기는 오른손으로 쥐고 귀에 대고 있으며 반대 손은 등 뒤에 있다. 칼은 벽에 박힌 날 끝으로 지지된다. 부유하거나 지지되지 않는 물체는 없다. 손의 떨림은 정지 화면에서 뚜렷하게 확인되지 않는다."
       },
       {
        "label": "B",
        "direction": "인물은 카메라 방향으로 몸을 숙이되 얼굴과 시선은 아래 바닥으로 향한다. 오른손은 전화기를 오른쪽 귀에 대고 있으며 관객에게는 주로 기기의 얇은 가장자리가 보이고 화면 내용은 보이지 않는다. 칼의 손잡이는 왼쪽 위를 향하고 날 끝은 오른쪽 아래 벽에 박혀 있다.",
        "built_space": "뒤쪽 왼편에 목재 책상 1개와 그 아래 일부 보이는 바퀴 의자 1개, 오른쪽 뒤 벽에 게시판 1개, 인물 뒤에 수납장 윗부분으로 보이는 구조가 있다. 오른쪽 돌출 벽의 칼은 1개다. 박철진은 책상 앞 빈 공간에서 서서 인사한다. 투톤 벽과 균열, 검은 걸레받이, 회색 바닥은 참조 장소와 일관된다. 정면 구도라 상체의 굽힘 각도는 단축되어 보인다. 불가능한 반사는 없다.",
        "entities": "한 명의 중년 동아시아계 남성이 보이며 짧게 정돈한 검은 머리, 체격, 검정 전투복과 붉은 완장, 검정 전투화가 박철진 참조에 부합한다. 얼굴은 아래로 가려져 정확한 얼굴 윤곽 비교가 어렵고 허리 장비도 상체에 상당 부분 가려진다. 귀 옆 휴대전화와 벽에 박힌 칼은 존재하지만 전화기 뒷면은 거의 드러나지 않는다. 게시물과 서류의 글씨는 읽히지 않는다. 중립적 실내 조명이며 밤을 직접 확인할 외부 풍경은 없다.",
        "hard_violations": [],
        "physics": "양발이 바닥에 완전히 닿아 있고 다리를 조금 벌린 상태에서 골반을 접어 상체를 숙인다. 깊은 인사가 가능한 자세이며 떠 있는 신체는 없다. 오른손이 전화기를 잡아 귀 옆에 유지하고, 칼은 벽에 박힌 부분으로 지지된다. 반대 팔은 몸 뒤로 가려져 있다. 손의 떨림은 명확히 판별되지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "측면 전신 와이드숏으로 약 90도의 허리 굽힘과 숙인 고개, 귀에 댄 휴대전화의 뒷면을 가장 명확하게 보여준다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "전신 통화와 깊은 인사는 충실하지만, 정면 단축으로 허리의 90도 굽힘이 덜 명확하고 휴대전화도 주로 얇은 모서리만 보인다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "몸통은 화면 오른쪽으로 거의 수평하게 숙여져 있고 얼굴과 시선은 발 앞쪽 바닥을 향한다. 오른손의 휴대전화는 귀에 밀착되어 있으며 화면은 머리 쪽, 뒷면과 카메라 부분은 관객 쪽이다. 칼은 손잡이가 왼쪽 위로 돌출되고 칼끝이 오른쪽 아래 벽 안으로 들어가 있어 사람을 겨누는 상태가 아니다.",
        "built_space": "뒤쪽 목재 책상 1개, 그 아래 일부 보이는 바퀴 의자 1개, 오른쪽 게시판 1개, 왼쪽의 높은 금속 수납장 1개와 뒤쪽에 일부 드러난 수납장 1개가 보인다. 오른쪽 돌출 벽에는 칼 1개가 박혀 있다. 인물은 책상 앞의 빈 바닥에 서 있으며 가구와 겹쳐 들어가지 않는다. 흰색·회색 투톤 벽, 균열, 검은 걸레받이와 회색 바닥이 이전 장면과 잘 이어진다. 반사는 없다.",
        "entities": "인물은 한 명뿐이며 짧은 검은 머리의 중년 동아시아계 남성으로 박철진의 외형과 부합한다. 숙인 얼굴 때문에 세부 얼굴 일치 여부는 제한적으로 보인다. 바랜 검정 전투복, 붉은 완장과 검정 표식, 허리 파우치, 검정 전투화가 참조와 맞는다. 검정 휴대전화와 회색 손잡이의 벽에 박힌 칼도 확인된다. 읽을 수 있는 글자는 없다. 중립적인 사무실 조명이며 밤이라는 외부 단서는 확실하지 않다.",
        "hard_violations": [],
        "physics": "두 전투화가 바닥에 닿아 체중을 지탱하고, 약간 굽힌 무릎과 뒤로 둔 골반이 앞으로 숙인 상체를 받친다. 전화기는 오른손으로 쥐고 귀에 대고 있으며 반대 손은 등 뒤에 있다. 칼은 벽에 박힌 날 끝으로 지지된다. 부유하거나 지지되지 않는 물체는 없다. 손의 떨림은 정지 화면에서 뚜렷하게 확인되지 않는다."
       },
       {
        "label": "A",
        "direction": "인물은 카메라 방향으로 몸을 숙이되 얼굴과 시선은 아래 바닥으로 향한다. 오른손은 전화기를 오른쪽 귀에 대고 있으며 관객에게는 주로 기기의 얇은 가장자리가 보이고 화면 내용은 보이지 않는다. 칼의 손잡이는 왼쪽 위를 향하고 날 끝은 오른쪽 아래 벽에 박혀 있다.",
        "built_space": "뒤쪽 왼편에 목재 책상 1개와 그 아래 일부 보이는 바퀴 의자 1개, 오른쪽 뒤 벽에 게시판 1개, 인물 뒤에 수납장 윗부분으로 보이는 구조가 있다. 오른쪽 돌출 벽의 칼은 1개다. 박철진은 책상 앞 빈 공간에서 서서 인사한다. 투톤 벽과 균열, 검은 걸레받이, 회색 바닥은 참조 장소와 일관된다. 정면 구도라 상체의 굽힘 각도는 단축되어 보인다. 불가능한 반사는 없다.",
        "entities": "한 명의 중년 동아시아계 남성이 보이며 짧게 정돈한 검은 머리, 체격, 검정 전투복과 붉은 완장, 검정 전투화가 박철진 참조에 부합한다. 얼굴은 아래로 가려져 정확한 얼굴 윤곽 비교가 어렵고 허리 장비도 상체에 상당 부분 가려진다. 귀 옆 휴대전화와 벽에 박힌 칼은 존재하지만 전화기 뒷면은 거의 드러나지 않는다. 게시물과 서류의 글씨는 읽히지 않는다. 중립적 실내 조명이며 밤을 직접 확인할 외부 풍경은 없다.",
        "hard_violations": [],
        "physics": "양발이 바닥에 완전히 닿아 있고 다리를 조금 벌린 상태에서 골반을 접어 상체를 숙인다. 깊은 인사가 가능한 자세이며 떠 있는 신체는 없다. 오른손이 전화기를 잡아 귀 옆에 유지하고, 칼은 벽에 박힌 부분으로 지지된다. 반대 팔은 몸 뒤로 가려져 있다. 손의 떨림은 명확히 판별되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학: 인물의 허리 뒤쪽에 팔과 연결되지 않은 기형적인 작은 손가락들이 떠 있는 형태로 렌더링됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1125
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "90도로 허리를 굽히고 비굴하게 통화하는 핵심 행동을 완벽히 연출했으며, 공간과 소품의 물리적 일관성도 매우 훌륭합니다."
   },
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "허리를 90도로 굽혀야 하는 지시를 따르지 않았으며, 벽을 향해 어색하게 서 있고 신체 렌더링 오류(허리 뒤쪽 손)가 심각합니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학: 인물의 허리 뒤쪽에 팔과 연결되지 않은 기형적인 작은 손가락들이 떠 있는 형태로 렌더링됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh6_sel.png",
    "asset_id": "3258ca24-c045-449f-813c-22c2f441d863",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-8b0e-767d-9eee-ff2ac7d9fa12",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S18sh6"
  }
 },
 "S43sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:00:51.454860+00:00",
  "fingerprint": "e158abc2c6b8de04ec84482b0fa7d468fab45f5d4c1311121d61942e1ab72a8d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S43sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S43sh4_sel.png",
  "source_sha256": "e24c3e2341d4e1dc664ca9586a0b90b6ebde877c9a8bf754147bd9e1401f48e1",
  "file": "S43sh4_cine.png",
  "staged_sha256": "70521446c9dbb0074886bc57818fa90ca0e17e4d1ab9034e8c5ebaf19bc4fa12",
  "latency_ms": 21626
 },
 "S43sh9::signage": {
  "fp": "b893f85bc7bfe54d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S43sh9": {
  "input_fingerprint": "f285a1072cc7ca6f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 야욕에 찬 핏발 선 눈으로 사악한 미소를 띤 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Inside the militia commander's office, near his subordinate after the phone call, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's neutral ambient treatment, keeping the eyes and forming smile legible through controlled tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the militia commander's office surfaces and nighttime interior lighting visible in the reference. Exclude the minister's sofa and furnishings from the other end of the telephone conversation.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He has ended the telephone call and remains in his office with an ambitious smile.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 야욕에 찬 핏발 선 눈으로 사악한 미소를 띤 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Inside the militia commander's office, near his subordinate after the phone call, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's neutral ambient treatment, keeping the eyes and forming smile legible through controlled tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the militia commander's office surfaces and nighttime interior lighting visible in the reference. Exclude the minister's sofa and furnishings from the other end of the telephone conversation.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He has ended the telephone call and remains in his office with an ambitious smile.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 야욕에 찬 핏발 선 눈으로 사악한 미소를 띤 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Inside the militia commander's office, near his subordinate after the phone call, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's neutral ambient treatment, keeping the eyes and forming smile legible through controlled tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the militia commander's office surfaces and nighttime interior lighting visible in the reference. Exclude the minister's sofa and furnishings from the other end of the telephone conversation.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He has ended the telephone call and remains in his office with an ambitious smile.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진이 화면 우측 전경에 있는 인물을 향해 시선을 고정하고 있음.",
    "built_space": "배경에 책상과 화이트보드가 보이며, 투톤으로 칠해진 벽면이 존재함.",
    "entities": "박철진(핏발 선 눈, 사악한 미소), 책상, 화이트보드, 샷 텍스트에 없는 인물의 어깨/뒷머리.",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 명시되지 않은 추가 인물(우측 전경) 등장",
     "[gemini-pro] 칼이 벽면에 온전히 꽂히지 않고 허공에 뜬 것처럼 어색하게 배치됨",
     "[gpt-high] 박철진 외에는 어떤 인물의 신체 일부도 허용되지 않는데 오른쪽 전경에 두 번째 남성의 머리와 어깨를 추가했다."
    ],
    "physics": "우측 전경의 인물 뒤로 칼이 허공에 뜬 것처럼 지지점이 불분명하게 묘사됨."
   },
   {
    "label": "B",
    "direction": "박철진이 화면 밖 우측을 향해 시선을 던지고 있음.",
    "built_space": "좌측 뒤편에 화이트보드가 있고, 우측에는 상단이 희고 하단이 짙은 회색인 벽면이 있음.",
    "entities": "박철진(핏발 선 눈, 사악한 미소, 정확한 의상), 벽에 박힌 칼, 화이트보드.",
    "hard_violations": [],
    "physics": "박철진은 안정적으로 서 있으며, 칼은 벽면 표면에 명확하게 꽂혀 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍을 정확히 지켰으며, 요구된 인물 단 한 명만 묘사하고 핏발 선 눈과 미소를 잘 표현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 명시되지 않은 타 인물의 뒷모습이 프레임에 포함되어 지시사항을 정면으로 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 화면 우측 전경에 있는 인물을 향해 시선을 고정하고 있음.",
        "built_space": "배경에 책상과 화이트보드가 보이며, 투톤으로 칠해진 벽면이 존재함.",
        "entities": "박철진(핏발 선 눈, 사악한 미소), 책상, 화이트보드, 샷 텍스트에 없는 인물의 어깨/뒷머리.",
        "hard_violations": [
         "샷 텍스트에 명시되지 않은 추가 인물(우측 전경) 등장",
         "칼이 벽면에 온전히 꽂히지 않고 허공에 뜬 것처럼 어색하게 배치됨"
        ],
        "physics": "우측 전경의 인물 뒤로 칼이 허공에 뜬 것처럼 지지점이 불분명하게 묘사됨."
       },
       {
        "label": "B",
        "direction": "박철진이 화면 밖 우측을 향해 시선을 던지고 있음.",
        "built_space": "좌측 뒤편에 화이트보드가 있고, 우측에는 상단이 희고 하단이 짙은 회색인 벽면이 있음.",
        "entities": "박철진(핏발 선 눈, 사악한 미소, 정확한 의상), 벽에 박힌 칼, 화이트보드.",
        "hard_violations": [],
        "physics": "박철진은 안정적으로 서 있으며, 칼은 벽면 표면에 명확하게 꽂혀 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍을 정확히 지켰으며, 요구된 인물 단 한 명만 묘사하고 핏발 선 눈과 미소를 잘 표현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 명시되지 않은 타 인물의 뒷모습이 프레임에 포함되어 지시사항을 정면으로 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 화면 우측 전경에 있는 인물을 향해 시선을 고정하고 있음.",
        "built_space": "배경에 책상과 화이트보드가 보이며, 투톤으로 칠해진 벽면이 존재함.",
        "entities": "박철진(핏발 선 눈, 사악한 미소), 책상, 화이트보드, 샷 텍스트에 없는 인물의 어깨/뒷머리.",
        "hard_violations": [
         "샷 텍스트에 명시되지 않은 추가 인물(우측 전경) 등장",
         "칼이 벽면에 온전히 꽂히지 않고 허공에 뜬 것처럼 어색하게 배치됨"
        ],
        "physics": "우측 전경의 인물 뒤로 칼이 허공에 뜬 것처럼 지지점이 불분명하게 묘사됨."
       },
       {
        "label": "B",
        "direction": "박철진이 화면 밖 우측을 향해 시선을 던지고 있음.",
        "built_space": "좌측 뒤편에 화이트보드가 있고, 우측에는 상단이 희고 하단이 짙은 회색인 벽면이 있음.",
        "entities": "박철진(핏발 선 눈, 사악한 미소, 정확한 의상), 벽에 박힌 칼, 화이트보드.",
        "hard_violations": [],
        "physics": "박철진은 안정적으로 서 있으며, 칼은 벽면 표면에 명확하게 꽂혀 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "박철진만의 얼굴 클로즈업으로 충혈된 눈과 야욕 어린 미소를 담고 사무실의 벽·칼·중성 조명을 유지해 가장 충실하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 두 번째 인물을 오른쪽 전경에 추가했으며, 얼굴 클로즈업을 상반신 중심의 대면 구도로 바꿨다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 눈은 렌즈보다 화면 왼쪽의 화면 밖 지점을 향한다. 시선 대상은 보이지 않으며, shot text는 특정 대상을 지정하지 않는다. 오른쪽 배경의 칼은 손잡이가 왼쪽 위, 칼끝이 오른쪽 아래의 벽을 향해 이전 장면과 같은 방향이다.",
        "built_space": "갈라지고 닳은 미색 상부·회색 하부 벽, 왼쪽 게시판 하나, 오른쪽 벽에 꽂힌 칼 하나가 보인다. 박철진의 얼굴과 어깨가 전경을 차지하고 벽의 물건들은 뒤에 놓인다. 책상과 바닥은 클로즈업 밖이며, 중복 설비나 반사는 없다. 배경의 상대적 배치는 변경된 카메라 각도로 설명 가능하다.",
        "entities": "인물은 박철진 한 명이다. 한국인 중년 남성으로 설정된 참조와 외관이 부합하며, 짧게 넘긴 검은 머리와 얼굴 윤곽, 낡은 짙은 회색 군복, 가장자리에 보이는 붉은 완장이 이어진다. 눈은 정상적인 홍채와 동공을 유지하면서 충혈되어 있고 입가에는 미소가 형성되어 있다. 벽의 칼과 게시판도 유지되며, 읽을 수 있는 글자나 통화 상대의 가구는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 있고 표정과 자세에 해부학적 불가능성은 없다. 하체 지지는 프레임 밖이므로 판단 대상이 아니다. 칼은 칼끝이 벽에 박혀 지지되는 것으로 보인다. 공중에 떠 있는 물체나 손 없이 들린 소품은 없다."
       },
       {
        "label": "B",
        "direction": "박철진은 화면 오른쪽 전경에 있는 다른 남성의 얼굴을 똑바로 바라보며 웃고, 그 남성도 박철진 쪽을 향한다. 시선 관계 자체는 맞지만 대상 인물의 등장은 이 샷의 허용 인물 조건에 어긋난다. 칼은 손잡이가 왼쪽 위, 칼끝이 오른쪽 아래 벽을 향한다.",
        "built_space": "미색·회색의 낡은 벽, 뒤쪽 게시판 하나, 오른쪽 벽의 칼 하나, 왼쪽 뒤의 책상 하나와 그 뒤 의자 등받이 하나가 보인다. 책상 위에는 검은 전화기와 종이류가 있다. 박철진은 책상 앞 벽 근처에 있고 다른 남성의 머리와 어깨가 오른쪽 전경을 점유한다. 장소 재질과 배치는 대체로 이어지지만 상반신과 상대 인물을 포함하도록 구도가 넓어졌다.",
        "entities": "박철진의 중년 남성 외관, 검은 머리, 짙은 군복과 붉은 완장은 참조에 대체로 부합한다. 충혈된 눈과 치아가 드러나는 음흉한 미소가 보인다. 그러나 오른쪽에 별개의 검은 머리 남성의 귀·머리·어깨가 추가되어 박철진만 허용한 조건을 위반한다. 칼은 벽에 남아 있고 전화기는 책상 위에 있으며, 읽히는 글자는 없다.",
        "hard_violations": [
         "박철진 외에는 어떤 인물의 신체 일부도 허용되지 않는데 오른쪽 전경에 두 번째 남성의 머리와 어깨를 추가했다."
        ],
        "physics": "박철진의 몸은 앞으로 조금 기울어 있지만 머리·목·몸통 연결이 자연스럽고, 보이는 자세에 물리적 불가능성은 없다. 두 인물의 하체는 프레임 밖이다. 칼은 벽에 박힌 끝으로 지지되고 전화기와 종이류는 책상 표면에 놓여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "박철진만의 얼굴 클로즈업으로 충혈된 눈과 야욕 어린 미소를 담고 사무실의 벽·칼·중성 조명을 유지해 가장 충실하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 두 번째 인물을 오른쪽 전경에 추가했으며, 얼굴 클로즈업을 상반신 중심의 대면 구도로 바꿨다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 눈은 렌즈보다 화면 왼쪽의 화면 밖 지점을 향한다. 시선 대상은 보이지 않으며, shot text는 특정 대상을 지정하지 않는다. 오른쪽 배경의 칼은 손잡이가 왼쪽 위, 칼끝이 오른쪽 아래의 벽을 향해 이전 장면과 같은 방향이다.",
        "built_space": "갈라지고 닳은 미색 상부·회색 하부 벽, 왼쪽 게시판 하나, 오른쪽 벽에 꽂힌 칼 하나가 보인다. 박철진의 얼굴과 어깨가 전경을 차지하고 벽의 물건들은 뒤에 놓인다. 책상과 바닥은 클로즈업 밖이며, 중복 설비나 반사는 없다. 배경의 상대적 배치는 변경된 카메라 각도로 설명 가능하다.",
        "entities": "인물은 박철진 한 명이다. 한국인 중년 남성으로 설정된 참조와 외관이 부합하며, 짧게 넘긴 검은 머리와 얼굴 윤곽, 낡은 짙은 회색 군복, 가장자리에 보이는 붉은 완장이 이어진다. 눈은 정상적인 홍채와 동공을 유지하면서 충혈되어 있고 입가에는 미소가 형성되어 있다. 벽의 칼과 게시판도 유지되며, 읽을 수 있는 글자나 통화 상대의 가구는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 있고 표정과 자세에 해부학적 불가능성은 없다. 하체 지지는 프레임 밖이므로 판단 대상이 아니다. 칼은 칼끝이 벽에 박혀 지지되는 것으로 보인다. 공중에 떠 있는 물체나 손 없이 들린 소품은 없다."
       },
       {
        "label": "A",
        "direction": "박철진은 화면 오른쪽 전경에 있는 다른 남성의 얼굴을 똑바로 바라보며 웃고, 그 남성도 박철진 쪽을 향한다. 시선 관계 자체는 맞지만 대상 인물의 등장은 이 샷의 허용 인물 조건에 어긋난다. 칼은 손잡이가 왼쪽 위, 칼끝이 오른쪽 아래 벽을 향한다.",
        "built_space": "미색·회색의 낡은 벽, 뒤쪽 게시판 하나, 오른쪽 벽의 칼 하나, 왼쪽 뒤의 책상 하나와 그 뒤 의자 등받이 하나가 보인다. 책상 위에는 검은 전화기와 종이류가 있다. 박철진은 책상 앞 벽 근처에 있고 다른 남성의 머리와 어깨가 오른쪽 전경을 점유한다. 장소 재질과 배치는 대체로 이어지지만 상반신과 상대 인물을 포함하도록 구도가 넓어졌다.",
        "entities": "박철진의 중년 남성 외관, 검은 머리, 짙은 군복과 붉은 완장은 참조에 대체로 부합한다. 충혈된 눈과 치아가 드러나는 음흉한 미소가 보인다. 그러나 오른쪽에 별개의 검은 머리 남성의 귀·머리·어깨가 추가되어 박철진만 허용한 조건을 위반한다. 칼은 벽에 남아 있고 전화기는 책상 위에 있으며, 읽히는 글자는 없다.",
        "hard_violations": [
         "박철진 외에는 어떤 인물의 신체 일부도 허용되지 않는데 오른쪽 전경에 두 번째 남성의 머리와 어깨를 추가했다."
        ],
        "physics": "박철진의 몸은 앞으로 조금 기울어 있지만 머리·목·몸통 연결이 자연스럽고, 보이는 자세에 물리적 불가능성은 없다. 두 인물의 하체는 프레임 밖이다. 칼은 벽에 박힌 끝으로 지지되고 전화기와 종이류는 책상 표면에 놓여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.651,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.401,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 샷 텍스트에 명시되지 않은 추가 인물(우측 전경) 등장",
     "[gemini-pro] 칼이 벽면에 온전히 꽂히지 않고 허공에 뜬 것처럼 어색하게 배치됨",
     "[gpt-high] 박철진 외에는 어떤 인물의 신체 일부도 허용되지 않는데 오른쪽 전경에 두 번째 남성의 머리와 어깨를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 401
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "클로즈업 프레이밍을 정확히 지켰으며, 요구된 인물 단 한 명만 묘사하고 핏발 선 눈과 미소를 잘 표현함."
   },
   {
    "label": "A",
    "score": 401,
    "verdict_ko": "샷 텍스트에 명시되지 않은 타 인물의 뒷모습이 프레임에 포함되어 지시사항을 정면으로 위반함.  ★위반: [gemini-pro] 샷 텍스트에 명시되지 않은 추가 인물(우측 전경) 등장 / [gemini-pro] 칼이 벽면에 온전히 꽂히지 않고 허공에 뜬 것처럼 어색하게 배치됨 / [gpt-high] 박철진 외에는 어떤 인물의 신체 일부도 허용되지 않는데 오른쪽 전경에 두 번째 남성의 머리와 어깨를 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh4_sel.png",
    "asset_id": "90356b08-c46e-4585-b3e0-b59c3c1bef68",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-8cbc-79de-bffe-83f686635731",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S43sh4"
  }
 },
 "S43sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:02:05.021801+00:00",
  "fingerprint": "1664335eeaeb2c64d2f95078ba585cc3d05bd30bf8d6b597e84907a9865d67f6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S43sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S43sh9_sel.png",
  "source_sha256": "ffca10ed6d40a97f69016fc8fc3acca6a4793c2ad5d274c6377069f47416d413",
  "file": "S43sh9_cine.png",
  "staged_sha256": "f6405986ff284a0b64ff3fc70f1c937075a4ec61075be45f4e02bcadf784b623",
  "latency_ms": 9171
 },
 "S44sh4::signage": {
  "fp": "94498938fa0032fc",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S44sh4": {
  "input_fingerprint": "658978ed3c0a8c9c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 어둠 속에서 먼지를 뒤집어쓴 채 서 있는 아담한 소형 캠핑카의 실루엣 전경.\n\nLOCATION (lock): In the dark vehicle-storage bay of a warehouse beside an abandoned factory, behind its raised shutter. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 소형 캠핑카 (Stationary inside the dark warehouse and covered in dust; its rear door has not yet been opened) — Seen diagonally with a side and the rear end readable, preserving the route toward the rear door; used as Primary reveal subject, contained within surrounding negative space rather than enlarged to fill the image; 창고 셔터 입구 (Raised to admit the group and camera) — Only a peripheral portion of the opening remains visible from just inside the entrance; used as Marks the completed threshold crossing and anchors the reveal to the approach path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse darkness, using only enough ambient tonal separation to read the camper's silhouette and dust-covered surfaces without inventing a beam or fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter has been raised, revealing a small camper in the darkness. The scavenged bags still contain supplies, medicine and the collected cash.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 어둠 속에서 먼지를 뒤집어쓴 채 서 있는 아담한 소형 캠핑카의 실루엣 전경.\n\nLOCATION (lock): In the dark vehicle-storage bay of a warehouse beside an abandoned factory, behind its raised shutter. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 소형 캠핑카 (Stationary inside the dark warehouse and covered in dust; its rear door has not yet been opened) — Seen diagonally with a side and the rear end readable, preserving the route toward the rear door; used as Primary reveal subject, contained within surrounding negative space rather than enlarged to fill the image; 창고 셔터 입구 (Raised to admit the group and camera) — Only a peripheral portion of the opening remains visible from just inside the entrance; used as Marks the completed threshold crossing and anchors the reveal to the approach path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse darkness, using only enough ambient tonal separation to read the camper's silhouette and dust-covered surfaces without inventing a beam or fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter has been raised, revealing a small camper in the darkness. The scavenged bags still contain supplies, medicine and the collected cash.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 어둠 속에서 먼지를 뒤집어쓴 채 서 있는 아담한 소형 캠핑카의 실루엣 전경.\n\nLOCATION (lock): In the dark vehicle-storage bay of a warehouse beside an abandoned factory, behind its raised shutter. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 소형 캠핑카 (Stationary inside the dark warehouse and covered in dust; its rear door has not yet been opened) — Seen diagonally with a side and the rear end readable, preserving the route toward the rear door; used as Primary reveal subject, contained within surrounding negative space rather than enlarged to fill the image; 창고 셔터 입구 (Raised to admit the group and camera) — Only a peripheral portion of the opening remains visible from just inside the entrance; used as Marks the completed threshold crossing and anchors the reveal to the approach path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse darkness, using only enough ambient tonal separation to read the camper's silhouette and dust-covered surfaces without inventing a beam or fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter has been raised, revealing a small camper in the darkness. The scavenged bags still contain supplies, medicine and the collected cash.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 창고 내부에서 캠핑카의 측면과 후면을 대각선으로 향해 바라봄.",
    "built_space": "어두운 창고 내부, 왼쪽 가장자리에 열린 셔터의 일부가 보이며 바닥에 가방들이 놓여 있음.",
    "entities": "먼지를 뒤집어쓴 소형 캠핑카, 바닥의 가방 묶음. 지시대로 인물 없음.",
    "hard_violations": [
     "[gpt-high] 장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
    ],
    "physics": "캠핑카와 가방들이 콘크리트 바닥에 자연스럽게 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "카메라가 창고 외부(또는 경계)에서 내부의 캠핑카를 정면으로 바라봄.",
    "built_space": "화면 테두리를 감싸고 있는 열린 셔터, 창고 내부 공간.",
    "entities": "먼지를 뒤집어쓴 소형 캠핑카. 프롬프트에 명시된 가방이 보이지 않음. 인물 없음.",
    "hard_violations": [
     "[gemini-pro] 카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함.",
     "[gpt-high] 입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
    ],
    "physics": "캠핑카가 바닥에 안정적으로 정차해 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "카메라 위치(내부에서 주변부 셔터만 보임), 캠핑카의 대각선 구도 및 먼지 덮인 질감, 바닥의 가방 등 모든 프롬프트 요구사항을 훌륭하게 충족함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 카메라 위치(내부)를 어기고 외부에서 정면으로 촬영했으며, 필수 요소인 바닥의 가방이 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 창고 내부에서 캠핑카의 측면과 후면을 대각선으로 향해 바라봄.",
        "built_space": "어두운 창고 내부, 왼쪽 가장자리에 열린 셔터의 일부가 보이며 바닥에 가방들이 놓여 있음.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카, 바닥의 가방 묶음. 지시대로 인물 없음.",
        "hard_violations": [],
        "physics": "캠핑카와 가방들이 콘크리트 바닥에 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "카메라가 창고 외부(또는 경계)에서 내부의 캠핑카를 정면으로 바라봄.",
        "built_space": "화면 테두리를 감싸고 있는 열린 셔터, 창고 내부 공간.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카. 프롬프트에 명시된 가방이 보이지 않음. 인물 없음.",
        "hard_violations": [
         "카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함."
        ],
        "physics": "캠핑카가 바닥에 안정적으로 정차해 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "카메라 위치(내부에서 주변부 셔터만 보임), 캠핑카의 대각선 구도 및 먼지 덮인 질감, 바닥의 가방 등 모든 프롬프트 요구사항을 훌륭하게 충족함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 카메라 위치(내부)를 어기고 외부에서 정면으로 촬영했으며, 필수 요소인 바닥의 가방이 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 창고 내부에서 캠핑카의 측면과 후면을 대각선으로 향해 바라봄.",
        "built_space": "어두운 창고 내부, 왼쪽 가장자리에 열린 셔터의 일부가 보이며 바닥에 가방들이 놓여 있음.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카, 바닥의 가방 묶음. 지시대로 인물 없음.",
        "hard_violations": [],
        "physics": "캠핑카와 가방들이 콘크리트 바닥에 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "카메라가 창고 외부(또는 경계)에서 내부의 캠핑카를 정면으로 바라봄.",
        "built_space": "화면 테두리를 감싸고 있는 열린 셔터, 창고 내부 공간.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카. 프롬프트에 명시된 가방이 보이지 않음. 인물 없음.",
        "hard_violations": [
         "카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함."
        ],
        "physics": "캠핑카가 바닥에 안정적으로 정차해 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "어둠 속 작은 캠핑카와 닫힌 후면 문은 잘 보이지만, 카메라가 창고 밖에서 입구 전체를 바라봐 입구를 이미 통과한 내부 시점을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "입구 안쪽 시점과 가장자리에 남은 셔터는 더 정확하지만, 지시되지 않은 상자·수납장·작업대 등으로 공간을 채워 엄격한 장소 제한을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 화면 왼쪽 안쪽을 향하고, 후면은 오른쪽의 카메라 쪽으로 향한다. 측면과 닫힌 후면 문이 함께 보이며 문 앞으로 접근할 바닥도 비어 있다. 사람이나 시선, 조준 대상은 없다.",
        "built_space": "셔터 입구 하나의 양쪽 기둥과 상단 셔터가 모두 보인다. 전경 바닥에서 문턱 너머 차량 보관실을 바라보는 외부 시점으로, 입구 안쪽에서 개구부 일부만 보이라는 조건과 다르다. 캠핑카 한 대 주변에는 충분한 어두운 여백이 있다.",
        "entities": "먼지와 때가 덮인 소형 캠핑카 한 대, 닫힌 후면 출입문 하나, 올라간 셔터 하나가 보인다. 사람은 없고 읽을 수 있는 글자도 없다. 별도 광원이나 빛줄기 없이 어두운 차체가 구분된다. 물품 가방과 그 내용물은 보이지 않으며, 이 구도에서 반드시 보여야 하는 대상은 아니다.",
        "hard_violations": [
         "입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 콘크리트 바닥에 닿아 차체를 지지한다. 차량은 정지해 있으며 떠 있는 물체나 불가능한 자세는 없다. 셔터는 입구의 측면 레일과 상부 구조에 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 왼쪽 안쪽으로, 후면은 오른쪽의 카메라 방향으로 향한다. 측면과 후면이 동시에 읽히며 후면으로 향하는 바닥 통로가 남아 있다. 다만 후면의 넓은 닫힌 패널은 출입문인지 창 덮개인지 명확하지 않다. 사람이나 조준 대상은 없다.",
        "built_space": "카메라는 창고 내부에 있고, 올라간 셔터 입구 하나가 왼쪽 가장자리에 일부 보인다. 외부의 어두운 하늘도 그 개구부로 보인다. 내부에는 철골 지붕과 벽 외에 여러 벽면 창, 수납장, 상자 더미와 작업대가 추가되어 있다. 캠핑카 주변 여백은 있지만 A보다 차체가 조금 크고 배경이 복잡하다.",
        "entities": "먼지로 덮인 소형 캠핑카 한 대와 올라간 셔터 하나가 보이고, 사람과 읽을 수 있는 글자는 없다. 왼쪽 전경에는 여러 가방이 놓여 있으나 약품·현금 등 내용물은 확인할 수 없다. 가방 자체는 지문에 언급되지만, 다수의 상자와 수납장 및 작업대는 지시되지 않은 추가 물체다. 밤의 어둠과 약한 주변광은 유지된다.",
        "hard_violations": [
         "장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
        ],
        "physics": "캠핑카의 보이는 앞뒤 바퀴가 바닥에 닿아 차량을 지지한다. 가방은 바닥에 놓여 있고, 상자와 적치물도 바닥이나 가구 위에 지지되어 있다. 셔터는 측면 레일과 상부 구조에 연결되어 있으며, 공중에 무지지 상태로 떠 있는 대상은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "어둠 속 작은 캠핑카와 닫힌 후면 문은 잘 보이지만, 카메라가 창고 밖에서 입구 전체를 바라봐 입구를 이미 통과한 내부 시점을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "입구 안쪽 시점과 가장자리에 남은 셔터는 더 정확하지만, 지시되지 않은 상자·수납장·작업대 등으로 공간을 채워 엄격한 장소 제한을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 화면 왼쪽 안쪽을 향하고, 후면은 오른쪽의 카메라 쪽으로 향한다. 측면과 닫힌 후면 문이 함께 보이며 문 앞으로 접근할 바닥도 비어 있다. 사람이나 시선, 조준 대상은 없다.",
        "built_space": "셔터 입구 하나의 양쪽 기둥과 상단 셔터가 모두 보인다. 전경 바닥에서 문턱 너머 차량 보관실을 바라보는 외부 시점으로, 입구 안쪽에서 개구부 일부만 보이라는 조건과 다르다. 캠핑카 한 대 주변에는 충분한 어두운 여백이 있다.",
        "entities": "먼지와 때가 덮인 소형 캠핑카 한 대, 닫힌 후면 출입문 하나, 올라간 셔터 하나가 보인다. 사람은 없고 읽을 수 있는 글자도 없다. 별도 광원이나 빛줄기 없이 어두운 차체가 구분된다. 물품 가방과 그 내용물은 보이지 않으며, 이 구도에서 반드시 보여야 하는 대상은 아니다.",
        "hard_violations": [
         "입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 콘크리트 바닥에 닿아 차체를 지지한다. 차량은 정지해 있으며 떠 있는 물체나 불가능한 자세는 없다. 셔터는 입구의 측면 레일과 상부 구조에 지지되어 있다."
       },
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 왼쪽 안쪽으로, 후면은 오른쪽의 카메라 방향으로 향한다. 측면과 후면이 동시에 읽히며 후면으로 향하는 바닥 통로가 남아 있다. 다만 후면의 넓은 닫힌 패널은 출입문인지 창 덮개인지 명확하지 않다. 사람이나 조준 대상은 없다.",
        "built_space": "카메라는 창고 내부에 있고, 올라간 셔터 입구 하나가 왼쪽 가장자리에 일부 보인다. 외부의 어두운 하늘도 그 개구부로 보인다. 내부에는 철골 지붕과 벽 외에 여러 벽면 창, 수납장, 상자 더미와 작업대가 추가되어 있다. 캠핑카 주변 여백은 있지만 A보다 차체가 조금 크고 배경이 복잡하다.",
        "entities": "먼지로 덮인 소형 캠핑카 한 대와 올라간 셔터 하나가 보이고, 사람과 읽을 수 있는 글자는 없다. 왼쪽 전경에는 여러 가방이 놓여 있으나 약품·현금 등 내용물은 확인할 수 없다. 가방 자체는 지문에 언급되지만, 다수의 상자와 수납장 및 작업대는 지시되지 않은 추가 물체다. 밤의 어둠과 약한 주변광은 유지된다.",
        "hard_violations": [
         "장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
        ],
        "physics": "캠핑카의 보이는 앞뒤 바퀴가 바닥에 닿아 차량을 지지한다. 가방은 바닥에 놓여 있고, 상자와 적치물도 바닥이나 가구 위에 지지되어 있다. 셔터는 측면 레일과 상부 구조에 연결되어 있으며, 공중에 무지지 상태로 떠 있는 대상은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.125
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.875
   },
   "violations": {
    "B": [
     "[gemini-pro] 카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함.",
     "[gpt-high] 입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
    ],
    "A": [
     "[gpt-high] 장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 875
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "카메라 위치(내부에서 주변부 셔터만 보임), 캠핑카의 대각선 구도 및 먼지 덮인 질감, 바닥의 가방 등 모든 프롬프트 요구사항을 훌륭하게 충족함.  ★위반: [gpt-high] 장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 875,
    "verdict_ko": "지정된 카메라 위치(내부)를 어기고 외부에서 정면으로 촬영했으며, 필수 요소인 바닥의 가방이 누락됨.  ★위반: [gemini-pro] 카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함. / [gpt-high] 입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-8e5f-7394-a92a-66acc660edd1",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S44sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:33:51.057437+00:00",
  "fingerprint": "77ec558b343017c29b5c0545c688bbc51ed5c28fc6cd86c0f04cab329ee9e339",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S44sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S44sh4_sel.png",
  "source_sha256": "6564b3d5174018a84441ae5bd4988e0c21093aa89a54685f6163cbfd1eb0e165",
  "file": "S44sh4_cine.png",
  "staged_sha256": "b1faf371ad06f14e387b2ca93c408b78e8cd95dbaf9903c7ebadfdc22c49616b",
  "latency_ms": 15186
 },
 "S44sh7::signage": {
  "fp": "caad744326f137ae",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S44sh7": {
  "input_fingerprint": "ba3da61f5d7ec76d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 캠핑카 안으로 시선을 고정한 채 안도하는 표정으로 미소 짓고 있는 현우의 옆얼굴.\n\nLOCATION (lock): Inside the dark warehouse beside the abandoned factory, immediately outside the camper's open rear doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Camper rear doorway (Open after 현우 has opened the rear door) — Seen obliquely beside his profile, with the opening leading toward the interior; used as Provides looking room and a narrow spatial boundary beside his face; Camper bed (Present inside the camper) — Only a small portion is visible through the rear opening; used as Gives concrete context to his relief without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse's established darkness while keeping the small change in his expression legible through controlled tonal separation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the dusty camper exterior and the surrounding dark storage warehouse. Exclude any assumption that the camper door remains closed; the rear door is now open.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter remains raised, and the camper's rear door is now open. Its interior includes a bed and enough room for four. 현우: He stands at the open rear of the camper with his facial bruises and leg wound unchanged. The contact card remains concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 캠핑카 안으로 시선을 고정한 채 안도하는 표정으로 미소 짓고 있는 현우의 옆얼굴.\n\nLOCATION (lock): Inside the dark warehouse beside the abandoned factory, immediately outside the camper's open rear doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Camper rear doorway (Open after 현우 has opened the rear door) — Seen obliquely beside his profile, with the opening leading toward the interior; used as Provides looking room and a narrow spatial boundary beside his face; Camper bed (Present inside the camper) — Only a small portion is visible through the rear opening; used as Gives concrete context to his relief without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse's established darkness while keeping the small change in his expression legible through controlled tonal separation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the dusty camper exterior and the surrounding dark storage warehouse. Exclude any assumption that the camper door remains closed; the rear door is now open.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter remains raised, and the camper's rear door is now open. Its interior includes a bed and enough room for four. 현우: He stands at the open rear of the camper with his facial bruises and leg wound unchanged. The contact card remains concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 캠핑카 안으로 시선을 고정한 채 안도하는 표정으로 미소 짓고 있는 현우의 옆얼굴.\n\nLOCATION (lock): Inside the dark warehouse beside the abandoned factory, immediately outside the camper's open rear doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Camper rear doorway (Open after 현우 has opened the rear door) — Seen obliquely beside his profile, with the opening leading toward the interior; used as Provides looking room and a narrow spatial boundary beside his face; Camper bed (Present inside the camper) — Only a small portion is visible through the rear opening; used as Gives concrete context to his relief without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse's established darkness while keeping the small change in his expression legible through controlled tonal separation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the dusty camper exterior and the surrounding dark storage warehouse. Exclude any assumption that the camper door remains closed; the rear door is now open.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter remains raised, and the camper's rear door is now open. Its interior includes a bed and enough room for four. 현우: He stands at the open rear of the camper with his facial bruises and leg wound unchanged. The contact card remains concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 왼쪽의 열린 캠핑카 내부를 향해 고정되어 있음.",
    "built_space": "어두운 창고 내부. 왼쪽에 내부가 보이는 열린 문이 있으나, 인물 뒤편으로 캠핑카의 외벽과 손잡이가 달린 외부 문짝이 비정상적으로 중복 등장하여 공간이 붕괴됨.",
    "entities": "현우(참조 이미지와 일치, 얼굴 타박상, 미소, 회색 셔츠). 캠핑카(구조가 완전히 왜곡됨).",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 배경 구조 (캠핑카 측면과 외부 문이 비정상적으로 중복되고 형태가 왜곡됨)"
    ],
    "physics": "지면에 서 있으며 물리적 자세 오류는 없음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 왼쪽의 열린 캠핑카 내부를 향해 고정되어 있음.",
    "built_space": "어두운 창고 내부. 왼쪽에 열린 문과 내부 침대가 보이며, 인물 뒤로 열린 문짝의 안쪽 면과 캠핑카 측면, 트럭 운전석이 올바른 원근감으로 이어짐.",
    "entities": "현우(참조 이미지와 일치, 앳된 얼굴, 헝클어진 머리, 얼굴 타박상, 회색 셔츠). 캠핑카(문이 열린 상태, 내부 침대 확인됨).",
    "hard_violations": [
     "[gpt-high] 후면 출입구 바로 밖이라는 지정 위치를 측면 출입구 밖으로 바꾸었습니다. 출입구 옆으로 운전석과 긴 측면 외벽이 이어져 위치 변경이 드러납니다."
    ],
    "physics": "지면에 안정적으로 서 있으며, 부자연스러운 물리적 요소는 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물의 안도하는 표정과 구도를 정확히 구현했으며, 캠핑카의 열린 문과 내부 침대 등 공간 구조가 자연스럽게 표현되었습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물의 표정은 지시를 따랐으나, 배경의 캠핑카 외벽 구조가 물리적으로 불가능하게 왜곡되어 심각한 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 왼쪽의 열린 캠핑카 내부를 향해 고정되어 있음.",
        "built_space": "어두운 창고 내부. 왼쪽에 열린 문과 내부 침대가 보이며, 인물 뒤로 열린 문짝의 안쪽 면과 캠핑카 측면, 트럭 운전석이 올바른 원근감으로 이어짐.",
        "entities": "현우(참조 이미지와 일치, 앳된 얼굴, 헝클어진 머리, 얼굴 타박상, 회색 셔츠). 캠핑카(문이 열린 상태, 내부 침대 확인됨).",
        "hard_violations": [],
        "physics": "지면에 안정적으로 서 있으며, 부자연스러운 물리적 요소는 없음."
       },
       {
        "label": "A",
        "direction": "현우의 시선이 왼쪽의 열린 캠핑카 내부를 향해 고정되어 있음.",
        "built_space": "어두운 창고 내부. 왼쪽에 내부가 보이는 열린 문이 있으나, 인물 뒤편으로 캠핑카의 외벽과 손잡이가 달린 외부 문짝이 비정상적으로 중복 등장하여 공간이 붕괴됨.",
        "entities": "현우(참조 이미지와 일치, 얼굴 타박상, 미소, 회색 셔츠). 캠핑카(구조가 완전히 왜곡됨).",
        "hard_violations": [
         "물리적으로 불가능한 배경 구조 (캠핑카 측면과 외부 문이 비정상적으로 중복되고 형태가 왜곡됨)"
        ],
        "physics": "지면에 서 있으며 물리적 자세 오류는 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물의 안도하는 표정과 구도를 정확히 구현했으며, 캠핑카의 열린 문과 내부 침대 등 공간 구조가 자연스럽게 표현되었습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물의 표정은 지시를 따랐으나, 배경의 캠핑카 외벽 구조가 물리적으로 불가능하게 왜곡되어 심각한 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 왼쪽의 열린 캠핑카 내부를 향해 고정되어 있음.",
        "built_space": "어두운 창고 내부. 왼쪽에 열린 문과 내부 침대가 보이며, 인물 뒤로 열린 문짝의 안쪽 면과 캠핑카 측면, 트럭 운전석이 올바른 원근감으로 이어짐.",
        "entities": "현우(참조 이미지와 일치, 앳된 얼굴, 헝클어진 머리, 얼굴 타박상, 회색 셔츠). 캠핑카(문이 열린 상태, 내부 침대 확인됨).",
        "hard_violations": [],
        "physics": "지면에 안정적으로 서 있으며, 부자연스러운 물리적 요소는 없음."
       },
       {
        "label": "A",
        "direction": "현우의 시선이 왼쪽의 열린 캠핑카 내부를 향해 고정되어 있음.",
        "built_space": "어두운 창고 내부. 왼쪽에 내부가 보이는 열린 문이 있으나, 인물 뒤편으로 캠핑카의 외벽과 손잡이가 달린 외부 문짝이 비정상적으로 중복 등장하여 공간이 붕괴됨.",
        "entities": "현우(참조 이미지와 일치, 얼굴 타박상, 미소, 회색 셔츠). 캠핑카(구조가 완전히 왜곡됨).",
        "hard_violations": [
         "물리적으로 불가능한 배경 구조 (캠핑카 측면과 외부 문이 비정상적으로 중복되고 형태가 왜곡됨)"
        ],
        "physics": "지면에 서 있으며 물리적 자세 오류는 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "실내를 바라보며 웃는 옆얼굴은 맞지만, 출입구가 후면이 아닌 운전석과 이어지는 측면에 놓였고 얼굴 클로즈업도 상대적으로 느슨합니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "더 밀착한 옆얼굴 클로즈업과 실내를 향한 시선, 안도의 미소, 얼굴 상처가 요구에 부합하지만 침대와 출입구의 화면 비중은 다소 큽니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽의 열린 출입구 안쪽을 향합니다. 카메라를 보지 않으며, 미소를 지은 채 캠핑카 내부를 바라보는 방향은 맞습니다.",
        "built_space": "출입구 하나, 오른쪽으로 열린 문짝 하나, 내부 침대 하나가 보입니다. 왼쪽 큰 창과 오른쪽 운전석 및 상부 돌출부가 출입구와 같은 긴 측면을 따라 이어져, 현우는 후면이 아니라 측면 출입문 밖에 서 있는 것으로 읽힙니다. 먼지 낀 외장과 어두운 주변은 유지되지만 창고 셔터는 프레임 밖입니다.",
        "entities": "젊은 동아시아계 남성 한 명이며, 검은 헝클어진 머리와 회색 셔츠는 현우 참조와 대체로 맞습니다. 뺨에 어두운 얼룩은 있으나 유지되어야 할 얼굴 타박상은 뚜렷하지 않습니다. 침대와 낡은 캠핑카는 식별됩니다. 다리 상처와 신발 속 카드는 프레임 밖이라 확인할 수 없고, 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "후면 출입구 바로 밖이라는 지정 위치를 측면 출입구 밖으로 바꾸었습니다. 출입구 옆으로 운전석과 긴 측면 외벽이 이어져 위치 변경이 드러납니다."
        ],
        "physics": "보이는 머리와 상체는 자연스럽게 연결되어 있고 서 있는 자세로 읽힙니다. 발은 클로즈업 밖이므로 지면 접촉은 확인되지 않지만 부유하는 모습은 아닙니다. 문짝은 경첩으로 차체에 연결되고 침구는 침대 받침 위에 놓여 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 코와 눈동자는 화면 왼쪽 출입구 너머의 실내를 향합니다. 눈높이의 내부 공간을 바라보면서 입꼬리를 올리고 있어, 열린 캠핑카 안에 시선을 고정한 안도의 순간으로 읽힙니다.",
        "built_space": "왼쪽에 열린 출입구 하나, 그 오른쪽에 경첩으로 연결된 문짝 하나, 안쪽에 침대 하나와 어두운 창 하나가 보입니다. 현우는 문 밖 오른쪽에 서 있고 문틀이 얼굴 옆 공간의 경계를 만듭니다. 비스듬한 차체 일부만 보여 후면 위치를 확정하기는 어렵지만, A처럼 운전석까지 펼쳐진 측면 구도는 아닙니다. 매트리스 전면과 받침까지 보여 침대가 요구된 작은 단편보다는 크게 드러납니다.",
        "entities": "검은 헝클어진 머리와 앳된 얼굴의 동아시아계 남성 한 명으로, 현우 참조의 인상과 회색 셔츠를 대체로 유지합니다. 코와 광대 주변의 멍 및 긁힌 자국이 명확합니다. 낡은 캠핑카와 실물 매트리스가 보이며 추가 인물이나 판독 가능한 문자는 없습니다. 다리와 신발은 프레임 밖이므로 상처와 카드의 상태는 확인 대상에서 제외됩니다.",
        "hard_violations": [],
        "physics": "목과 어깨가 자연스럽게 연결된 서 있는 상체이며, 발이 잘렸다는 것 외에 지지 없는 부유의 징후는 없습니다. 열린 문은 보이는 경첩에 매달려 있고 매트리스는 목재 받침 위에 안정적으로 놓여 있습니다. 손에 든 물건이나 지지 없는 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "실내를 바라보며 웃는 옆얼굴은 맞지만, 출입구가 후면이 아닌 운전석과 이어지는 측면에 놓였고 얼굴 클로즈업도 상대적으로 느슨합니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "더 밀착한 옆얼굴 클로즈업과 실내를 향한 시선, 안도의 미소, 얼굴 상처가 요구에 부합하지만 침대와 출입구의 화면 비중은 다소 큽니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽의 열린 출입구 안쪽을 향합니다. 카메라를 보지 않으며, 미소를 지은 채 캠핑카 내부를 바라보는 방향은 맞습니다.",
        "built_space": "출입구 하나, 오른쪽으로 열린 문짝 하나, 내부 침대 하나가 보입니다. 왼쪽 큰 창과 오른쪽 운전석 및 상부 돌출부가 출입구와 같은 긴 측면을 따라 이어져, 현우는 후면이 아니라 측면 출입문 밖에 서 있는 것으로 읽힙니다. 먼지 낀 외장과 어두운 주변은 유지되지만 창고 셔터는 프레임 밖입니다.",
        "entities": "젊은 동아시아계 남성 한 명이며, 검은 헝클어진 머리와 회색 셔츠는 현우 참조와 대체로 맞습니다. 뺨에 어두운 얼룩은 있으나 유지되어야 할 얼굴 타박상은 뚜렷하지 않습니다. 침대와 낡은 캠핑카는 식별됩니다. 다리 상처와 신발 속 카드는 프레임 밖이라 확인할 수 없고, 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "후면 출입구 바로 밖이라는 지정 위치를 측면 출입구 밖으로 바꾸었습니다. 출입구 옆으로 운전석과 긴 측면 외벽이 이어져 위치 변경이 드러납니다."
        ],
        "physics": "보이는 머리와 상체는 자연스럽게 연결되어 있고 서 있는 자세로 읽힙니다. 발은 클로즈업 밖이므로 지면 접촉은 확인되지 않지만 부유하는 모습은 아닙니다. 문짝은 경첩으로 차체에 연결되고 침구는 침대 받침 위에 놓여 있습니다."
       },
       {
        "label": "A",
        "direction": "현우의 코와 눈동자는 화면 왼쪽 출입구 너머의 실내를 향합니다. 눈높이의 내부 공간을 바라보면서 입꼬리를 올리고 있어, 열린 캠핑카 안에 시선을 고정한 안도의 순간으로 읽힙니다.",
        "built_space": "왼쪽에 열린 출입구 하나, 그 오른쪽에 경첩으로 연결된 문짝 하나, 안쪽에 침대 하나와 어두운 창 하나가 보입니다. 현우는 문 밖 오른쪽에 서 있고 문틀이 얼굴 옆 공간의 경계를 만듭니다. 비스듬한 차체 일부만 보여 후면 위치를 확정하기는 어렵지만, A처럼 운전석까지 펼쳐진 측면 구도는 아닙니다. 매트리스 전면과 받침까지 보여 침대가 요구된 작은 단편보다는 크게 드러납니다.",
        "entities": "검은 헝클어진 머리와 앳된 얼굴의 동아시아계 남성 한 명으로, 현우 참조의 인상과 회색 셔츠를 대체로 유지합니다. 코와 광대 주변의 멍 및 긁힌 자국이 명확합니다. 낡은 캠핑카와 실물 매트리스가 보이며 추가 인물이나 판독 가능한 문자는 없습니다. 다리와 신발은 프레임 밖이므로 상처와 카드의 상태는 확인 대상에서 제외됩니다.",
        "hard_violations": [],
        "physics": "목과 어깨가 자연스럽게 연결된 서 있는 상체이며, 발이 잘렸다는 것 외에 지지 없는 부유의 징후는 없습니다. 열린 문은 보이는 경첩에 매달려 있고 매트리스는 목재 받침 위에 안정적으로 놓여 있습니다. 손에 든 물건이나 지지 없는 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 배경 구조 (캠핑카 측면과 외부 문이 비정상적으로 중복되고 형태가 왜곡됨)"
    ],
    "B": [
     "[gpt-high] 후면 출입구 바로 밖이라는 지정 위치를 측면 출입구 밖으로 바꾸었습니다. 출입구 옆으로 운전석과 긴 측면 외벽이 이어져 위치 변경이 드러납니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "인물의 안도하는 표정과 구도를 정확히 구현했으며, 캠핑카의 열린 문과 내부 침대 등 공간 구조가 자연스럽게 표현되었습니다.  ★위반: [gpt-high] 후면 출입구 바로 밖이라는 지정 위치를 측면 출입구 밖으로 바꾸었습니다. 출입구 옆으로 운전석과 긴 측면 외벽이 이어져 위치 변경이 드러납니다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "인물의 표정은 지시를 따랐으나, 배경의 캠핑카 외벽 구조가 물리적으로 불가능하게 왜곡되어 심각한 오류가 발생했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 배경 구조 (캠핑카 측면과 외부 문이 비정상적으로 중복되고 형태가 왜곡됨)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S44sh4_sel.png",
    "asset_id": "c2695594-fdd0-410f-9f81-0a1a52a57d8f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-8ff8-7d1d-9e44-80b8e76d1286",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S44sh4"
  }
 },
 "S44sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:03:15.900402+00:00",
  "fingerprint": "92ec39e78c11135ffcd9751099f026c561a608d0910558fbbb319757781c6582",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S44sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S44sh7_sel.png",
  "source_sha256": "1ef0027b7e4dcb99b72f06fd28d252ef4b741119200262e5a094105abcf467e1",
  "file": "S44sh7_cine.png",
  "staged_sha256": "f55052833a3733ade1d13a764efec5ce325124e981aab88d942a7ed9ef2dda4b",
  "latency_ms": 11932
 },
 "S45sh1::signage": {
  "fp": "fdd9b2611d824981",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S45sh1": {
  "input_fingerprint": "4732513fd5f88956",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거센 비가 내리는 텅 빈 도로 위를 헤드라이트가 꺼진 채 달리는 낡은 캠핑카의 외관.\n\nLOCATION (lock): On a nearly deserted road at the city's outskirts at night, where the camper travels through heavy rain without headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Traveling through heavy rain with its headlights off) — Front and passenger-side quarter visible from above; used as Small moving subject surrounded by exposed road space; Outlying road (Empty within the frame during the downpour) — Extends diagonally past the camper in its direction of travel; used as Establishes isolation and the forward travel axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the rainy night dark and the headlights unlit, using restrained tonal separation to distinguish the moving camper from the road.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper travels through heavy nighttime rain with its headlights off and one small rear-compartment lamp lit. The scavenged luggage and supplies are stowed inside.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거센 비가 내리는 텅 빈 도로 위를 헤드라이트가 꺼진 채 달리는 낡은 캠핑카의 외관.\n\nLOCATION (lock): On a nearly deserted road at the city's outskirts at night, where the camper travels through heavy rain without headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Traveling through heavy rain with its headlights off) — Front and passenger-side quarter visible from above; used as Small moving subject surrounded by exposed road space; Outlying road (Empty within the frame during the downpour) — Extends diagonally past the camper in its direction of travel; used as Establishes isolation and the forward travel axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the rainy night dark and the headlights unlit, using restrained tonal separation to distinguish the moving camper from the road.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper travels through heavy nighttime rain with its headlights off and one small rear-compartment lamp lit. The scavenged luggage and supplies are stowed inside.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거센 비가 내리는 텅 빈 도로 위를 헤드라이트가 꺼진 채 달리는 낡은 캠핑카의 외관.\n\nLOCATION (lock): On a nearly deserted road at the city's outskirts at night, where the camper travels through heavy rain without headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Traveling through heavy rain with its headlights off) — Front and passenger-side quarter visible from above; used as Small moving subject surrounded by exposed road space; Outlying road (Empty within the frame during the downpour) — Extends diagonally past the camper in its direction of travel; used as Establishes isolation and the forward travel axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the rainy night dark and the headlights unlit, using restrained tonal separation to distinguish the moving camper from the road.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper travels through heavy nighttime rain with its headlights off and one small rear-compartment lamp lit. The scavenged luggage and supplies are stowed inside.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "캠핑카가 프레임을 가로지르는 대각선 도로를 따라 앞을 향해 이동 중임.",
    "built_space": "가드레일과 우측 하단 구조물이 있는 도로가 높은 곳에서 내려다보이는 부감 샷으로 구성됨.",
    "entities": "낡은 캠핑카의 정면과 우측(조수석) 측면이 보이며, 헤드라이트가 꺼진 상태로 내부 뒷면 램프만 켜져 있음.",
    "hard_violations": [],
    "physics": "차량 바퀴가 젖은 도로 노면에 자연스럽게 닿아 무게감을 지탱함."
   },
   {
    "label": "B",
    "direction": "캠핑카가 도로의 차선을 따라 전방으로 주행하고 있음.",
    "built_space": "지상 눈높이에서 굽은 도로와 가드레일, 멀리 배경의 도시 조명이 보임.",
    "entities": "캠핑카의 정면과 좌측(운전석) 측면이 크게 노출되었으며, 헤드라이트는 꺼져 있고 내부 조명이 켜짐.",
    "hard_violations": [
     "[gpt-high] 앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "캠핑카가 빗물이 고인 도로 위에 안정적으로 서서 이동함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요청된 '위에서 내려다본 조수석 측면 뷰'와 '넓은 여백 속 작은 피사체'라는 프레이밍 지침을 정확히 구현하여 높은 점수를 부여함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "'위에서 내려다본 구도'와 '조수석 측면 노출'이라는 명시적인 카메라 및 앵글 지시를 위반하고 피사체를 너무 크게 배치함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 프레임을 가로지르는 대각선 도로를 따라 앞을 향해 이동 중임.",
        "built_space": "가드레일과 우측 하단 구조물이 있는 도로가 높은 곳에서 내려다보이는 부감 샷으로 구성됨.",
        "entities": "낡은 캠핑카의 정면과 우측(조수석) 측면이 보이며, 헤드라이트가 꺼진 상태로 내부 뒷면 램프만 켜져 있음.",
        "hard_violations": [],
        "physics": "차량 바퀴가 젖은 도로 노면에 자연스럽게 닿아 무게감을 지탱함."
       },
       {
        "label": "B",
        "direction": "캠핑카가 도로의 차선을 따라 전방으로 주행하고 있음.",
        "built_space": "지상 눈높이에서 굽은 도로와 가드레일, 멀리 배경의 도시 조명이 보임.",
        "entities": "캠핑카의 정면과 좌측(운전석) 측면이 크게 노출되었으며, 헤드라이트는 꺼져 있고 내부 조명이 켜짐.",
        "hard_violations": [],
        "physics": "캠핑카가 빗물이 고인 도로 위에 안정적으로 서서 이동함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요청된 '위에서 내려다본 조수석 측면 뷰'와 '넓은 여백 속 작은 피사체'라는 프레이밍 지침을 정확히 구현하여 높은 점수를 부여함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "'위에서 내려다본 구도'와 '조수석 측면 노출'이라는 명시적인 카메라 및 앵글 지시를 위반하고 피사체를 너무 크게 배치함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 프레임을 가로지르는 대각선 도로를 따라 앞을 향해 이동 중임.",
        "built_space": "가드레일과 우측 하단 구조물이 있는 도로가 높은 곳에서 내려다보이는 부감 샷으로 구성됨.",
        "entities": "낡은 캠핑카의 정면과 우측(조수석) 측면이 보이며, 헤드라이트가 꺼진 상태로 내부 뒷면 램프만 켜져 있음.",
        "hard_violations": [],
        "physics": "차량 바퀴가 젖은 도로 노면에 자연스럽게 닿아 무게감을 지탱함."
       },
       {
        "label": "B",
        "direction": "캠핑카가 도로의 차선을 따라 전방으로 주행하고 있음.",
        "built_space": "지상 눈높이에서 굽은 도로와 가드레일, 멀리 배경의 도시 조명이 보임.",
        "entities": "캠핑카의 정면과 좌측(운전석) 측면이 크게 노출되었으며, 헤드라이트는 꺼져 있고 내부 조명이 켜짐.",
        "hard_violations": [],
        "physics": "캠핑카가 빗물이 고인 도로 위에 안정적으로 서서 이동함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "폭우와 소등 상태는 맞지만, 차량이 크고 카메라가 낮아 지정된 하이앵글 와이드 구도에서 벗어나며 번호판 숫자가 읽힌다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "높은 시점에서 작은 캠핑카와 넓은 대각선 도로를 담아 핵심 구도를 가장 충실히 구현했지만, 보이는 측면은 요구된 조수석 쪽과 반대다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 화면 왼쪽 아래를 향하고 도로는 오른쪽 뒤로 이어져 진행축이 일치한다. 전조등의 전방 광선은 없다. 사람의 시선은 식별되지 않는다. 앞부분과 함께 보이는 차체 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽에 해당해 조수석 측면 요구와 다르다.",
        "built_space": "노란 중앙선이 있는 2차로 도로, 양쪽 가장자리의 가드레일 두 줄, 여러 전신주와 전선, 수목이 보인다. 참조의 젖은 아스팔트와 금속 가드레일은 유지하지만 멀리 도시 건물과 다수의 불빛이 더 두드러진다. 차량은 화면 폭의 약 3분의 1을 차지하며 지붕이 조금만 보여, 도로에 둘러싸인 작은 피사체를 위에서 보는 구도와 거리가 있다.",
        "entities": "낡고 얼룩진 흰색 오버캡 캠핑카 한 대가 있으며 다른 차량이나 식별 가능한 사람은 없다. 헤드라이트는 꺼져 있다. 객실 창 안에 따뜻한 갓등 하나가 뚜렷하게 보이고 뒤쪽 창에도 약한 빛이 있다. 외부에 실린 짐은 없다. 밤과 강한 비는 구현되었으나 앞 번호판의 숫자열은 읽을 수 있어 무문자 조건에 어긋난다.",
        "hard_violations": [
         "앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 노면에 닿아 차체를 지지한다. 바퀴 주변의 물보라와 젖은 노면은 빗속 주행으로 가능한 모습이다. 실내 갓등은 창 아래에 놓인 것으로 보이며 공중에 떠 있다는 증거는 없다. 노면의 반사도 젖은 표면에서 가능한 범위다."
       },
       {
        "label": "B",
        "direction": "캠핑카는 화면 왼쪽 아래를 향하고, 도로 역시 그 방향으로 대각선으로 뻗어 진행축이 맞는다. 전조등은 빛을 쏘지 않는다. 사람의 시선은 식별되지 않는다. 다만 앞부분과 함께 노출된 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽이어서 조수석 측면 요구와 반대다.",
        "built_space": "노란 중앙선으로 나뉜 2차로 도로와 양쪽 가드레일 두 줄, 도로변 전신주 여러 개와 전선, 어두운 수목이 보인다. 참조의 주요 도로 재료와 시설은 대체로 유지된다. 오른쪽 가장자리는 참조보다 높고 교량 난간처럼 보이는 차이가 있다. 높은 외부 시점에서 지붕과 앞부분을 내려다보며, 차량 주변으로 넓은 빈 노면이 확보되어 지정된 와이드 구도에 잘 맞는다.",
        "entities": "낡은 흰색 오버캡 캠핑카 한 대만 보이며 다른 차량이나 식별 가능한 사람은 없다. 앞 헤드라이트는 꺼져 있고 후방 객실의 작은 창에 광원 하나가 켜져 있다. 옆의 큰 창에는 약한 실내 빛만 보여 별도의 추가 램프로 단정할 수 없다. 짐과 보급품이 외부에 노출되지 않는다. 강한 비와 어두운 밤이 표현되었으며 확실히 읽히는 문자는 없다.",
        "hard_violations": [],
        "physics": "앞뒤 바퀴가 도로에 닿아 차량을 지지하며 차체가 떠 있지 않다. 타이어 주변의 물보라가 빗길 주행을 뒷받침한다. 객실 불빛은 창 안의 실내 광원으로 읽힌다. 빗줄기, 물웅덩이와 노면 반사는 물리적으로 가능한 모습이다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "폭우와 소등 상태는 맞지만, 차량이 크고 카메라가 낮아 지정된 하이앵글 와이드 구도에서 벗어나며 번호판 숫자가 읽힌다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "높은 시점에서 작은 캠핑카와 넓은 대각선 도로를 담아 핵심 구도를 가장 충실히 구현했지만, 보이는 측면은 요구된 조수석 쪽과 반대다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 화면 왼쪽 아래를 향하고 도로는 오른쪽 뒤로 이어져 진행축이 일치한다. 전조등의 전방 광선은 없다. 사람의 시선은 식별되지 않는다. 앞부분과 함께 보이는 차체 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽에 해당해 조수석 측면 요구와 다르다.",
        "built_space": "노란 중앙선이 있는 2차로 도로, 양쪽 가장자리의 가드레일 두 줄, 여러 전신주와 전선, 수목이 보인다. 참조의 젖은 아스팔트와 금속 가드레일은 유지하지만 멀리 도시 건물과 다수의 불빛이 더 두드러진다. 차량은 화면 폭의 약 3분의 1을 차지하며 지붕이 조금만 보여, 도로에 둘러싸인 작은 피사체를 위에서 보는 구도와 거리가 있다.",
        "entities": "낡고 얼룩진 흰색 오버캡 캠핑카 한 대가 있으며 다른 차량이나 식별 가능한 사람은 없다. 헤드라이트는 꺼져 있다. 객실 창 안에 따뜻한 갓등 하나가 뚜렷하게 보이고 뒤쪽 창에도 약한 빛이 있다. 외부에 실린 짐은 없다. 밤과 강한 비는 구현되었으나 앞 번호판의 숫자열은 읽을 수 있어 무문자 조건에 어긋난다.",
        "hard_violations": [
         "앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 노면에 닿아 차체를 지지한다. 바퀴 주변의 물보라와 젖은 노면은 빗속 주행으로 가능한 모습이다. 실내 갓등은 창 아래에 놓인 것으로 보이며 공중에 떠 있다는 증거는 없다. 노면의 반사도 젖은 표면에서 가능한 범위다."
       },
       {
        "label": "A",
        "direction": "캠핑카는 화면 왼쪽 아래를 향하고, 도로 역시 그 방향으로 대각선으로 뻗어 진행축이 맞는다. 전조등은 빛을 쏘지 않는다. 사람의 시선은 식별되지 않는다. 다만 앞부분과 함께 노출된 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽이어서 조수석 측면 요구와 반대다.",
        "built_space": "노란 중앙선으로 나뉜 2차로 도로와 양쪽 가드레일 두 줄, 도로변 전신주 여러 개와 전선, 어두운 수목이 보인다. 참조의 주요 도로 재료와 시설은 대체로 유지된다. 오른쪽 가장자리는 참조보다 높고 교량 난간처럼 보이는 차이가 있다. 높은 외부 시점에서 지붕과 앞부분을 내려다보며, 차량 주변으로 넓은 빈 노면이 확보되어 지정된 와이드 구도에 잘 맞는다.",
        "entities": "낡은 흰색 오버캡 캠핑카 한 대만 보이며 다른 차량이나 식별 가능한 사람은 없다. 앞 헤드라이트는 꺼져 있고 후방 객실의 작은 창에 광원 하나가 켜져 있다. 옆의 큰 창에는 약한 실내 빛만 보여 별도의 추가 램프로 단정할 수 없다. 짐과 보급품이 외부에 노출되지 않는다. 강한 비와 어두운 밤이 표현되었으며 확실히 읽히는 문자는 없다.",
        "hard_violations": [],
        "physics": "앞뒤 바퀴가 도로에 닿아 차량을 지지하며 차체가 떠 있지 않다. 타이어 주변의 물보라가 빗길 주행을 뒷받침한다. 객실 불빛은 창 안의 실내 광원으로 읽힌다. 빗줄기, 물웅덩이와 노면 반사는 물리적으로 가능한 모습이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.946
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.696
   },
   "violations": {
    "B": [
     "[gpt-high] 앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 696
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요청된 '위에서 내려다본 조수석 측면 뷰'와 '넓은 여백 속 작은 피사체'라는 프레이밍 지침을 정확히 구현하여 높은 점수를 부여함."
   },
   {
    "label": "B",
    "score": 696,
    "verdict_ko": "'위에서 내려다본 구도'와 '조수석 측면 노출'이라는 명시적인 카메라 및 앵글 지시를 위반하고 피사체를 너무 크게 배치함.  ★위반: [gpt-high] 앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B04.png",
    "asset_id": "4c3502ac-ef92-4502-9e15-01f19a5f689e",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-919e-702e-98b0-574458402d6a",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S45sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:35:53.678191+00:00",
  "fingerprint": "249c067adb1bc5f50ed135e0b915f3e9e0f4bcad8f458d119a3b851066081f44",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S45sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S45sh1_sel.png",
  "source_sha256": "141ec6e614c0d5532afcf6a0d18e3d54a2f846f3037baa6c51191aaaa26af67e",
  "file": "S45sh1_cine.png",
  "staged_sha256": "b7efd4c215214e712189839560f568aa85e5f0ad057d510fe1883490157d73d1",
  "latency_ms": 17020
 },
 "S45sh4::confined_fp_apt": {
  "applies": true,
  "reason_ko": "캠핑카 운전석이라는 제한된 내부 공간에서 운전자인 현우가 백미러를 통해 뒷좌석을 주시하는 상황입니다. 운전석, 백미러, 뒷좌석의 정확한 위치 관계와 시선 방향을 올바르게 연출해야만 하는 샷이므로 평면도 보조가 유용합니다.",
  "input_fingerprint": "6de6aca07def2239"
 },
 "S45sh4::signage": {
  "fp": "ac469a439fca3474",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::c0923a6850ca": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_c0923a6850ca.png",
  "place_text": "Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind.",
  "input_fingerprint": "76c4295ae003be69"
 },
 "S45sh4::confined_fp": {
  "reads": {
   "controls": "The steering wheel is located at the front left seat.",
   "mirrors": "A rearview mirror is mounted at the top center of the windshield area, facing backward toward the rear cabin.",
   "camera": "The camera is located between the driver and passenger seats, pointing forward and upward directly at the rearview mirror.",
   "occupants": "현우 occupies the front left driver's seat. The passenger seat is empty."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned between the front seats, aiming forward and slightly upward. The rearview mirror occupies the upper center of the frame, its reflective glass facing directly back at the lens. From this angle, the mirror's surface physically reflects the face of 현우, who sits in the front left driver's seat. The back of 현우's head and the steering wheel are situated in the lower left foreground. The unoccupied passenger seat is visible in the lower right foreground. The rear lamp is located behind the camera and does not appear in this forward-facing shot.",
  "fixed": false,
  "input_fingerprint": "a95da29a919b3e74"
 },
 "era_assess::1e11ec86026bde7c": {
  "subjects": [],
  "subject_text": "캠핑카 내부\n앞쪽 운전석과 조수석 뒤로 침대와 뒷좌석이 이어지는 소형 이동식 주거 공간. 작은 전등과 측면 창문, 뒤쪽 출입문이 있다.",
  "identity": "canonical",
  "scope_id": "L199",
  "scope_role": "location_interior",
  "scope_sha": "efdbc327adcdaedd"
 },
 "S45sh4": {
  "input_fingerprint": "c48370caa9aec348",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 백미러를 통해 뒷좌석을 차갑게 노려보는 현우의 매서운 눈매 클로즈업.\n\nLOCATION (lock): Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rearview mirror (Reflecting 현우's narrowed eyes) — Its reflective face is visible with a small rim retained; the reflected face is mirror-reversed and its eyeline runs toward the rear seats rather than the lens; used as Frames the indirect view and separates his hostile scrutiny from direct audience address; Front cabin (Occupied driving compartment) — Only soft peripheral fragments remain around the mirror; used as Maintains an interior spatial reference around the reflected close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the dark cabin and restrained illumination from the small rear-seat light, allowing the reflected eyes to remain readable without an added source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy rain continues around the moving camper; its headlights remain off and the small rear lamp remains lit. The luggage and supplies remain inside. 현우: He remains at the wheel with facial bruises and the untreated leg wound. Yoon's contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, aiming forward and slightly upward. The rearview mirror occupies the upper center of the frame, its reflective glass facing directly back at the lens. From this angle, the mirror's surface physically reflects the face of 현우, who sits in the front left driver's seat. The back of 현우's head and the steering wheel are situated in the lower left foreground. The unoccupied passenger seat is visible in the lower right foreground. The rear lamp is located behind the camera and does not appear in this forward-facing shot.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 백미러를 통해 뒷좌석을 차갑게 노려보는 현우의 매서운 눈매 클로즈업.\n\nLOCATION (lock): Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the dark cabin and restrained illumination from the small rear-seat light, allowing the reflected eyes to remain readable without an added source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy rain continues around the moving camper; its headlights remain off and the small rear lamp remains lit. The luggage and supplies remain inside. 현우: He remains at the wheel with facial bruises and the untreated leg wound. Yoon's contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, aiming forward and slightly upward. The rearview mirror occupies the upper center of the frame, its reflective glass facing directly back at the lens. From this angle, the mirror's surface physically reflects the face of 현우, who sits in the front left driver's seat. The back of 현우's head and the steering wheel are situated in the lower left foreground. The unoccupied passenger seat is visible in the lower right foreground. The rear lamp is located behind the camera and does not appear in this forward-facing shot.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 백미러를 통해 뒷좌석을 차갑게 노려보는 현우의 매서운 눈매 클로즈업.\n\nLOCATION (lock): Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the dark cabin and restrained illumination from the small rear-seat light, allowing the reflected eyes to remain readable without an added source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy rain continues around the moving camper; its headlights remain off and the small rear lamp remains lit. The luggage and supplies remain inside. 현우: He remains at the wheel with facial bruises and the untreated leg wound. Yoon's contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh4_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh4_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "백미러에 비친 현우의 시선이 뒷좌석(카메라 렌즈 방향)을 명확히 향하고 있음.",
    "built_space": "차량 운전석 구조가 확인되며, 스티어링 휠과 백미러, 조수석이 올바른 위치에 배치됨.",
    "entities": "현우의 외모는 레퍼런스와 일치하나, 요구된 얼굴의 멍 자국은 보이지 않음.",
    "hard_violations": [],
    "physics": "인물이 운전석에 자연스럽게 앉아 자세를 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "백미러 속 시선이 뒷좌석 쪽을 향하고 있음.",
    "built_space": "대시보드와 차량 전면부가 넓게 보이며 백미러와 좌석이 위치하지만, 요구된 클로즈업보다 배경이 너무 많이 노출됨.",
    "entities": "현우의 외모가 일치하며 얼굴에 멍 자국이 표현됨.",
    "hard_violations": [
     "[gemini-pro] 차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
    ],
    "physics": "인물은 운전석에 앉아 있으나, 실내 허공과 물체 표면에 빗방울이 떠 있어 물리적으로 불가능한 상태임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 클로즈업 프레이밍에 더 가깝게 연출되었으나, 프롬프트에 명시된 얼굴의 멍과 내리는 비의 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴의 멍은 묘사되었으나 프레이밍이 너무 넓고, 차량 내부에 빗방울이 맺히는 치명적인 렌더링 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백미러에 비친 현우의 시선이 뒷좌석(카메라 렌즈 방향)을 명확히 향하고 있음.",
        "built_space": "차량 운전석 구조가 확인되며, 스티어링 휠과 백미러, 조수석이 올바른 위치에 배치됨.",
        "entities": "현우의 외모는 레퍼런스와 일치하나, 요구된 얼굴의 멍 자국은 보이지 않음.",
        "hard_violations": [],
        "physics": "인물이 운전석에 자연스럽게 앉아 자세를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "백미러 속 시선이 뒷좌석 쪽을 향하고 있음.",
        "built_space": "대시보드와 차량 전면부가 넓게 보이며 백미러와 좌석이 위치하지만, 요구된 클로즈업보다 배경이 너무 많이 노출됨.",
        "entities": "현우의 외모가 일치하며 얼굴에 멍 자국이 표현됨.",
        "hard_violations": [
         "차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
        ],
        "physics": "인물은 운전석에 앉아 있으나, 실내 허공과 물체 표면에 빗방울이 떠 있어 물리적으로 불가능한 상태임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 클로즈업 프레이밍에 더 가깝게 연출되었으나, 프롬프트에 명시된 얼굴의 멍과 내리는 비의 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴의 멍은 묘사되었으나 프레이밍이 너무 넓고, 차량 내부에 빗방울이 맺히는 치명적인 렌더링 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "백미러에 비친 현우의 시선이 뒷좌석(카메라 렌즈 방향)을 명확히 향하고 있음.",
        "built_space": "차량 운전석 구조가 확인되며, 스티어링 휠과 백미러, 조수석이 올바른 위치에 배치됨.",
        "entities": "현우의 외모는 레퍼런스와 일치하나, 요구된 얼굴의 멍 자국은 보이지 않음.",
        "hard_violations": [],
        "physics": "인물이 운전석에 자연스럽게 앉아 자세를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "백미러 속 시선이 뒷좌석 쪽을 향하고 있음.",
        "built_space": "대시보드와 차량 전면부가 넓게 보이며 백미러와 좌석이 위치하지만, 요구된 클로즈업보다 배경이 너무 많이 노출됨.",
        "entities": "현우의 외모가 일치하며 얼굴에 멍 자국이 표현됨.",
        "hard_violations": [
         "차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
        ],
        "physics": "인물은 운전석에 앉아 있으나, 실내 허공과 물체 표면에 빗방울이 떠 있어 물리적으로 불가능한 상태임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "야간 폭우와 얼굴 타박상은 보이지만, 운전실 전체를 넓게 담아 핵심인 백미러 속 눈매 클로즈업을 놓쳤다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "거울 속 눈매를 더 크게 담고 작은 후방등과 어두운 실내를 살렸지만, 여전히 구도가 넓고 시선도 렌즈와 충분히 분리되지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 전방을 향해 앉아 있고, 거울 속 좁힌 눈은 거울을 통해 차량 뒤쪽 중앙, 카메라에 가까운 방향을 바라본다. 화면 밖 뒷좌석의 특정 지점을 겨냥한다기보다 관객을 정면으로 보는 인상이 강하다. 차량의 실제 진행 방향은 정지 화면만으로 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개, 중앙 송풍구 두 개가 보인다. 좌측 운전석 배치는 평면도와 부합하고, 중앙에서 거울에 운전자의 얼굴이 보이는 반사는 가능한 배치다. 다만 운전자의 뒤통수와 좌석, 대시보드, 앞유리가 넓게 노출되어 운전실이 부드러운 주변 파편으로만 남아야 한다는 지시와 다르다. 후방의 작은 등은 식별되지 않는다.",
        "entities": "인물은 현우 한 명이며 거울상은 추가 인물이 아니다. 검은 헝클어진 머리와 동아시아계 젊은 남성의 외형, 어두운 상의가 참조와 대체로 맞고 거울 속 볼에 붉은 타박상이 보인다. 정확한 나이나 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 백미러와 빗물이 흐르는 앞유리가 보이며 밤으로 읽힌다. 다리 상처, 신발 속 카드, 짐은 구도 밖이라 판단하지 않는다. 판독 가능한 문구는 없다.",
        "hard_violations": [],
        "physics": "현우의 몸은 운전석에 앉아 지지되고 머리는 목과 몸통에 연결되어 있다. 백미러는 위쪽 장착대로 고정되고 운전대와 좌석도 차체에 연결되어 있다. 손은 명확히 보이지 않아 운전대를 잡았는지는 확인할 수 없지만, 떠 있는 신체나 무지지 물체는 확인되지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 운전석에서 앞을 향하고 거울 속 눈은 약간 화면 왼쪽으로 치우친 채 차량 뒤쪽 중앙을 주시한다. 차갑게 좁힌 눈매는 보이지만 시선의 도착점이 카메라 근처여서, 렌즈가 아닌 뒷좌석을 노려본다는 구분은 명확하지 않다. 차량의 이동 자체는 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 일부 노출된 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개가 보인다. 거울 오른쪽에는 켜진 작은 후방등 하나가 반사되어 있다. 좌측 운전자와 중앙 거울의 배치는 평면도와 양립하며 반사도 불가능하다고 볼 근거는 없다. 거울은 A보다 크게 잡혔지만 뒤통수와 두 좌석, 앞유리가 여전히 화면 상당 부분을 차지해 눈매 중심 클로즈업에는 못 미친다.",
        "entities": "현우 한 명의 뒤통수와 거울 속 얼굴이 보인다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리는 인물 참조와 대체로 일치한다. 의복은 어둡고 흐린 일부만 보여 정확한 색과 형태를 확정하기 어렵다. 눈은 정상적인 사람의 눈이며 표정으로 경계심을 표현한다. 뚜렷한 얼굴 타박상은 A보다 약하고, 폭우도 명확하게 드러나지 않는다. 작은 후방등과 어두운 야간 실내는 보인다. 상처 난 다리와 숨겨진 카드, 짐은 구도 밖이며 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 운전석에 앉아 있고 목과 머리의 연결도 자연스럽다. 거울 위의 어두운 연결부가 차체 쪽으로 이어지며, 좌석과 운전대는 고정된 차량 부품으로 보인다. 거울 속 등은 후방 벽면에 붙은 조명으로 읽힌다. 손의 운전대 접촉은 화면에서 확인되지 않지만, 지지 없이 떠 있는 물체나 불가능한 신체 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "야간 폭우와 얼굴 타박상은 보이지만, 운전실 전체를 넓게 담아 핵심인 백미러 속 눈매 클로즈업을 놓쳤다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "거울 속 눈매를 더 크게 담고 작은 후방등과 어두운 실내를 살렸지만, 여전히 구도가 넓고 시선도 렌즈와 충분히 분리되지 않는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 전방을 향해 앉아 있고, 거울 속 좁힌 눈은 거울을 통해 차량 뒤쪽 중앙, 카메라에 가까운 방향을 바라본다. 화면 밖 뒷좌석의 특정 지점을 겨냥한다기보다 관객을 정면으로 보는 인상이 강하다. 차량의 실제 진행 방향은 정지 화면만으로 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개, 중앙 송풍구 두 개가 보인다. 좌측 운전석 배치는 평면도와 부합하고, 중앙에서 거울에 운전자의 얼굴이 보이는 반사는 가능한 배치다. 다만 운전자의 뒤통수와 좌석, 대시보드, 앞유리가 넓게 노출되어 운전실이 부드러운 주변 파편으로만 남아야 한다는 지시와 다르다. 후방의 작은 등은 식별되지 않는다.",
        "entities": "인물은 현우 한 명이며 거울상은 추가 인물이 아니다. 검은 헝클어진 머리와 동아시아계 젊은 남성의 외형, 어두운 상의가 참조와 대체로 맞고 거울 속 볼에 붉은 타박상이 보인다. 정확한 나이나 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 백미러와 빗물이 흐르는 앞유리가 보이며 밤으로 읽힌다. 다리 상처, 신발 속 카드, 짐은 구도 밖이라 판단하지 않는다. 판독 가능한 문구는 없다.",
        "hard_violations": [],
        "physics": "현우의 몸은 운전석에 앉아 지지되고 머리는 목과 몸통에 연결되어 있다. 백미러는 위쪽 장착대로 고정되고 운전대와 좌석도 차체에 연결되어 있다. 손은 명확히 보이지 않아 운전대를 잡았는지는 확인할 수 없지만, 떠 있는 신체나 무지지 물체는 확인되지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 운전석에서 앞을 향하고 거울 속 눈은 약간 화면 왼쪽으로 치우친 채 차량 뒤쪽 중앙을 주시한다. 차갑게 좁힌 눈매는 보이지만 시선의 도착점이 카메라 근처여서, 렌즈가 아닌 뒷좌석을 노려본다는 구분은 명확하지 않다. 차량의 이동 자체는 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 일부 노출된 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개가 보인다. 거울 오른쪽에는 켜진 작은 후방등 하나가 반사되어 있다. 좌측 운전자와 중앙 거울의 배치는 평면도와 양립하며 반사도 불가능하다고 볼 근거는 없다. 거울은 A보다 크게 잡혔지만 뒤통수와 두 좌석, 앞유리가 여전히 화면 상당 부분을 차지해 눈매 중심 클로즈업에는 못 미친다.",
        "entities": "현우 한 명의 뒤통수와 거울 속 얼굴이 보인다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리는 인물 참조와 대체로 일치한다. 의복은 어둡고 흐린 일부만 보여 정확한 색과 형태를 확정하기 어렵다. 눈은 정상적인 사람의 눈이며 표정으로 경계심을 표현한다. 뚜렷한 얼굴 타박상은 A보다 약하고, 폭우도 명확하게 드러나지 않는다. 작은 후방등과 어두운 야간 실내는 보인다. 상처 난 다리와 숨겨진 카드, 짐은 구도 밖이며 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 운전석에 앉아 있고 목과 머리의 연결도 자연스럽다. 거울 위의 어두운 연결부가 차체 쪽으로 이어지며, 좌석과 운전대는 고정된 차량 부품으로 보인다. 거울 속 등은 후방 벽면에 붙은 조명으로 읽힌다. 손의 운전대 접촉은 화면에서 확인되지 않지만, 지지 없이 떠 있는 물체나 불가능한 신체 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.167
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.917
   },
   "violations": {
    "B": [
     "[gemini-pro] 차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 917
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 프레이밍에 더 가깝게 연출되었으나, 프롬프트에 명시된 얼굴의 멍과 내리는 비의 묘사가 누락되었습니다."
   },
   {
    "label": "B",
    "score": 917,
    "verdict_ko": "얼굴의 멍은 묘사되었으나 프레이밍이 너무 넓고, 차량 내부에 빗방울이 맺히는 치명적인 렌더링 오류가 발생했습니다.  ★위반: [gemini-pro] 차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh4_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-934b-7dba-81e6-a89c38c6a30e",
  "confined_fp": {
   "base_key": "confinedfp::c0923a6850ca",
   "apt_reason": "캠핑카 운전석이라는 제한된 내부 공간에서 운전자인 현우가 백미러를 통해 뒷좌석을 주시하는 상황입니다. 운전석, 백미러, 뒷좌석의 정확한 위치 관계와 시선 방향을 올바르게 연출해야만 하는 샷이므로 평면도 보조가 유용합니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S45sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:38:16.303796+00:00",
  "fingerprint": "67ccc9251672d609d2e07e25dc1e4aff686e3349ab392e8e0b07fb0e3a552200",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S45sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S45sh4_sel.png",
  "source_sha256": "a0b1518b5b6a5cef0e5cfa74be48fa6fb15e8f4e0dee623dcf2a870aa5e1a8a7",
  "file": "S45sh4_cine.png",
  "staged_sha256": "c186db603dbde576ccca476b765a590a6e82926d6dbf5fe00af8e84987fbfee3",
  "latency_ms": 11821
 },
 "S45sh9::signage": {
  "fp": "40fbd5b51c65c2e1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S45sh9::bgfirst_bg": {
  "input_fingerprint": "e4e360d1aa9b3fbf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh9__bgfirst_bg.png",
  "asset_id": "dc0bb94a-03a6-4948-805a-1b44a4d89d2d",
  "input_asset_ids": [
   "b921f333-63fe-415e-a03d-af2cd1337f45",
   "4be906f8-5377-4122-9528-ee6aaccbd804"
  ]
 },
 "S45sh9": {
  "input_fingerprint": "74a50c3e4a40679c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rainwater leaks through the camper's roof and window frames despite cloth pressed against the openings. The headlights remain off and the small rear lamp remains lit. 앰버: She presses clothing against a leaking window frame.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rainwater leaks through the camper's roof and window frames despite cloth pressed against the openings. The headlights remain off and the small rear lamp remains lit. 앰버: She presses clothing against a leaking window frame.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rainwater leaks through the camper's roof and window frames despite cloth pressed against the openings. The headlights remain off and the small rear lamp remains lit. 앰버: She presses clothing against a leaking window frame.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh9__bgfirst_bg.png",
     "asset_id": "dc0bb94a-03a6-4948-805a-1b44a4d89d2d",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S45sh9.png",
     "asset_id": "b921f333-63fe-415e-a03d-af2cd1337f45",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
     "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선과 양손이 창문 틈새에 밀착시킨 헝겊을 정확히 향하고 있음.",
    "built_space": "캠핑카 내부의 뒷좌석 소파와 창문이 나타나나, 레퍼런스와 달리 오른쪽에 커튼이 추가로 묘사됨.",
    "entities": "앰버의 외모(금발, 파란 티셔츠, 멜빵바지)와 헝겊이 레퍼런스 및 프롬프트와 일치함.",
    "hard_violations": [],
    "physics": "하체로 무게를 지탱하며 상체를 앞으로 기울여 창문을 강하게 누르는 자세가 자연스럽고 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "시선과 양손이 창문에 댄 헝겊을 향하고 있음.",
    "built_space": "뒷좌석 소파와 조명이 보이나 레퍼런스의 공간 구조와 비교해 위치와 비율이 왜곡됨.",
    "entities": "앰버의 인상착의와 파란색 헝겊이 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 창문 유리에 앰버가 아닌 정체불명의 짧은 머리 성인 얼굴이 반사되어 나타남 (invented people/impossible reflection)"
    ],
    "physics": "무릎을 꿇고 양손으로 헝겊을 누르고 있으나, 창문에 비친 반사체가 물리적 상황과 전혀 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트의 요구대로 앰버가 창문 틈새로 들어오는 비를 막기 위해 헝겊을 누르며 찡그리는 모습을 사실적인 미디엄 샷으로 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창문 유리에 등장인물과 전혀 다른 정체불명의 인물 얼굴이 반사되어 나타나는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선과 양손이 창문 틈새에 밀착시킨 헝겊을 정확히 향하고 있음.",
        "built_space": "캠핑카 내부의 뒷좌석 소파와 창문이 나타나나, 레퍼런스와 달리 오른쪽에 커튼이 추가로 묘사됨.",
        "entities": "앰버의 외모(금발, 파란 티셔츠, 멜빵바지)와 헝겊이 레퍼런스 및 프롬프트와 일치함.",
        "hard_violations": [],
        "physics": "하체로 무게를 지탱하며 상체를 앞으로 기울여 창문을 강하게 누르는 자세가 자연스럽고 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "시선과 양손이 창문에 댄 헝겊을 향하고 있음.",
        "built_space": "뒷좌석 소파와 조명이 보이나 레퍼런스의 공간 구조와 비교해 위치와 비율이 왜곡됨.",
        "entities": "앰버의 인상착의와 파란색 헝겊이 묘사됨.",
        "hard_violations": [
         "창문 유리에 앰버가 아닌 정체불명의 짧은 머리 성인 얼굴이 반사되어 나타남 (invented people/impossible reflection)"
        ],
        "physics": "무릎을 꿇고 양손으로 헝겊을 누르고 있으나, 창문에 비친 반사체가 물리적 상황과 전혀 맞지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트의 요구대로 앰버가 창문 틈새로 들어오는 비를 막기 위해 헝겊을 누르며 찡그리는 모습을 사실적인 미디엄 샷으로 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창문 유리에 등장인물과 전혀 다른 정체불명의 인물 얼굴이 반사되어 나타나는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선과 양손이 창문 틈새에 밀착시킨 헝겊을 정확히 향하고 있음.",
        "built_space": "캠핑카 내부의 뒷좌석 소파와 창문이 나타나나, 레퍼런스와 달리 오른쪽에 커튼이 추가로 묘사됨.",
        "entities": "앰버의 외모(금발, 파란 티셔츠, 멜빵바지)와 헝겊이 레퍼런스 및 프롬프트와 일치함.",
        "hard_violations": [],
        "physics": "하체로 무게를 지탱하며 상체를 앞으로 기울여 창문을 강하게 누르는 자세가 자연스럽고 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "시선과 양손이 창문에 댄 헝겊을 향하고 있음.",
        "built_space": "뒷좌석 소파와 조명이 보이나 레퍼런스의 공간 구조와 비교해 위치와 비율이 왜곡됨.",
        "entities": "앰버의 인상착의와 파란색 헝겊이 묘사됨.",
        "hard_violations": [
         "창문 유리에 앰버가 아닌 정체불명의 짧은 머리 성인 얼굴이 반사되어 나타남 (invented people/impossible reflection)"
        ],
        "physics": "무릎을 꿇고 양손으로 헝겊을 누르고 있으나, 창문에 비친 반사체가 물리적 상황과 전혀 맞지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찡그린 상체와 양손의 압박, 실내로 흐르는 누수, 뒤쪽 좌석과 켜진 단일 벽등의 관계를 가장 충실하게 구현한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "상체 중심 구도와 힘주는 표정은 적절하지만, 뒤쪽 좌석에서 떨어진 위치와 불분명한 실내 누수·후미등 점등 상태가 장소 및 행동 조건을 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 오른쪽 창문을 향해 몸을 기울이고, 찡그린 눈과 얼굴을 두 손의 헝겊 접촉부로 향한다. 양손은 헝겊을 창문 하단 틀 쪽으로 밀고 있다. 위쪽에서 내려오는 물줄기는 손 바로 위 창면을 따라 흐르며, 실제 틈의 위치는 완전히 선명하지 않다.",
        "built_space": "격자무늬 후면 벤치 하나, 오른쪽 측면 창문 하나, 켜진 작은 벽등 하나와 벽면 스위치가 보인다. 앰버는 벤치 오른쪽 끝에 앉아 옆 창문에 손을 뻗고 있어 좌석·창문·후미등의 배치가 참조 장소와 대체로 맞는다. 창틀 안쪽 가장자리가 비스듬히 보인다. 구도는 상체와 허벅지 일부를 포함해 요구된 미디엄 숏보다 약간 넉넉하다.",
        "entities": "인물은 금발의 어린 여자아이 한 명으로, 둥근 얼굴과 체격이 앰버 설정에 대체로 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 남색 반소매와 낡은 갈색 멜빵 작업복은 참조와 맞지만, 머리 위 보호장비는 보이지 않는다. 양손 사이에는 짙은 청회색 옷감이 있고, 젖은 창문과 실내로 떨어지는 물방울이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 벤치 좌면에 지지되고, 앞으로 기울인 몸에서 뻗은 두 팔이 창틀에 압력을 전달한다. 헝겊은 양손과 창틀 사이에 눌려 있으며 아래쪽은 중력 방향으로 늘어진다. 창면의 물줄기와 헝겊 아래 물방울도 아래로 떨어진다. 지지 없이 떠 있는 신체나 물건은 없다."
       },
       {
        "label": "B",
        "direction": "앰버의 얼굴과 좁힌 눈은 오른쪽 창문의 헝겊 쪽을 향한다. 양손은 헝겊을 창문 하단 가까이에 밀어붙이지만, 주된 접촉 면은 틈 자체보다 유리면으로 읽힌다. 창의 빗방울은 뚜렷하나 그 접촉부를 통해 실내로 들어오는 물줄기는 분명하지 않다.",
        "built_space": "후면 벤치 하나가 인물 뒤로 떨어져 보이고, 좌우 측면 창문 두 개, 커튼, 왼쪽 전경 좌석 일부가 보인다. 앰버는 뒤쪽 벤치 위가 아니라 그보다 앞쪽 창가에 자리한 것으로 읽혀 지정된 후면 좌석 옆 위치와의 일치가 약하다. 오른쪽 창 옆 작은 벽등 형태는 보이지만 밝게 켜진 상태는 확인하기 어렵고, 창밖의 따뜻한 불빛도 보인다. 상체 중심 미디엄 구도와 비스듬한 창틀 표현은 적절하다.",
        "entities": "금발의 어린 여자아이 한 명이며, 둥근 얼굴과 아동 체격은 앰버 설정에 대체로 맞는다. 혼혈 배경은 이미지로 확정할 수 없다. 남색 반소매, 갈색 멜빵 작업복과 허리 도구가 참조 복장을 따른다. 머리 위 보호장비는 없다. 회색 헝겊, 젖은 손과 팔, 빗방울 맺힌 창문은 보이지만 실내 누수의 발생 지점은 불명확하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손이 헝겊을 잡고 창 쪽으로 눌러 지지하며, 남은 옷감은 아래로 처진다. 팔꿈치와 어깨의 굽힘은 힘주는 동작으로 가능하다. 하체의 지지점은 화면 밖이어서 앉았는지 서 있는지 확정할 수 없지만, 몸이 공중에 떠 있다고 볼 근거도 없다. 젖은 유리와 피부의 물방울은 물리적으로 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "찡그린 상체와 양손의 압박, 실내로 흐르는 누수, 뒤쪽 좌석과 켜진 단일 벽등의 관계를 가장 충실하게 구현한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "상체 중심 구도와 힘주는 표정은 적절하지만, 뒤쪽 좌석에서 떨어진 위치와 불분명한 실내 누수·후미등 점등 상태가 장소 및 행동 조건을 약화한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 오른쪽 창문을 향해 몸을 기울이고, 찡그린 눈과 얼굴을 두 손의 헝겊 접촉부로 향한다. 양손은 헝겊을 창문 하단 틀 쪽으로 밀고 있다. 위쪽에서 내려오는 물줄기는 손 바로 위 창면을 따라 흐르며, 실제 틈의 위치는 완전히 선명하지 않다.",
        "built_space": "격자무늬 후면 벤치 하나, 오른쪽 측면 창문 하나, 켜진 작은 벽등 하나와 벽면 스위치가 보인다. 앰버는 벤치 오른쪽 끝에 앉아 옆 창문에 손을 뻗고 있어 좌석·창문·후미등의 배치가 참조 장소와 대체로 맞는다. 창틀 안쪽 가장자리가 비스듬히 보인다. 구도는 상체와 허벅지 일부를 포함해 요구된 미디엄 숏보다 약간 넉넉하다.",
        "entities": "인물은 금발의 어린 여자아이 한 명으로, 둥근 얼굴과 체격이 앰버 설정에 대체로 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 남색 반소매와 낡은 갈색 멜빵 작업복은 참조와 맞지만, 머리 위 보호장비는 보이지 않는다. 양손 사이에는 짙은 청회색 옷감이 있고, 젖은 창문과 실내로 떨어지는 물방울이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 벤치 좌면에 지지되고, 앞으로 기울인 몸에서 뻗은 두 팔이 창틀에 압력을 전달한다. 헝겊은 양손과 창틀 사이에 눌려 있으며 아래쪽은 중력 방향으로 늘어진다. 창면의 물줄기와 헝겊 아래 물방울도 아래로 떨어진다. 지지 없이 떠 있는 신체나 물건은 없다."
       },
       {
        "label": "A",
        "direction": "앰버의 얼굴과 좁힌 눈은 오른쪽 창문의 헝겊 쪽을 향한다. 양손은 헝겊을 창문 하단 가까이에 밀어붙이지만, 주된 접촉 면은 틈 자체보다 유리면으로 읽힌다. 창의 빗방울은 뚜렷하나 그 접촉부를 통해 실내로 들어오는 물줄기는 분명하지 않다.",
        "built_space": "후면 벤치 하나가 인물 뒤로 떨어져 보이고, 좌우 측면 창문 두 개, 커튼, 왼쪽 전경 좌석 일부가 보인다. 앰버는 뒤쪽 벤치 위가 아니라 그보다 앞쪽 창가에 자리한 것으로 읽혀 지정된 후면 좌석 옆 위치와의 일치가 약하다. 오른쪽 창 옆 작은 벽등 형태는 보이지만 밝게 켜진 상태는 확인하기 어렵고, 창밖의 따뜻한 불빛도 보인다. 상체 중심 미디엄 구도와 비스듬한 창틀 표현은 적절하다.",
        "entities": "금발의 어린 여자아이 한 명이며, 둥근 얼굴과 아동 체격은 앰버 설정에 대체로 맞는다. 혼혈 배경은 이미지로 확정할 수 없다. 남색 반소매, 갈색 멜빵 작업복과 허리 도구가 참조 복장을 따른다. 머리 위 보호장비는 없다. 회색 헝겊, 젖은 손과 팔, 빗방울 맺힌 창문은 보이지만 실내 누수의 발생 지점은 불명확하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손이 헝겊을 잡고 창 쪽으로 눌러 지지하며, 남은 옷감은 아래로 처진다. 팔꿈치와 어깨의 굽힘은 힘주는 동작으로 가능하다. 하체의 지지점은 화면 밖이어서 앉았는지 서 있는지 확정할 수 없지만, 몸이 공중에 떠 있다고 볼 근거도 없다. 젖은 유리와 피부의 물방울은 물리적으로 자연스럽다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 창문 유리에 앰버가 아닌 정체불명의 짧은 머리 성인 얼굴이 반사되어 나타남 (invented people/impossible reflection)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트의 요구대로 앰버가 창문 틈새로 들어오는 비를 막기 위해 헝겊을 누르며 찡그리는 모습을 사실적인 미디엄 샷으로 잘 구현했습니다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "창문 유리에 등장인물과 전혀 다른 정체불명의 인물 얼굴이 반사되어 나타나는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 창문 유리에 앰버가 아닌 정체불명의 짧은 머리 성인 얼굴이 반사되어 나타남 (invented people/impossible reflection)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
    "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-9669-7520-9e2b-1d8691e48f84",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh9__bgfirst_bg.png",
   "bg_asset_id": "dc0bb94a-03a6-4948-805a-1b44a4d89d2d",
   "bg_record_key": "S45sh9::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S45sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:04:26.622226+00:00",
  "fingerprint": "2600df857deb0cb09e2eba61e28a00c8b602a82b8d422b7b28d1b07638f23294",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S45sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S45sh9_sel.png",
  "source_sha256": "af5e3cfac8b3a178b6b0340e4a027c54d0c9ae641a4c7c2dc6799849767c5c43",
  "file": "S45sh9_cine.png",
  "staged_sha256": "7c1915659384924416dbea0bcb12242f22dc46a8eb1516340bc981bd052710fc",
  "latency_ms": 10832
 },
 "S46sh1::signage": {
  "fp": "014702876d03f367",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S46sh1": {
  "input_fingerprint": "56f98fbf45873d89",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 금이 가고 마른 넝쿨이 잔뜩 얽힌 낡은 미추홀도서관 외벽 앞에 멈춰 선 캠핑카 전경.\n\nLOCATION (lock): Outside the abandoned library, in front of its cracked, vine-covered facade, where the camper stops in nighttime rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Stopped in front of the library) — Seen obliquely from above, with its side and roof visible; used as Foreground scale reference and connection to the previous road scene; Library facade (Large cracks across the walls, with dry vines clinging densely to them) — The building front recedes obliquely across the upper frame; used as Establishes the abandoned destination without overwhelming the vehicle-ground relationship; Library sign (Old) — The lettered face is visible to camera and reads '미추홀도서관, 인천'; used as Identifies the destination within the establishing composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the rainy nighttime setting with subdued ambient exposure and controlled contrast, without specifying an unsupported exterior light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The leaking camper is stopped outside the library in nighttime rain. Dry vines cover the deeply cracked exterior beneath the aged sign reading “미추홀도서관, 인천.”\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 금이 가고 마른 넝쿨이 잔뜩 얽힌 낡은 미추홀도서관 외벽 앞에 멈춰 선 캠핑카 전경.\n\nLOCATION (lock): Outside the abandoned library, in front of its cracked, vine-covered facade, where the camper stops in nighttime rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Stopped in front of the library) — Seen obliquely from above, with its side and roof visible; used as Foreground scale reference and connection to the previous road scene; Library facade (Large cracks across the walls, with dry vines clinging densely to them) — The building front recedes obliquely across the upper frame; used as Establishes the abandoned destination without overwhelming the vehicle-ground relationship; Library sign (Old) — The lettered face is visible to camera and reads '미추홀도서관, 인천'; used as Identifies the destination within the establishing composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the rainy nighttime setting with subdued ambient exposure and controlled contrast, without specifying an unsupported exterior light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The leaking camper is stopped outside the library in nighttime rain. Dry vines cover the deeply cracked exterior beneath the aged sign reading “미추홀도서관, 인천.”\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 금이 가고 마른 넝쿨이 잔뜩 얽힌 낡은 미추홀도서관 외벽 앞에 멈춰 선 캠핑카 전경.\n\nLOCATION (lock): Outside the abandoned library, in front of its cracked, vine-covered facade, where the camper stops in nighttime rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Stopped in front of the library) — Seen obliquely from above, with its side and roof visible; used as Foreground scale reference and connection to the previous road scene; Library facade (Large cracks across the walls, with dry vines clinging densely to them) — The building front recedes obliquely across the upper frame; used as Establishes the abandoned destination without overwhelming the vehicle-ground relationship; Library sign (Old) — The lettered face is visible to camera and reads '미추홀도서관, 인천'; used as Identifies the destination within the establishing composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the rainy nighttime setting with subdued ambient exposure and controlled contrast, without specifying an unsupported exterior light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The leaking camper is stopped outside the library in nighttime rain. Dry vines cover the deeply cracked exterior beneath the aged sign reading “미추홀도서관, 인천.”\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 도서관 입구와 주차된 캠핑카를 비스듬히 위에서 내려다보고 있음.",
    "built_space": "레퍼런스와 일치하는 도서관 건물, 유리문 입구, 주차장 구획선이 올바르게 배치됨.",
    "entities": "비 오는 밤, 금이 가고 넝쿨이 얽힌 도서관 외벽, 지붕과 측면이 보이는 낡은 캠핑카(모터홈). 단, 간판 텍스트는 프롬프트와 달리 '시립도서관'으로 나타남.",
    "hard_violations": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ],
    "physics": "캠핑카가 젖은 아스팔트 바닥에 안정적으로 정차해 있으며 빗방울과 바닥 반사가 자연스러움."
   },
   {
    "label": "B",
    "direction": "카메라가 도서관 건물과 트레일러를 비스듬히 위에서 내려다보고 있음.",
    "built_space": "도서관 건물 구조는 유사하나, 좌측 원경의 도시 건물들이 레퍼런스보다 과장되게 가깝게 변경됨.",
    "entities": "비 오는 밤, 넝쿨이 얽힌 도서관 외벽. 운전석이 없는 트레일러 형태의 캠핑카. 간판은 역시 '시립도서관'으로 출력됨.",
    "hard_violations": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ],
    "physics": "트레일러가 바닥에 놓여 있으나, 정차한 상태를 설명할 만한 견인 차량이 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 간판 텍스트('미추홀도서관, 인천') 대신 레퍼런스의 글씨('시립도서관')가 그대로 출력되었으나, 지정된 부감 앵글과 낡은 캠핑카(모터홈)의 묘사가 프롬프트의 지시를 충실히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "간판 텍스트 요구사항을 실패했으며, 캠핑카가 견인차 없는 트레일러 형태로 묘사되어 '도착해서 멈춰 선' 상황에 어울리지 않고 배경 원경이 레퍼런스에 비해 과장되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 도서관 입구와 주차된 캠핑카를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "레퍼런스와 일치하는 도서관 건물, 유리문 입구, 주차장 구획선이 올바르게 배치됨.",
        "entities": "비 오는 밤, 금이 가고 넝쿨이 얽힌 도서관 외벽, 지붕과 측면이 보이는 낡은 캠핑카(모터홈). 단, 간판 텍스트는 프롬프트와 달리 '시립도서관'으로 나타남.",
        "hard_violations": [],
        "physics": "캠핑카가 젖은 아스팔트 바닥에 안정적으로 정차해 있으며 빗방울과 바닥 반사가 자연스러움."
       },
       {
        "label": "B",
        "direction": "카메라가 도서관 건물과 트레일러를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "도서관 건물 구조는 유사하나, 좌측 원경의 도시 건물들이 레퍼런스보다 과장되게 가깝게 변경됨.",
        "entities": "비 오는 밤, 넝쿨이 얽힌 도서관 외벽. 운전석이 없는 트레일러 형태의 캠핑카. 간판은 역시 '시립도서관'으로 출력됨.",
        "hard_violations": [],
        "physics": "트레일러가 바닥에 놓여 있으나, 정차한 상태를 설명할 만한 견인 차량이 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 간판 텍스트('미추홀도서관, 인천') 대신 레퍼런스의 글씨('시립도서관')가 그대로 출력되었으나, 지정된 부감 앵글과 낡은 캠핑카(모터홈)의 묘사가 프롬프트의 지시를 충실히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "간판 텍스트 요구사항을 실패했으며, 캠핑카가 견인차 없는 트레일러 형태로 묘사되어 '도착해서 멈춰 선' 상황에 어울리지 않고 배경 원경이 레퍼런스에 비해 과장되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 도서관 입구와 주차된 캠핑카를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "레퍼런스와 일치하는 도서관 건물, 유리문 입구, 주차장 구획선이 올바르게 배치됨.",
        "entities": "비 오는 밤, 금이 가고 넝쿨이 얽힌 도서관 외벽, 지붕과 측면이 보이는 낡은 캠핑카(모터홈). 단, 간판 텍스트는 프롬프트와 달리 '시립도서관'으로 나타남.",
        "hard_violations": [],
        "physics": "캠핑카가 젖은 아스팔트 바닥에 안정적으로 정차해 있으며 빗방울과 바닥 반사가 자연스러움."
       },
       {
        "label": "B",
        "direction": "카메라가 도서관 건물과 트레일러를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "도서관 건물 구조는 유사하나, 좌측 원경의 도시 건물들이 레퍼런스보다 과장되게 가깝게 변경됨.",
        "entities": "비 오는 밤, 넝쿨이 얽힌 도서관 외벽. 운전석이 없는 트레일러 형태의 캠핑카. 간판은 역시 '시립도서관'으로 출력됨.",
        "hard_violations": [],
        "physics": "트레일러가 바닥에 놓여 있으나, 정차한 상태를 설명할 만한 견인 차량이 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "야간 비와 외벽의 균열·넝쿨은 맞지만, 차량 하부가 크게 잘려 캠핑카와 지면의 관계가 약하고 판독 가능한 간판이 최종 문자 금지 지시를 위반합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캠핑카의 운전석·측면·지붕과 접지 바퀴를 함께 보여 지정된 부감 전경에 더 충실하지만, 읽히는 간판 때문에 최종 문자 금지 지시는 충족하지 못합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물이나 조준 물체는 없습니다. 차량의 넓은 측면과 둥근 끝부분이 카메라를 향하지만 운전석이 없어 진행 방향은 확정하기 어렵습니다. 도서관 정면은 화면 오른쪽 가까운 곳에서 왼쪽 뒤로 물러나며, 간판의 글자 면은 카메라에 노출되어 있습니다.",
        "built_space": "참조의 2층 콘크리트 외벽, 왼쪽 수직 창탑 하나, 가로 창열 두 층, 중앙 오른쪽 유리 출입구 한 구역과 그 위 녹슨 간판 하나, 앞쪽 화단이 보입니다. 출입 계단은 차량에 상당 부분 가려집니다. 차량은 화면 하단을 크게 차지하고 하부가 프레임 밖으로 잘려, 요구한 차량과 지면의 관계보다 건물과 차량 상부가 강조됩니다.",
        "entities": "사람은 없으므로 불필요한 인물 추가는 없습니다. 낡고 얼룩진 숙박용 차량 한 대는 보이지만 동력 캠핑카인지 견인식 카라반인지 확인하기 어렵습니다. 외벽의 큰 균열과 빽빽한 마른 넝쿨, 비 내리는 밤과 젖은 포장은 일치합니다. 간판은 요청된 ‘미추홀도서관, 인천’이 아니라 참조의 ‘시립도서관’으로 읽히며, 마지막의 모든 문자 비가독화 지시에도 어긋납니다. 실내 누수는 확인되지 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "차량 바퀴와 지면 접점은 하단 크롭 밖이라 지지 상태를 직접 확인할 수 없지만, 공중에 떠 있다고 볼 시각적 근거도 없습니다. 지붕 설비는 지붕에 붙어 있고 간판은 출입구 구조에 고정되어 있습니다. 젖은 포장 반사는 가능한 모습이며 움직임이나 충돌은 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "인물이나 조준 물체는 없습니다. 캠핑카 운전석과 앞바퀴는 화면 왼쪽을 향하고 차량은 도서관 앞에 정지해 있습니다. 카메라는 비스듬한 위쪽에서 차량 측면과 지붕을 봅니다. 도서관 정면은 화면 상부에서 왼쪽 뒤로 물러나고 간판 전면은 카메라를 향합니다.",
        "built_space": "참조와 대응하는 왼쪽 창탑 하나, 두 층의 가로 창열, 중앙 오른쪽 유리 출입구 한 구역, 그 위 녹슨 간판 하나, 전면 화단과 출입 계단·난간이 보입니다. 차량은 오른쪽 아래 전경에 놓이고 왼쪽에는 넓은 젖은 주차장이 남아 건물·차량·지면의 크기 관계가 드러납니다. 차량 뒤쪽과 하부 일부는 잘렸지만 앞바퀴의 접지는 보입니다.",
        "entities": "인물은 없습니다. 운전실과 상부 침상 돌출부를 갖춘 낡은 캠핑카 한 대가 명확합니다. 외벽의 깊은 균열, 마른 넝쿨, 노후한 간판, 야간 빗줄기와 물 고인 포장이 보입니다. 간판은 ‘시립도서관’으로 읽혀 지정 명칭과 다르고 최종 문자 비가독화 지시도 어깁니다. 지붕의 물 고임은 보이지만 실내 누수까지 확인되지는 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "보이는 앞바퀴가 포장면에 닿아 차량을 지지하며 정차 상태가 자연스럽습니다. 지붕 설비와 난간은 차체에 부착되어 있고 물은 지붕과 포장면에 고여 있습니다. 출입구 간판은 건물 구조가 지지합니다. 떠 있는 물체나 불가능한 반사는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "야간 비와 외벽의 균열·넝쿨은 맞지만, 차량 하부가 크게 잘려 캠핑카와 지면의 관계가 약하고 판독 가능한 간판이 최종 문자 금지 지시를 위반합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캠핑카의 운전석·측면·지붕과 접지 바퀴를 함께 보여 지정된 부감 전경에 더 충실하지만, 읽히는 간판 때문에 최종 문자 금지 지시는 충족하지 못합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "인물이나 조준 물체는 없습니다. 차량의 넓은 측면과 둥근 끝부분이 카메라를 향하지만 운전석이 없어 진행 방향은 확정하기 어렵습니다. 도서관 정면은 화면 오른쪽 가까운 곳에서 왼쪽 뒤로 물러나며, 간판의 글자 면은 카메라에 노출되어 있습니다.",
        "built_space": "참조의 2층 콘크리트 외벽, 왼쪽 수직 창탑 하나, 가로 창열 두 층, 중앙 오른쪽 유리 출입구 한 구역과 그 위 녹슨 간판 하나, 앞쪽 화단이 보입니다. 출입 계단은 차량에 상당 부분 가려집니다. 차량은 화면 하단을 크게 차지하고 하부가 프레임 밖으로 잘려, 요구한 차량과 지면의 관계보다 건물과 차량 상부가 강조됩니다.",
        "entities": "사람은 없으므로 불필요한 인물 추가는 없습니다. 낡고 얼룩진 숙박용 차량 한 대는 보이지만 동력 캠핑카인지 견인식 카라반인지 확인하기 어렵습니다. 외벽의 큰 균열과 빽빽한 마른 넝쿨, 비 내리는 밤과 젖은 포장은 일치합니다. 간판은 요청된 ‘미추홀도서관, 인천’이 아니라 참조의 ‘시립도서관’으로 읽히며, 마지막의 모든 문자 비가독화 지시에도 어긋납니다. 실내 누수는 확인되지 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "차량 바퀴와 지면 접점은 하단 크롭 밖이라 지지 상태를 직접 확인할 수 없지만, 공중에 떠 있다고 볼 시각적 근거도 없습니다. 지붕 설비는 지붕에 붙어 있고 간판은 출입구 구조에 고정되어 있습니다. 젖은 포장 반사는 가능한 모습이며 움직임이나 충돌은 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "인물이나 조준 물체는 없습니다. 캠핑카 운전석과 앞바퀴는 화면 왼쪽을 향하고 차량은 도서관 앞에 정지해 있습니다. 카메라는 비스듬한 위쪽에서 차량 측면과 지붕을 봅니다. 도서관 정면은 화면 상부에서 왼쪽 뒤로 물러나고 간판 전면은 카메라를 향합니다.",
        "built_space": "참조와 대응하는 왼쪽 창탑 하나, 두 층의 가로 창열, 중앙 오른쪽 유리 출입구 한 구역, 그 위 녹슨 간판 하나, 전면 화단과 출입 계단·난간이 보입니다. 차량은 오른쪽 아래 전경에 놓이고 왼쪽에는 넓은 젖은 주차장이 남아 건물·차량·지면의 크기 관계가 드러납니다. 차량 뒤쪽과 하부 일부는 잘렸지만 앞바퀴의 접지는 보입니다.",
        "entities": "인물은 없습니다. 운전실과 상부 침상 돌출부를 갖춘 낡은 캠핑카 한 대가 명확합니다. 외벽의 깊은 균열, 마른 넝쿨, 노후한 간판, 야간 빗줄기와 물 고인 포장이 보입니다. 간판은 ‘시립도서관’으로 읽혀 지정 명칭과 다르고 최종 문자 비가독화 지시도 어깁니다. 지붕의 물 고임은 보이지만 실내 누수까지 확인되지는 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "보이는 앞바퀴가 포장면에 닿아 차량을 지지하며 정차 상태가 자연스럽습니다. 지붕 설비와 난간은 차체에 부착되어 있고 물은 지붕과 포장면에 고여 있습니다. 출입구 간판은 건물 구조가 지지합니다. 떠 있는 물체나 불가능한 반사는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.417
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "B": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ],
    "A": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "요구된 간판 텍스트('미추홀도서관, 인천') 대신 레퍼런스의 글씨('시립도서관')가 그대로 출력되었으나, 지정된 부감 앵글과 낡은 캠핑카(모터홈)의 묘사가 프롬프트의 지시를 충실히 따랐습니다.  ★위반: [gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "간판 텍스트 요구사항을 실패했으며, 캠핑카가 견인차 없는 트레일러 형태로 묘사되어 '도착해서 멈춰 선' 상황에 어울리지 않고 배경 원경이 레퍼런스에 비해 과장되었습니다.  ★위반: [gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B03.png",
    "asset_id": "1b1737dc-7622-466c-bcf0-357121348d7e",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-99b3-7e03-b45c-1f8a5ae18f91",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S46sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:41:29.922748+00:00",
  "fingerprint": "3e22f9109e14f9305fcb372f0ac4f3480faf25d2dd9b9e9ecb5f513bd3810291",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S46sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S46sh1_sel.png",
  "source_sha256": "49341a5cef17c7fc9ca0be998e239607602606e8e8633428bf3649835f2a9d9f",
  "file": "S46sh1_cine.png",
  "staged_sha256": "2da0014c8e824495f79baea7bad1eecafb6be941e913bea23364a3b7723e5487",
  "latency_ms": 10363
 },
 "S46sh17::signage": {
  "fp": "affc0ecbd612238c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9e4648ab22ad1f33": {
  "subjects": [
   {
    "subject_native": "인천 미추홀도서관 열람실 및 천창",
    "search_terms_native": [
     "미추홀도서관 열람실",
     "미추홀도서관 내부",
     "미추홀도서관 종합자료실",
     "인천 미추홀도서관 천창"
    ],
    "language_lock_native": "모든 검색어는 반드시 한국어로만 작성해야 하며 다른 언어로 번역하거나 추가해서는 안 됩니다.",
    "reason_ko": "인천 미추홀도서관 특유의 현대식 천창 구조와 한국 공공도서관 열람실 형태 대신 서구식 고전 도서관 폐허로 왜곡될 가능성이 큼."
   }
  ],
  "subject_text": "미추홀도서관 열람실\n먼지와 거미줄이 가득한 버려진 열람실. 부서진 책장과 낡은 책, 고물 컴퓨터가 남아 있고 깨진 천창으로 잿빛 햇살이 들어온다.",
  "identity": "canonical",
  "scope_id": "L202",
  "scope_role": "location_interior",
  "scope_sha": "b3dc6c070f833894"
 },
 "era_fail::06265b5c57a63a76": {
  "stage": "research",
  "subject": "인천 미추홀도서관 열람실 및 천창",
  "terms": [
   "미추홀도서관 열람실",
   "미추홀도서관 내부",
   "미추홀도서관 종합자료실",
   "인천 미추홀도서관 천창"
  ],
  "status": "no_usable",
  "queries": [
   [
    "인천 미추홀도서관 내부 종합자료실 열람실 천창",
    "미추홀도서관 천창 실내 건축"
   ],
   [
    "\"미추홀도서관\" \"천창\"",
    "\"미추홀도서관\" \"종합자료실\" \"층\""
   ]
  ],
  "candidate_urls": [
   "https://cdn.welfarehello.com/naver-blog/production/iseogu/2023-05/223097596981/iseogu_223097596981_10.jpg",
   "https://biz.namdong.go.kr/images/bbs/archive/2021/michuholdoseogwan_kkumnamuteo1_aywc0467.jpg",
   "https://cdn.welfarehello.com/naver-blog/production/tong_namgu/2024-07/223524882298/tong_namgu_223524882298_14.jpg",
   "https://cdn.welfarehello.com/naver-blog/production/namdongdistrict/2023-07/223154007596/namdongdistrict_223154007596_5.jpg"
  ],
  "coarse": {
   "eligible": [],
   "chosen_index": 0,
   "reason": "종류·보임·기준을 다 만족하는 후보가 없다 — 이 라운드에선 안 고른다",
   "single_judge": true,
   "rejected_judges": {}
  },
  "verdicts": [
   {
    "index": 1,
    "object_type_match": "yes",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 95
   },
   {
    "index": 2,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 30
   },
   {
    "index": 3,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 40
   },
   {
    "index": 4,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 40
   }
  ],
  "chosen_reason_ko": "",
  "attempts": 12
 },
 "S46sh17::bgfirst_bg": {
  "input_fingerprint": "45002603facd4e3e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17__bgfirst_bg.png",
  "asset_id": "423c83f7-5368-4b62-8c21-328accdd47f4",
  "input_asset_ids": [
   "c2cd0399-5f87-4274-8987-1f4261def294",
   "247ce4f2-a8a9-4191-9f73-361363704c8d"
  ]
 },
 "S46sh17": {
  "input_fingerprint": "0f203e4465dfee82",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The abandoned reading room remains dusty and cobwebbed, with disordered shelves and books. Opened canned food and the scavenged bags remain at the makeshift eating place, while the Ubik book is open to its next-generation robot chapter. 현우: He remains at the eating place with facial bruises and the untreated leg wound. The contact card is still inside his shoe, and the scavenged disposable phone remains among his belongings. 찰리: He holds the Ubik book open for display, retaining his worn body, earlier disguise and store blanket.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리, 현우 right now, so 찰리, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The abandoned reading room remains dusty and cobwebbed, with disordered shelves and books. Opened canned food and the scavenged bags remain at the makeshift eating place, while the Ubik book is open to its next-generation robot chapter. 현우: He remains at the eating place with facial bruises and the untreated leg wound. The contact card is still inside his shoe, and the scavenged disposable phone remains among his belongings. 찰리: He holds the Ubik book open for display, retaining his worn body, earlier disguise and store blanket.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리, 현우 right now, so 찰리, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The abandoned reading room remains dusty and cobwebbed, with disordered shelves and books. Opened canned food and the scavenged bags remain at the makeshift eating place, while the Ubik book is open to its next-generation robot chapter. 현우: He remains at the eating place with facial bruises and the untreated leg wound. The contact card is still inside his shoe, and the scavenged disposable phone remains among his belongings. 찰리: He holds the Ubik book open for display, retaining his worn body, earlier disguise and store blanket.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리, 현우 right now, so 찰리, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17__bgfirst_bg.png",
     "asset_id": "423c83f7-5368-4b62-8c21-328accdd47f4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S46sh17.png",
     "asset_id": "c2cd0399-5f87-4274-8987-1f4261def294",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 유빅사 책: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1126012>",
     "asset_id": "6f976831-f189-4330-b56f-fdd0f9951db8",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B04.png",
     "asset_id": "247ce4f2-a8a9-4191-9f73-361363704c8d",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 유빅사 책: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1126012>",
     "asset_id": "6f976831-f189-4330-b56f-fdd0f9951db8",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 강하게 뻗은 오른팔이 찰리가 들고 있는 책의 페이지를 정확히 향해 타격하고 있습니다.",
    "built_space": "어두운 밤의 버려진 도서관 열람실로, 배경에 낡은 책장들이 있고 앞쪽 책상 위에는 통조림 캔, 가방, 스마트폰 등 레퍼런스의 소품들이 알맞은 위치에 배치되어 있습니다.",
    "entities": "현우는 상처 입은 얼굴과 회색 셔츠 등 레퍼런스와 일치하는 인상착의를 가집니다. 찰리는 마스크와 담요를 덮고 있으나 비정상적인 손 개수를 가집니다. 책은 레퍼런스에 제시된 고릴라 로봇 일러스트를 정확히 보여줍니다.",
    "hard_violations": [
     "[gemini-pro] 찰리의 신체 구조 오류: 아래쪽에 놓인 거대한 팔 외에, 책을 쥐고 있는 작은 사이즈의 로봇 손 2개가 추가로 생성되어 총 3개의 손과 팔이 존재함.",
     "[gpt-high] 책의 장 제목과 통조림 표면에 판독 가능한 글자가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
    ],
    "physics": "현우의 주먹이 책 페이지에 맞닿아 접히는 물리적인 타격 순간이 잘 묘사되어 있습니다. 그러나 책을 들고 있는 찰리의 두 손은 아래에 놓인 거대한 원래 팔과 완전히 분리된 엉뚱한 크기의 여분의 손으로, 물리적 지탱 구조가 불가능합니다."
   },
   {
    "label": "B",
    "direction": "현우의 강렬한 시선과 주먹이 허공에 흩날리는 책 상단을 향해 뻗어 타격하고 있습니다.",
    "built_space": "먼지와 거미줄이 낀 도서관 열람실 내부이며, 뒤쪽의 책장과 앞쪽 책상의 통조림, 가방 세팅이 레퍼런스 공간의 특징을 잘 반영하고 있습니다.",
    "entities": "현우는 레퍼런스의 인물 및 복장과 일치합니다. 찰리는 고릴라형 로봇의 얼굴과 담요를 가졌으나 구조가 심하게 망가졌습니다. 책은 텍스트만 보일 뿐 지정된 고릴라 일러스트가 확인되지 않습니다.",
    "hard_violations": [
     "[gemini-pro] 사물 중복: 타격받고 있는 상단의 책과 별개로, 찰리의 아래쪽 손이 쥐고 있는 온전한 책이 화면에 동시에 존재함.",
     "[gemini-pro] 물리적으로 불가능한 신체 구조: 찰리의 상단 팔과 하단 팔이 구조적으로 연결되지 않고 해부학적으로 완전히 무너져 있음.",
     "[gpt-high] 하나여야 할 열린 책이 찰리가 받치는 아래 책과 현우가 타격하는 위 책으로 분리·중복되어 핵심 소품과 접촉 관계를 위반한다.",
     "[gpt-high] 통조림에 '참치' 등 판독 가능한 글자가 노출되어 읽을 수 있는 글자 금지 조건을 위반한다."
    ],
    "physics": "주먹의 타격으로 책이 심하게 휘어지고 움직이는 역동성이 묘사되었습니다. 하지만 현우의 팔 아래로 온전한 책을 받치고 있는 찰리의 하단 로봇 팔과, 타격받는 상단 책을 잡고 있는 기형적인 상단 로봇 팔이 물리적인 일관성 없이 분리되어 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "타격 순간과 지정된 로봇 일러스트 책의 묘사는 우수하나, 찰리에게 여분의 작은 로봇 손 2개가 추가로 생성되는 치명적인 신체 구조 오류가 발생하여 사용할 수 없습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "타격의 역동성은 표현되었으나 지정된 책의 일러스트가 누락되었으며, 화면 내에 책이 2개로 중복 생성되고 로봇의 팔 구조가 무너지는 심각한 물리적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 강하게 뻗은 오른팔이 찰리가 들고 있는 책의 페이지를 정확히 향해 타격하고 있습니다.",
        "built_space": "어두운 밤의 버려진 도서관 열람실로, 배경에 낡은 책장들이 있고 앞쪽 책상 위에는 통조림 캔, 가방, 스마트폰 등 레퍼런스의 소품들이 알맞은 위치에 배치되어 있습니다.",
        "entities": "현우는 상처 입은 얼굴과 회색 셔츠 등 레퍼런스와 일치하는 인상착의를 가집니다. 찰리는 마스크와 담요를 덮고 있으나 비정상적인 손 개수를 가집니다. 책은 레퍼런스에 제시된 고릴라 로봇 일러스트를 정확히 보여줍니다.",
        "hard_violations": [
         "찰리의 신체 구조 오류: 아래쪽에 놓인 거대한 팔 외에, 책을 쥐고 있는 작은 사이즈의 로봇 손 2개가 추가로 생성되어 총 3개의 손과 팔이 존재함."
        ],
        "physics": "현우의 주먹이 책 페이지에 맞닿아 접히는 물리적인 타격 순간이 잘 묘사되어 있습니다. 그러나 책을 들고 있는 찰리의 두 손은 아래에 놓인 거대한 원래 팔과 완전히 분리된 엉뚱한 크기의 여분의 손으로, 물리적 지탱 구조가 불가능합니다."
       },
       {
        "label": "B",
        "direction": "현우의 강렬한 시선과 주먹이 허공에 흩날리는 책 상단을 향해 뻗어 타격하고 있습니다.",
        "built_space": "먼지와 거미줄이 낀 도서관 열람실 내부이며, 뒤쪽의 책장과 앞쪽 책상의 통조림, 가방 세팅이 레퍼런스 공간의 특징을 잘 반영하고 있습니다.",
        "entities": "현우는 레퍼런스의 인물 및 복장과 일치합니다. 찰리는 고릴라형 로봇의 얼굴과 담요를 가졌으나 구조가 심하게 망가졌습니다. 책은 텍스트만 보일 뿐 지정된 고릴라 일러스트가 확인되지 않습니다.",
        "hard_violations": [
         "사물 중복: 타격받고 있는 상단의 책과 별개로, 찰리의 아래쪽 손이 쥐고 있는 온전한 책이 화면에 동시에 존재함.",
         "물리적으로 불가능한 신체 구조: 찰리의 상단 팔과 하단 팔이 구조적으로 연결되지 않고 해부학적으로 완전히 무너져 있음."
        ],
        "physics": "주먹의 타격으로 책이 심하게 휘어지고 움직이는 역동성이 묘사되었습니다. 하지만 현우의 팔 아래로 온전한 책을 받치고 있는 찰리의 하단 로봇 팔과, 타격받는 상단 책을 잡고 있는 기형적인 상단 로봇 팔이 물리적인 일관성 없이 분리되어 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "타격 순간과 지정된 로봇 일러스트 책의 묘사는 우수하나, 찰리에게 여분의 작은 로봇 손 2개가 추가로 생성되는 치명적인 신체 구조 오류가 발생하여 사용할 수 없습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "타격의 역동성은 표현되었으나 지정된 책의 일러스트가 누락되었으며, 화면 내에 책이 2개로 중복 생성되고 로봇의 팔 구조가 무너지는 심각한 물리적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 강하게 뻗은 오른팔이 찰리가 들고 있는 책의 페이지를 정확히 향해 타격하고 있습니다.",
        "built_space": "어두운 밤의 버려진 도서관 열람실로, 배경에 낡은 책장들이 있고 앞쪽 책상 위에는 통조림 캔, 가방, 스마트폰 등 레퍼런스의 소품들이 알맞은 위치에 배치되어 있습니다.",
        "entities": "현우는 상처 입은 얼굴과 회색 셔츠 등 레퍼런스와 일치하는 인상착의를 가집니다. 찰리는 마스크와 담요를 덮고 있으나 비정상적인 손 개수를 가집니다. 책은 레퍼런스에 제시된 고릴라 로봇 일러스트를 정확히 보여줍니다.",
        "hard_violations": [
         "찰리의 신체 구조 오류: 아래쪽에 놓인 거대한 팔 외에, 책을 쥐고 있는 작은 사이즈의 로봇 손 2개가 추가로 생성되어 총 3개의 손과 팔이 존재함."
        ],
        "physics": "현우의 주먹이 책 페이지에 맞닿아 접히는 물리적인 타격 순간이 잘 묘사되어 있습니다. 그러나 책을 들고 있는 찰리의 두 손은 아래에 놓인 거대한 원래 팔과 완전히 분리된 엉뚱한 크기의 여분의 손으로, 물리적 지탱 구조가 불가능합니다."
       },
       {
        "label": "B",
        "direction": "현우의 강렬한 시선과 주먹이 허공에 흩날리는 책 상단을 향해 뻗어 타격하고 있습니다.",
        "built_space": "먼지와 거미줄이 낀 도서관 열람실 내부이며, 뒤쪽의 책장과 앞쪽 책상의 통조림, 가방 세팅이 레퍼런스 공간의 특징을 잘 반영하고 있습니다.",
        "entities": "현우는 레퍼런스의 인물 및 복장과 일치합니다. 찰리는 고릴라형 로봇의 얼굴과 담요를 가졌으나 구조가 심하게 망가졌습니다. 책은 텍스트만 보일 뿐 지정된 고릴라 일러스트가 확인되지 않습니다.",
        "hard_violations": [
         "사물 중복: 타격받고 있는 상단의 책과 별개로, 찰리의 아래쪽 손이 쥐고 있는 온전한 책이 화면에 동시에 존재함.",
         "물리적으로 불가능한 신체 구조: 찰리의 상단 팔과 하단 팔이 구조적으로 연결되지 않고 해부학적으로 완전히 무너져 있음."
        ],
        "physics": "주먹의 타격으로 책이 심하게 휘어지고 움직이는 역동성이 묘사되었습니다. 하지만 현우의 팔 아래로 온전한 책을 받치고 있는 찰리의 하단 로봇 팔과, 타격받는 상단 책을 잡고 있는 기형적인 상단 로봇 팔이 물리적인 일관성 없이 분리되어 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우의 타격 대상과 찰리가 받쳐 든 책이 별개의 두 책으로 갈라져 핵심 접촉 관계를 위반하며, 통조림의 글자도 읽힌다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리가 잡은 로봇 책에 현우의 주먹이 닿는 미디엄 숏은 정확하지만, 책의 장 제목과 통조림 글자가 읽혀 무문자 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 오른쪽으로 팔을 뻗어 위쪽에서 펼쳐진 책을 주먹으로 치며, 시선도 그 방향이다. 찰리의 얼굴은 접촉 부근을 향하지만 두 손은 아래쪽의 별도 책을 받친다. 따라서 타격이 찰리 손에 있는 책으로 연결되지 않는다. 아래 책의 인쇄면은 현우와 카메라 쪽으로 비스듬히 열려 있다.",
        "built_space": "전경에 식사 탁자 하나와 일부 의자 등받이가 있고, 왼쪽 벽면 서가와 중앙의 높은 서가 두 줄, 오른쪽 창 아래 낮은 서가 일부가 보인다. 오른쪽 검은 창틀과 낡은 벽, 천장 등기구는 장소 참조와 부합한다. 현우는 탁자 건너편 왼쪽, 찰리는 오른쪽 가까이에 있다. 서가와 거미줄이 비교적 선명하여 요구한 부드러운 배경보다 주목도가 높다.",
        "entities": "등장 개체는 현우와 찰리 둘이다. 현우는 검은 머리의 젊은 동아시아계 남성으로 회색 셔츠와 얼굴 상처가 보이지만 참조보다 다소 성숙한 인상이다. 찰리는 샌드 베이지 장갑, 흰 기계식 얼굴, 큰 팔과 회색 줄무늬 담요를 유지한다. 책은 하나가 아니라 두 개처럼 분리되어 보이고, 로봇 삽화는 확인되지 않는다. 탁자에는 가방, 통조림, 금속 컵과 천이 있으며 통조림의 '참치'가 읽힌다. 다리 상처와 신발 속 카드는 화면 밖이라 판단할 수 없다.",
        "hard_violations": [
         "하나여야 할 열린 책이 찰리가 받치는 아래 책과 현우가 타격하는 위 책으로 분리·중복되어 핵심 소품과 접촉 관계를 위반한다.",
         "통조림에 '참치' 등 판독 가능한 글자가 노출되어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "현우의 몸통 전진과 뻗은 팔은 타격 동작으로 가능하며, 하체는 탁자에 가려져 있다. 아래 책은 찰리의 두 손이 지지한다. 위 책에는 현우의 주먹이 접촉하고 페이지가 날리므로 충격에 의한 순간 운동으로는 해석할 수 있지만, 찰리의 손이 그 책을 계속 잡고 있다는 지지는 보이지 않는다. 담요는 찰리의 어깨에, 식사 물품은 탁자에 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "현우의 시선과 오른팔이 찰리가 든 책의 로봇 삽화 쪽으로 향하고 주먹이 해당 페이지에 직접 닿는다. 찰리도 책과 접촉점을 내려다본다. 열린 인쇄면은 왼쪽의 현우를 향하며 카메라에는 비스듬히 보여 소품 방향이 맞는다.",
        "built_space": "식사 탁자 하나가 전경에 있고 의자 등받이 일부가 왼쪽과 아래에 보인다. 왼쪽 벽면 서가와 중앙의 높은 서가 두 줄, 뒤쪽 서가 일부, 오른쪽 창과 낡은 기둥이 참조 공간을 유지한다. 현우는 탁자 건너편 왼쪽, 찰리는 오른쪽에 있으며 책은 두 인물 사이에 놓인다. 미디엄 숏은 충족하지만 배경 서가가 상당히 선명하고 탁자의 비중도 크다.",
        "entities": "현우는 앳된 동아시아계 남성의 얼굴, 검은 머리, 회색 셔츠와 얼굴 멍을 갖추어 참조에 가깝다. 소매는 걷혀 있다. 찰리의 육중한 기계 몸체, 샌드 베이지 장갑판, 흰 마스크형 얼굴과 줄무늬 담요가 유지된다. 열린 책 한 권에 참조와 유사한 고릴라형 로봇 삽화가 있다. 다만 장 제목의 '제3장'과 '차세대'가 읽힌다. 탁자에는 가방, 통조림, 컵, 천과 휴대전화가 있고 통조림에도 판독 가능한 글자가 있다. 가려진 다리와 신발 속 물품은 평가할 수 없다.",
        "hard_violations": [
         "책의 장 제목과 통조림 표면에 판독 가능한 글자가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "현우의 어깨에서 주먹까지 이어지는 팔이 책에 닿고, 찰리의 손가락은 책 아래쪽을 받쳐 충격 중에도 책이 손에 남아 있다. 들린 페이지는 타격으로 벌어진 순간으로 해석 가능하다. 다른 큰 손은 책 아래 가까이에 있으며 중복된 팔은 보이지 않는다. 하체 접지는 프레임과 탁자에 가려져 있지만 공중 부유의 징후는 없다. 담요는 어깨에 걸리고 휴대전화와 식사 물품은 탁자에 놓여 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우의 타격 대상과 찰리가 받쳐 든 책이 별개의 두 책으로 갈라져 핵심 접촉 관계를 위반하며, 통조림의 글자도 읽힌다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리가 잡은 로봇 책에 현우의 주먹이 닿는 미디엄 숏은 정확하지만, 책의 장 제목과 통조림 글자가 읽혀 무문자 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 오른쪽으로 팔을 뻗어 위쪽에서 펼쳐진 책을 주먹으로 치며, 시선도 그 방향이다. 찰리의 얼굴은 접촉 부근을 향하지만 두 손은 아래쪽의 별도 책을 받친다. 따라서 타격이 찰리 손에 있는 책으로 연결되지 않는다. 아래 책의 인쇄면은 현우와 카메라 쪽으로 비스듬히 열려 있다.",
        "built_space": "전경에 식사 탁자 하나와 일부 의자 등받이가 있고, 왼쪽 벽면 서가와 중앙의 높은 서가 두 줄, 오른쪽 창 아래 낮은 서가 일부가 보인다. 오른쪽 검은 창틀과 낡은 벽, 천장 등기구는 장소 참조와 부합한다. 현우는 탁자 건너편 왼쪽, 찰리는 오른쪽 가까이에 있다. 서가와 거미줄이 비교적 선명하여 요구한 부드러운 배경보다 주목도가 높다.",
        "entities": "등장 개체는 현우와 찰리 둘이다. 현우는 검은 머리의 젊은 동아시아계 남성으로 회색 셔츠와 얼굴 상처가 보이지만 참조보다 다소 성숙한 인상이다. 찰리는 샌드 베이지 장갑, 흰 기계식 얼굴, 큰 팔과 회색 줄무늬 담요를 유지한다. 책은 하나가 아니라 두 개처럼 분리되어 보이고, 로봇 삽화는 확인되지 않는다. 탁자에는 가방, 통조림, 금속 컵과 천이 있으며 통조림의 '참치'가 읽힌다. 다리 상처와 신발 속 카드는 화면 밖이라 판단할 수 없다.",
        "hard_violations": [
         "하나여야 할 열린 책이 찰리가 받치는 아래 책과 현우가 타격하는 위 책으로 분리·중복되어 핵심 소품과 접촉 관계를 위반한다.",
         "통조림에 '참치' 등 판독 가능한 글자가 노출되어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "현우의 몸통 전진과 뻗은 팔은 타격 동작으로 가능하며, 하체는 탁자에 가려져 있다. 아래 책은 찰리의 두 손이 지지한다. 위 책에는 현우의 주먹이 접촉하고 페이지가 날리므로 충격에 의한 순간 운동으로는 해석할 수 있지만, 찰리의 손이 그 책을 계속 잡고 있다는 지지는 보이지 않는다. 담요는 찰리의 어깨에, 식사 물품은 탁자에 지지되어 있다."
       },
       {
        "label": "A",
        "direction": "현우의 시선과 오른팔이 찰리가 든 책의 로봇 삽화 쪽으로 향하고 주먹이 해당 페이지에 직접 닿는다. 찰리도 책과 접촉점을 내려다본다. 열린 인쇄면은 왼쪽의 현우를 향하며 카메라에는 비스듬히 보여 소품 방향이 맞는다.",
        "built_space": "식사 탁자 하나가 전경에 있고 의자 등받이 일부가 왼쪽과 아래에 보인다. 왼쪽 벽면 서가와 중앙의 높은 서가 두 줄, 뒤쪽 서가 일부, 오른쪽 창과 낡은 기둥이 참조 공간을 유지한다. 현우는 탁자 건너편 왼쪽, 찰리는 오른쪽에 있으며 책은 두 인물 사이에 놓인다. 미디엄 숏은 충족하지만 배경 서가가 상당히 선명하고 탁자의 비중도 크다.",
        "entities": "현우는 앳된 동아시아계 남성의 얼굴, 검은 머리, 회색 셔츠와 얼굴 멍을 갖추어 참조에 가깝다. 소매는 걷혀 있다. 찰리의 육중한 기계 몸체, 샌드 베이지 장갑판, 흰 마스크형 얼굴과 줄무늬 담요가 유지된다. 열린 책 한 권에 참조와 유사한 고릴라형 로봇 삽화가 있다. 다만 장 제목의 '제3장'과 '차세대'가 읽힌다. 탁자에는 가방, 통조림, 컵, 천과 휴대전화가 있고 통조림에도 판독 가능한 글자가 있다. 가려진 다리와 신발 속 물품은 평가할 수 없다.",
        "hard_violations": [
         "책의 장 제목과 통조림 표면에 판독 가능한 글자가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "현우의 어깨에서 주먹까지 이어지는 팔이 책에 닿고, 찰리의 손가락은 책 아래쪽을 받쳐 충격 중에도 책이 손에 남아 있다. 들린 페이지는 타격으로 벌어진 순간으로 해석 가능하다. 다른 큰 손은 책 아래 가까이에 있으며 중복된 팔은 보이지 않는다. 하체 접지는 프레임과 탁자에 가려져 있지만 공중 부유의 징후는 없다. 담요는 어깨에 걸리고 휴대전화와 식사 물품은 탁자에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.167
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.917
   },
   "violations": {
    "A": [
     "[gemini-pro] 찰리의 신체 구조 오류: 아래쪽에 놓인 거대한 팔 외에, 책을 쥐고 있는 작은 사이즈의 로봇 손 2개가 추가로 생성되어 총 3개의 손과 팔이 존재함.",
     "[gpt-high] 책의 장 제목과 통조림 표면에 판독 가능한 글자가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 사물 중복: 타격받고 있는 상단의 책과 별개로, 찰리의 아래쪽 손이 쥐고 있는 온전한 책이 화면에 동시에 존재함.",
     "[gemini-pro] 물리적으로 불가능한 신체 구조: 찰리의 상단 팔과 하단 팔이 구조적으로 연결되지 않고 해부학적으로 완전히 무너져 있음.",
     "[gpt-high] 하나여야 할 열린 책이 찰리가 받치는 아래 책과 현우가 타격하는 위 책으로 분리·중복되어 핵심 소품과 접촉 관계를 위반한다.",
     "[gpt-high] 통조림에 '참치' 등 판독 가능한 글자가 노출되어 읽을 수 있는 글자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 917
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "타격 순간과 지정된 로봇 일러스트 책의 묘사는 우수하나, 찰리에게 여분의 작은 로봇 손 2개가 추가로 생성되는 치명적인 신체 구조 오류가 발생하여 사용할 수 없습니다.  ★위반: [gemini-pro] 찰리의 신체 구조 오류: 아래쪽에 놓인 거대한 팔 외에, 책을 쥐고 있는 작은 사이즈의 로봇 손 2개가 추가로 생성되어 총 3개의 손과 팔이 존재함. / [gpt-high] 책의 장 제목과 통조림 표면에 판독 가능한 글자가 있어 읽을 수 있는 글자 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 917,
    "verdict_ko": "타격의 역동성은 표현되었으나 지정된 책의 일러스트가 누락되었으며, 화면 내에 책이 2개로 중복 생성되고 로봇의 팔 구조가 무너지는 심각한 물리적 오류가 있습니다.  ★위반: [gemini-pro] 사물 중복: 타격받고 있는 상단의 책과 별개로, 찰리의 아래쪽 손이 쥐고 있는 온전한 책이 화면에 동시에 존재함. / [gemini-pro] 물리적으로 불가능한 신체 구조: 찰리의 상단 팔과 하단 팔이 구조적으로 연결되지 않고 해부학적으로 완전히 무너져 있음. / [gpt-high] 하나여야 할 열린 책이 찰리가 받치는 아래 책과 현우가 타격하는 위 책으로 분리·중복되어 핵심 소품과 접촉 관계를 위반한다. / [gpt-high] 통조림에 '참치' 등 판독 가능한 글자가 노출되어 읽을 수 있는 글자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B04.png",
    "asset_id": "247ce4f2-a8a9-4191-9f73-361363704c8d",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 유빅사 책: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:1126012>",
    "asset_id": "6f976831-f189-4330-b56f-fdd0f9951db8",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40b-9b75-7b1c-8b34-5df83fc40e85",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17__bgfirst_bg.png",
   "bg_asset_id": "423c83f7-5368-4b62-8c21-328accdd47f4",
   "bg_record_key": "S46sh17::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S46sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:06:42.667597+00:00",
  "fingerprint": "93e04cff03e3b80ac2d77daeb2a4781ceb96cd82a3573ca6710f126b088b006e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S46sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S46sh17_sel.png",
  "source_sha256": "1697386b01374656f9d048bb23d8c83f58209cf8bbe8921e79e55c485eae1882",
  "file": "S46sh17_cine.png",
  "staged_sha256": "c046cd637043df41cae326c157da640a9a36f1080255fecd36f159b8b9ef89de",
  "latency_ms": 11065
 },
 "S46sh28::signage": {
  "fp": "9d890661c8b50cc9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S46sh28": {
  "input_fingerprint": "28ba053293b81967",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면의 불빛이 비친 채 입 모양을 크고 둥글게 벌리고 집중하는 찰리의 낡은 금속 얼굴.\n\nLOCATION (lock): At an old computer among the abandoned library's dusty bookshelves, with monitor light illuminating the otherwise dark area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old computer monitor (In use during 찰리's speech practice) — Only a narrow side edge is visible; the display face points toward 찰리 outside the crop; used as Locates his attention without showing invented screen content; Library books and shelving (Dust-covered) — Indistinct fragments remain behind his head; used as Retains the abandoned reading-room context at shallow focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the monitor's established light articulate 찰리's worn metal face against the subdued room, without assigning the display an unsupported color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The dusty library shelves and books remain disordered, and an old computer is now running. Boxes and books have been arranged as makeshift bedding in one corner. 찰리: He is awake at the old computer, practicing speech after gathering data. His worn body and retained disguise are unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면의 불빛이 비친 채 입 모양을 크고 둥글게 벌리고 집중하는 찰리의 낡은 금속 얼굴.\n\nLOCATION (lock): At an old computer among the abandoned library's dusty bookshelves, with monitor light illuminating the otherwise dark area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old computer monitor (In use during 찰리's speech practice) — Only a narrow side edge is visible; the display face points toward 찰리 outside the crop; used as Locates his attention without showing invented screen content; Library books and shelving (Dust-covered) — Indistinct fragments remain behind his head; used as Retains the abandoned reading-room context at shallow focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the monitor's established light articulate 찰리's worn metal face against the subdued room, without assigning the display an unsupported color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The dusty library shelves and books remain disordered, and an old computer is now running. Boxes and books have been arranged as makeshift bedding in one corner. 찰리: He is awake at the old computer, practicing speech after gathering data. His worn body and retained disguise are unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면의 불빛이 비친 채 입 모양을 크고 둥글게 벌리고 집중하는 찰리의 낡은 금속 얼굴.\n\nLOCATION (lock): At an old computer among the abandoned library's dusty bookshelves, with monitor light illuminating the otherwise dark area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old computer monitor (In use during 찰리's speech practice) — Only a narrow side edge is visible; the display face points toward 찰리 outside the crop; used as Locates his attention without showing invented screen content; Library books and shelving (Dust-covered) — Indistinct fragments remain behind his head; used as Retains the abandoned reading-room context at shallow focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the monitor's established light articulate 찰리's worn metal face against the subdued room, without assigning the display an unsupported color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The dusty library shelves and books remain disordered, and an old computer is now running. Boxes and books have been arranged as makeshift bedding in one corner. 찰리: He is awake at the old computer, practicing speech after gathering data. His worn body and retained disguise are unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선이 화면 오른쪽에 위치한 모니터를 향하고 있습니다.",
    "built_space": "흐릿하게 처리된 낡은 도서관 서가가 배경에 있고, 오른쪽 전면에 구형 모니터의 측면이 배치되어 있습니다.",
    "entities": "레퍼런스와 동일한 외형의 찰리가 스트라이프 패턴의 담요를 두르고 있으며, 지시문대로 입을 벌리고 있습니다.",
    "hard_violations": [],
    "physics": "찰리가 책상 앞에 자연스럽게 위치하여 모니터를 바라보는 자세가 안정적입니다."
   },
   {
    "label": "B",
    "direction": "찰리의 시선과 얼굴 방향이 오른쪽의 모니터를 향하고 있습니다.",
    "built_space": "배경에 도서관 서가와 창문이 흐릿하게 보이며, 우측에 모니터 측면이 자리 잡고 있습니다.",
    "entities": "찰리의 로봇 외형과 입을 벌린 표정은 나타나지만, 이전 샷에서 지정된 줄무늬 담요 위장이 없습니다.",
    "hard_violations": [],
    "physics": "모니터 앞에 위치한 상반신의 자세가 안정적으로 유지되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍, 모니터 불빛에 비친 얼굴, 크게 벌린 입, 그리고 이전 샷에서 고정된 담요 위장을 모두 정확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "표정과 조명, 배경 구성은 준수하나, 이전 샷 및 레퍼런스에서 필수적으로 유지해야 하는 담요 위장이 완전히 누락되어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선이 화면 오른쪽에 위치한 모니터를 향하고 있습니다.",
        "built_space": "흐릿하게 처리된 낡은 도서관 서가가 배경에 있고, 오른쪽 전면에 구형 모니터의 측면이 배치되어 있습니다.",
        "entities": "레퍼런스와 동일한 외형의 찰리가 스트라이프 패턴의 담요를 두르고 있으며, 지시문대로 입을 벌리고 있습니다.",
        "hard_violations": [],
        "physics": "찰리가 책상 앞에 자연스럽게 위치하여 모니터를 바라보는 자세가 안정적입니다."
       },
       {
        "label": "B",
        "direction": "찰리의 시선과 얼굴 방향이 오른쪽의 모니터를 향하고 있습니다.",
        "built_space": "배경에 도서관 서가와 창문이 흐릿하게 보이며, 우측에 모니터 측면이 자리 잡고 있습니다.",
        "entities": "찰리의 로봇 외형과 입을 벌린 표정은 나타나지만, 이전 샷에서 지정된 줄무늬 담요 위장이 없습니다.",
        "hard_violations": [],
        "physics": "모니터 앞에 위치한 상반신의 자세가 안정적으로 유지되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍, 모니터 불빛에 비친 얼굴, 크게 벌린 입, 그리고 이전 샷에서 고정된 담요 위장을 모두 정확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "표정과 조명, 배경 구성은 준수하나, 이전 샷 및 레퍼런스에서 필수적으로 유지해야 하는 담요 위장이 완전히 누락되어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선이 화면 오른쪽에 위치한 모니터를 향하고 있습니다.",
        "built_space": "흐릿하게 처리된 낡은 도서관 서가가 배경에 있고, 오른쪽 전면에 구형 모니터의 측면이 배치되어 있습니다.",
        "entities": "레퍼런스와 동일한 외형의 찰리가 스트라이프 패턴의 담요를 두르고 있으며, 지시문대로 입을 벌리고 있습니다.",
        "hard_violations": [],
        "physics": "찰리가 책상 앞에 자연스럽게 위치하여 모니터를 바라보는 자세가 안정적입니다."
       },
       {
        "label": "B",
        "direction": "찰리의 시선과 얼굴 방향이 오른쪽의 모니터를 향하고 있습니다.",
        "built_space": "배경에 도서관 서가와 창문이 흐릿하게 보이며, 우측에 모니터 측면이 자리 잡고 있습니다.",
        "entities": "찰리의 로봇 외형과 입을 벌린 표정은 나타나지만, 이전 샷에서 지정된 줄무늬 담요 위장이 없습니다.",
        "hard_violations": [],
        "physics": "모니터 앞에 위치한 상반신의 자세가 안정적으로 유지되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "모니터를 향해 입을 벌린 낡은 금속 얼굴은 구현했지만, 모니터가 좁은 가장자리보다 훨씬 크게 노출되고 기존 줄무늬 망토가 사라졌으며 입도 둥글기보다 길고 각지다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "모니터 노출 폭은 지시보다 크지만, 둥글게 벌린 입과 집중하는 얼굴, 기존 줄무늬 망토와 밧줄 위장을 유지해 A보다 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리와 보이는 눈은 화면 오른쪽의 모니터를 향한다. 모니터의 표시 면은 찰리 쪽으로 향하고 카메라에는 옆면과 후면이 보여 사용 방향은 맞다. 입은 크게 열렸지만 둥근 발음 모양보다는 세로로 긴 각진 개구부에 가깝다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대, 머리 뒤 왼쪽에 책과 선반 일부, 오른쪽 뒤에 어두운 창틀이 보인다. 책장은 흐리게 남아 도서관 맥락을 유지한다. 다만 모니터 몸통이 화면 너비의 약 3분의 1을 차지해 '좁은 측면 가장자리만' 보이라는 구도 지시를 어긴다. 얼굴뿐 아니라 양 어깨와 가슴도 상당히 포함된다. 중복 집기나 불가능한 반사는 없다.",
        "entities": "찰리 한 개체만 보이며 다른 사람과 읽을 수 있는 문자는 없다. 샌드 베이지 장갑판, 흰 각진 금속 얼굴, 원형 귀 장치와 안테나는 참조의 로봇 정체성과 대체로 맞는다. 그러나 참조에서 어깨와 몸통을 감싼 회색 줄무늬 망토와 밧줄이, 해당 부위가 보이는데도 없다. 금속의 마모와 실제 구형 모니터, 책은 확인된다. 하체와 침구는 구도 밖이므로 평가하지 않는다.",
        "hard_violations": [],
        "physics": "머리는 노출된 기계식 목과 몸통에 연결되어 있고 열린 아래턱도 얼굴 측면의 기구에 이어져 있다. 공중에 떠 있는 신체나 물체는 없다. 모니터 하단과 몸통의 좌석·바닥 접점은 잘려 있어 직접 확인할 수 없지만, 지지 없이 떠 있다고 판단할 근거는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 오른쪽 앞의 모니터 쪽으로 얼굴을 기울이고 두 눈을 그 방향에 집중한다. 모니터 표시 면은 찰리를 향하며 카메라에는 측후면만 보이므로 화면 내용이 노출되지 않는다. 열린 입은 아래턱과 내부 테두리가 둥근 형태를 만들어 A보다 크게 둥글게 발음하는 순간에 가깝다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대가 있고, 뒤에는 책이 꽂힌 선반 여러 구획과 낡은 수직 서가 측판이 흐리게 보인다. 어두운 도서관과 얼굴을 비추는 모니터 조명은 장소 조건에 부합한다. 그러나 모니터의 넓은 몸통이 화면 오른쪽 약 4분의 1을 차지해 좁은 가장자리만 보여야 한다는 지시는 충족하지 못한다. 머리와 어깨 외에 아래쪽 팔과 손도 일부 들어온다. 중복 집기나 불가능한 반사는 없다.",
        "entities": "참조와 같은 비인간 로봇 찰리만 등장한다. 낡은 베이지 장갑판, 흰 마스크형 얼굴, 원형 귀 장치와 안테나가 유지되며, 회색·흰색 줄무늬 망토와 어깨의 밧줄 매듭도 남아 있다. 입을 벌린 얼굴은 금속 부품으로 표현된다. 책과 구형 모니터가 보이고 읽을 수 있는 글이나 추가 인물은 없다. 하체와 구석 침구는 구도 밖이다.",
        "hard_violations": [],
        "physics": "머리는 기계식 목으로 몸통에 연결되고 아래턱도 측면 관절 구조에 붙어 있다. 망토는 어깨에 걸쳐 몸통을 따라 처지며 밧줄이 이를 묶는다. 아래쪽 손과 팔은 서로 연결된 채 몸 앞에 놓여 있어 떠 있는 절단 신체처럼 보이지 않는다. 모니터 받침과 좌석은 화면 밖이며, 보이는 범위에서 무지지 부유나 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "모니터를 향해 입을 벌린 낡은 금속 얼굴은 구현했지만, 모니터가 좁은 가장자리보다 훨씬 크게 노출되고 기존 줄무늬 망토가 사라졌으며 입도 둥글기보다 길고 각지다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "모니터 노출 폭은 지시보다 크지만, 둥글게 벌린 입과 집중하는 얼굴, 기존 줄무늬 망토와 밧줄 위장을 유지해 A보다 충실하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리와 보이는 눈은 화면 오른쪽의 모니터를 향한다. 모니터의 표시 면은 찰리 쪽으로 향하고 카메라에는 옆면과 후면이 보여 사용 방향은 맞다. 입은 크게 열렸지만 둥근 발음 모양보다는 세로로 긴 각진 개구부에 가깝다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대, 머리 뒤 왼쪽에 책과 선반 일부, 오른쪽 뒤에 어두운 창틀이 보인다. 책장은 흐리게 남아 도서관 맥락을 유지한다. 다만 모니터 몸통이 화면 너비의 약 3분의 1을 차지해 '좁은 측면 가장자리만' 보이라는 구도 지시를 어긴다. 얼굴뿐 아니라 양 어깨와 가슴도 상당히 포함된다. 중복 집기나 불가능한 반사는 없다.",
        "entities": "찰리 한 개체만 보이며 다른 사람과 읽을 수 있는 문자는 없다. 샌드 베이지 장갑판, 흰 각진 금속 얼굴, 원형 귀 장치와 안테나는 참조의 로봇 정체성과 대체로 맞는다. 그러나 참조에서 어깨와 몸통을 감싼 회색 줄무늬 망토와 밧줄이, 해당 부위가 보이는데도 없다. 금속의 마모와 실제 구형 모니터, 책은 확인된다. 하체와 침구는 구도 밖이므로 평가하지 않는다.",
        "hard_violations": [],
        "physics": "머리는 노출된 기계식 목과 몸통에 연결되어 있고 열린 아래턱도 얼굴 측면의 기구에 이어져 있다. 공중에 떠 있는 신체나 물체는 없다. 모니터 하단과 몸통의 좌석·바닥 접점은 잘려 있어 직접 확인할 수 없지만, 지지 없이 떠 있다고 판단할 근거는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 오른쪽 앞의 모니터 쪽으로 얼굴을 기울이고 두 눈을 그 방향에 집중한다. 모니터 표시 면은 찰리를 향하며 카메라에는 측후면만 보이므로 화면 내용이 노출되지 않는다. 열린 입은 아래턱과 내부 테두리가 둥근 형태를 만들어 A보다 크게 둥글게 발음하는 순간에 가깝다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대가 있고, 뒤에는 책이 꽂힌 선반 여러 구획과 낡은 수직 서가 측판이 흐리게 보인다. 어두운 도서관과 얼굴을 비추는 모니터 조명은 장소 조건에 부합한다. 그러나 모니터의 넓은 몸통이 화면 오른쪽 약 4분의 1을 차지해 좁은 가장자리만 보여야 한다는 지시는 충족하지 못한다. 머리와 어깨 외에 아래쪽 팔과 손도 일부 들어온다. 중복 집기나 불가능한 반사는 없다.",
        "entities": "참조와 같은 비인간 로봇 찰리만 등장한다. 낡은 베이지 장갑판, 흰 마스크형 얼굴, 원형 귀 장치와 안테나가 유지되며, 회색·흰색 줄무늬 망토와 어깨의 밧줄 매듭도 남아 있다. 입을 벌린 얼굴은 금속 부품으로 표현된다. 책과 구형 모니터가 보이고 읽을 수 있는 글이나 추가 인물은 없다. 하체와 구석 침구는 구도 밖이다.",
        "hard_violations": [],
        "physics": "머리는 기계식 목으로 몸통에 연결되고 아래턱도 측면 관절 구조에 붙어 있다. 망토는 어깨에 걸쳐 몸통을 따라 처지며 밧줄이 이를 묶는다. 아래쪽 손과 팔은 서로 연결된 채 몸 앞에 놓여 있어 떠 있는 절단 신체처럼 보이지 않는다. 모니터 받침과 좌석은 화면 밖이며, 보이는 범위에서 무지지 부유나 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.286
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.286
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1286
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 프레이밍, 모니터 불빛에 비친 얼굴, 크게 벌린 입, 그리고 이전 샷에서 고정된 담요 위장을 모두 정확하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1286,
    "verdict_ko": "표정과 조명, 배경 구성은 준수하나, 이전 샷 및 레퍼런스에서 필수적으로 유지해야 하는 담요 위장이 완전히 누락되어 감점되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17_sel.png",
    "asset_id": "6b3f3798-4303-498d-84fc-66efa084996f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-d2d4-7e6e-a1b6-2fb49a5126df",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S46sh17"
  }
 },
 "S46sh28::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:07:40.183430+00:00",
  "fingerprint": "19688091389cc588690c950caf715c480e855a89346795208b75a61dce308ee6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S46sh28_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S46sh28_sel.png",
  "source_sha256": "e8a3397b4b08d749f0f6a301a20c6429b70a9947b692dd9fcdeef063f9929b83",
  "file": "S46sh28_cine.png",
  "staged_sha256": "f3de97c970ea28a9898c68e9461f9735ac6c4e5d37466bad7766d32da11b7628",
  "latency_ms": 9322
 },
 "S47sh3::signage": {
  "fp": "0fb0867d39ad412d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S47sh3": {
  "input_fingerprint": "de52827095deea78",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 화려한 바다와 인물들이 그려진 거대한 벽화 앞에서 뭉툭한 크레파스를 쥔 손을 벽면에 댄 찰리의 낡은 뒷모습.\n\nLOCATION (lock): Along the mural-covered inner wall of the abandoned library's reading room, in early-morning ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Crayon mural (Extensive and still being drawn) — The illustrated wall face is seen obliquely, showing the blue sea, marine life, palms, and drawn figures of 찰리, 현우, 라울, 앰버, 페드로, their mother, and the priest around the imagined Haenam arrival; used as Colorful narrative backdrop separating the drawing hand from 찰리's silhouette; the depicted people remain drawings, not additional physical figures; Blunt crayon (Held against the wall in 찰리's hand) — Tip meets the mural beyond the outline of his torso; used as Small, clearly separated action detail establishing authorship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained early-morning ambient light, allowing the mural's blue sea and other explicitly rich colors to provide the scene's selective chromatic warmth and tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the early-morning library, a crayon mural already covers a wall: a blue sea leads to a colorful Haenam paradise with marine life, palms and a camper carrying Charlie, Hyunwoo, Raul and Amber, with Pedro, Miyeon and the priest also depicted. The written message about Haenam, Soyoung and Giant Charlie accompanies the pictures. 찰리: He stands at the mural holding a crayon and continuing the drawing. His worn metal body and retained disguise remain unchanged.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 화려한 바다와 인물들이 그려진 거대한 벽화 앞에서 뭉툭한 크레파스를 쥔 손을 벽면에 댄 찰리의 낡은 뒷모습.\n\nLOCATION (lock): Along the mural-covered inner wall of the abandoned library's reading room, in early-morning ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Crayon mural (Extensive and still being drawn) — The illustrated wall face is seen obliquely, showing the blue sea, marine life, palms, and drawn figures of 찰리, 현우, 라울, 앰버, 페드로, their mother, and the priest around the imagined Haenam arrival; used as Colorful narrative backdrop separating the drawing hand from 찰리's silhouette; the depicted people remain drawings, not additional physical figures; Blunt crayon (Held against the wall in 찰리's hand) — Tip meets the mural beyond the outline of his torso; used as Small, clearly separated action detail establishing authorship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained early-morning ambient light, allowing the mural's blue sea and other explicitly rich colors to provide the scene's selective chromatic warmth and tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the early-morning library, a crayon mural already covers a wall: a blue sea leads to a colorful Haenam paradise with marine life, palms and a camper carrying Charlie, Hyunwoo, Raul and Amber, with Pedro, Miyeon and the priest also depicted. The written message about Haenam, Soyoung and Giant Charlie accompanies the pictures. 찰리: He stands at the mural holding a crayon and continuing the drawing. His worn metal body and retained disguise remain unchanged.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 화려한 바다와 인물들이 그려진 거대한 벽화 앞에서 뭉툭한 크레파스를 쥔 손을 벽면에 댄 찰리의 낡은 뒷모습.\n\nLOCATION (lock): Along the mural-covered inner wall of the abandoned library's reading room, in early-morning ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Crayon mural (Extensive and still being drawn) — The illustrated wall face is seen obliquely, showing the blue sea, marine life, palms, and drawn figures of 찰리, 현우, 라울, 앰버, 페드로, their mother, and the priest around the imagined Haenam arrival; used as Colorful narrative backdrop separating the drawing hand from 찰리's silhouette; the depicted people remain drawings, not additional physical figures; Blunt crayon (Held against the wall in 찰리's hand) — Tip meets the mural beyond the outline of his torso; used as Small, clearly separated action detail establishing authorship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained early-morning ambient light, allowing the mural's blue sea and other explicitly rich colors to provide the scene's selective chromatic warmth and tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the early-morning library, a crayon mural already covers a wall: a blue sea leads to a colorful Haenam paradise with marine life, palms and a camper carrying Charlie, Hyunwoo, Raul and Amber, with Pedro, Miyeon and the priest also depicted. The written message about Haenam, Soyoung and Giant Charlie accompanies the pictures. 찰리: He stands at the mural holding a crayon and continuing the drawing. His worn metal body and retained disguise remain unchanged.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 오른손에 뭉툭한 크레파스를 쥐고 벽화 속 인물의 머리 부분에 대고 있다.",
    "built_space": "창문이 있는 텅 빈 타일 바닥의 방으로, 이전 샷에서 지정된 도서관의 나무 책장들이 전혀 존재하지 않는다.",
    "entities": "찰리의 외형과 장비는 레퍼런스와 일치한다. 벽화에는 바다, 야자수, 캠핑카, 여러 인물과 신부가 그려져 있으나 지시된 찰리의 그림이 빠져 있다. 좌측 상단 깃발 그림에 읽을 수 있는 문자가 포함되어 있다.",
    "hard_violations": [
     "[gemini-pro] 이전 샷의 장소 고정 위반 (도서관 책장 누락 및 전혀 다른 공간 묘사)",
     "[gemini-pro] 텍스트 포함 금지 위반 (벽화 내 깃발에 문자 표기)"
    ],
    "physics": "찰리는 두 발로 바닥을 딛고 서 있으며, 손으로 크레파스를 쥐고 벽에 대는 자세가 물리적으로 자연스럽게 지탱되고 있다."
   },
   {
    "label": "B",
    "direction": "찰리가 오른손에 크레파스를 쥐고 벽화의 하늘과 바다 경계선 부분에 대고 선을 긋고 있다.",
    "built_space": "이전 샷의 배경인 낡은 도서관 내부로, 화면 좌측에 나무 책장들이 늘어서 있어 장소의 일관성이 정확히 유지되고 있다.",
    "entities": "찰리의 낡은 뒷모습, 마스크, 담요가 레퍼런스와 정확히 일치한다. 벽화에는 바다, 해양 생물, 야자수, 아이들, 신부뿐만 아니라 찰리 본인의 모습도 지시대로 모두 그려져 있다.",
    "hard_violations": [],
    "physics": "찰리는 바닥에 안정적으로 서서 오른팔을 뻗어 크레파스를 쥐고 있으며, 벽면에 크레파스가 맞닿아 있는 물리적 접촉이 사실적으로 표현되었다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 외형은 잘 구현되었으나, 이전 샷에서 고정된 장소(도서관 책장)를 무시하고 텍스트 제약을 위반하여 실격입니다."
       },
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "찰리의 뒷모습, 이전 샷의 도서관 배경, 벽화의 디테일(찰리 본인의 그림 포함)을 모두 지시사항에 맞게 완벽하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 오른손에 뭉툭한 크레파스를 쥐고 벽화 속 인물의 머리 부분에 대고 있다.",
        "built_space": "창문이 있는 텅 빈 타일 바닥의 방으로, 이전 샷에서 지정된 도서관의 나무 책장들이 전혀 존재하지 않는다.",
        "entities": "찰리의 외형과 장비는 레퍼런스와 일치한다. 벽화에는 바다, 야자수, 캠핑카, 여러 인물과 신부가 그려져 있으나 지시된 찰리의 그림이 빠져 있다. 좌측 상단 깃발 그림에 읽을 수 있는 문자가 포함되어 있다.",
        "hard_violations": [
         "이전 샷의 장소 고정 위반 (도서관 책장 누락 및 전혀 다른 공간 묘사)",
         "텍스트 포함 금지 위반 (벽화 내 깃발에 문자 표기)"
        ],
        "physics": "찰리는 두 발로 바닥을 딛고 서 있으며, 손으로 크레파스를 쥐고 벽에 대는 자세가 물리적으로 자연스럽게 지탱되고 있다."
       },
       {
        "label": "B",
        "direction": "찰리가 오른손에 크레파스를 쥐고 벽화의 하늘과 바다 경계선 부분에 대고 선을 긋고 있다.",
        "built_space": "이전 샷의 배경인 낡은 도서관 내부로, 화면 좌측에 나무 책장들이 늘어서 있어 장소의 일관성이 정확히 유지되고 있다.",
        "entities": "찰리의 낡은 뒷모습, 마스크, 담요가 레퍼런스와 정확히 일치한다. 벽화에는 바다, 해양 생물, 야자수, 아이들, 신부뿐만 아니라 찰리 본인의 모습도 지시대로 모두 그려져 있다.",
        "hard_violations": [],
        "physics": "찰리는 바닥에 안정적으로 서서 오른팔을 뻗어 크레파스를 쥐고 있으며, 벽면에 크레파스가 맞닿아 있는 물리적 접촉이 사실적으로 표현되었다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 외형은 잘 구현되었으나, 이전 샷에서 고정된 장소(도서관 책장)를 무시하고 텍스트 제약을 위반하여 실격입니다."
       },
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "찰리의 뒷모습, 이전 샷의 도서관 배경, 벽화의 디테일(찰리 본인의 그림 포함)을 모두 지시사항에 맞게 완벽하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 오른손에 뭉툭한 크레파스를 쥐고 벽화 속 인물의 머리 부분에 대고 있다.",
        "built_space": "창문이 있는 텅 빈 타일 바닥의 방으로, 이전 샷에서 지정된 도서관의 나무 책장들이 전혀 존재하지 않는다.",
        "entities": "찰리의 외형과 장비는 레퍼런스와 일치한다. 벽화에는 바다, 야자수, 캠핑카, 여러 인물과 신부가 그려져 있으나 지시된 찰리의 그림이 빠져 있다. 좌측 상단 깃발 그림에 읽을 수 있는 문자가 포함되어 있다.",
        "hard_violations": [
         "이전 샷의 장소 고정 위반 (도서관 책장 누락 및 전혀 다른 공간 묘사)",
         "텍스트 포함 금지 위반 (벽화 내 깃발에 문자 표기)"
        ],
        "physics": "찰리는 두 발로 바닥을 딛고 서 있으며, 손으로 크레파스를 쥐고 벽에 대는 자세가 물리적으로 자연스럽게 지탱되고 있다."
       },
       {
        "label": "B",
        "direction": "찰리가 오른손에 크레파스를 쥐고 벽화의 하늘과 바다 경계선 부분에 대고 선을 긋고 있다.",
        "built_space": "이전 샷의 배경인 낡은 도서관 내부로, 화면 좌측에 나무 책장들이 늘어서 있어 장소의 일관성이 정확히 유지되고 있다.",
        "entities": "찰리의 낡은 뒷모습, 마스크, 담요가 레퍼런스와 정확히 일치한다. 벽화에는 바다, 해양 생물, 야자수, 아이들, 신부뿐만 아니라 찰리 본인의 모습도 지시대로 모두 그려져 있다.",
        "hard_violations": [],
        "physics": "찰리는 바닥에 안정적으로 서서 오른팔을 뻗어 크레파스를 쥐고 있으며, 벽면에 크레파스가 맞닿아 있는 물리적 접촉이 사실적으로 표현되었다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "낡은 찰리의 뒷모습과 벽에 닿는 손은 구현했지만, 상체 중심으로 잘린 구도가 지정된 와이드 숏보다 좁고 크레파스도 지나치게 크다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "와이드 숏에서 찰리의 뒷모습과 몸통 밖으로 분리된 크레파스 접점을 명확히 보여주며, 비스듬한 대형 벽화와 아침빛도 지시에 더 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라에 등을 보이고 머리를 오른쪽 벽화의 작업 지점으로 돌린다. 오른손에 잡힌 갈색 크레파스의 오른쪽 끝이 인물 그림 위쪽의 파란 선 부근을 향해 벽에 닿아 있다. 손과 접점은 몸통 윤곽 밖에 분리되어 보인다.",
        "built_space": "오른쪽 벽화 벽이 왼쪽 뒤로 이어지고, 왼쪽에는 뒤쪽 책장 한 구획과 옆벽을 따라 이어지는 격자형 책장, 전경의 낮은 목재 가구 일부가 보인다. 천장에는 길쭉한 조명 기구들이 있다. 낡은 벽과 책장은 이전 사진의 도서관 재질과 대체로 연결되지만, 이전 사진만으로 정확한 배치까지 검증할 수는 없다. 찰리의 하체가 잘리고 상체가 화면을 크게 차지해 요청한 와이드 숏보다 가깝다.",
        "entities": "실제 등장 개체는 찰리 하나이고, 벽의 다른 인물들은 그림이다. 찰리의 샌드 베이지 장갑, 긴 기계 팔, 안테나, 회색 줄무늬 천과 밧줄은 참조와 부합한다. 얼굴은 대부분 가려져 정면 마스크의 세부는 검증할 수 없다. 벽화에는 푸른 바다, 야자수, 여러 해양 생물, 찰리 그림과 성직자를 포함한 사람 그림들이 있으나 각 인물의 신원을 모두 확정하기 어렵고 캠핑카는 보이지 않는다. 크레파스는 손에 비해 상당히 굵고 길다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "오른팔은 어깨와 팔꿈치 관절을 통해 들어 올려져 있고 기계 손가락이 크레파스를 감싸 잡는다. 크레파스는 손으로 지지되며 벽 접점도 보인다. 천은 어깨에 걸쳐 아래로 처진다. 발은 화면 밖이므로 바닥 접촉은 확인할 수 없지만, 몸이 공중에 떠 있다는 증거는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 등을 보인 채 오른쪽 벽화의 작업 지점을 바라보는 방향으로 머리를 돌리고 있다. 뻗은 오른손의 짧은 크레파스 끝은 옅게 그려진 인물의 머리 옆 벽면에 닿는다. 그 접점이 몸통 바깥으로 명확히 분리되어 그림을 이어 그리는 행동이 읽힌다.",
        "built_space": "대형 벽화가 오른쪽 벽을 덮고 왼쪽 뒤로 비스듬히 이어진다. 왼쪽에는 창 한 조, 천장에는 일부 잘린 조명 한 개와 어두운 개구부 한 개, 벽 아래에는 걸레받이가 보인다. 찰리는 벽 바로 앞 바닥에 서 있으며 전신과 벽화의 넓은 면적을 함께 담은 와이드 숏이다. 이전 사진의 책장은 이 방향에서 보이지 않아 같은 방의 배치를 확증하기 어렵지만, 드러난 구조적 모순은 없다.",
        "entities": "실제 등장 개체는 찰리 하나이며 벽화의 인물들은 모두 표면에 그린 그림이다. 찰리의 육중한 몸통, 긴 팔, 짧은 다리, 마모된 베이지 장갑, 안테나와 줄무늬 천은 참조에 대체로 맞는다. 밧줄의 등쪽 묶음은 이전 사진에서 확인되지 않는 세부다. 벽화에는 푸른 바다, 해양 생물, 야자수, 캠핑카와 성직자를 포함한 여섯 사람 그림이 보인다. 별도의 찰리 그림과 지명된 일곱 인물 전체의 대응은 명확하지 않으며, 캠핑카에 사람들이 탑승한 모습도 아니다. 왼쪽 리본에는 문자 같은 흔적이 있으나 확실하게 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 두 기계 발은 바닥에 닿아 몸을 지지하고, 다리를 벌린 자세에서 오른팔을 벽으로 내민다. 크레파스는 오른손 손가락 사이에 잡혀 있고 끝이 벽면에 닿아 있어 지지와 사용 방향이 자연스럽다. 천은 어깨와 몸통에 걸리고 밧줄에 고정되어 아래로 늘어진다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "낡은 찰리의 뒷모습과 벽에 닿는 손은 구현했지만, 상체 중심으로 잘린 구도가 지정된 와이드 숏보다 좁고 크레파스도 지나치게 크다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "와이드 숏에서 찰리의 뒷모습과 몸통 밖으로 분리된 크레파스 접점을 명확히 보여주며, 비스듬한 대형 벽화와 아침빛도 지시에 더 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 카메라에 등을 보이고 머리를 오른쪽 벽화의 작업 지점으로 돌린다. 오른손에 잡힌 갈색 크레파스의 오른쪽 끝이 인물 그림 위쪽의 파란 선 부근을 향해 벽에 닿아 있다. 손과 접점은 몸통 윤곽 밖에 분리되어 보인다.",
        "built_space": "오른쪽 벽화 벽이 왼쪽 뒤로 이어지고, 왼쪽에는 뒤쪽 책장 한 구획과 옆벽을 따라 이어지는 격자형 책장, 전경의 낮은 목재 가구 일부가 보인다. 천장에는 길쭉한 조명 기구들이 있다. 낡은 벽과 책장은 이전 사진의 도서관 재질과 대체로 연결되지만, 이전 사진만으로 정확한 배치까지 검증할 수는 없다. 찰리의 하체가 잘리고 상체가 화면을 크게 차지해 요청한 와이드 숏보다 가깝다.",
        "entities": "실제 등장 개체는 찰리 하나이고, 벽의 다른 인물들은 그림이다. 찰리의 샌드 베이지 장갑, 긴 기계 팔, 안테나, 회색 줄무늬 천과 밧줄은 참조와 부합한다. 얼굴은 대부분 가려져 정면 마스크의 세부는 검증할 수 없다. 벽화에는 푸른 바다, 야자수, 여러 해양 생물, 찰리 그림과 성직자를 포함한 사람 그림들이 있으나 각 인물의 신원을 모두 확정하기 어렵고 캠핑카는 보이지 않는다. 크레파스는 손에 비해 상당히 굵고 길다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "오른팔은 어깨와 팔꿈치 관절을 통해 들어 올려져 있고 기계 손가락이 크레파스를 감싸 잡는다. 크레파스는 손으로 지지되며 벽 접점도 보인다. 천은 어깨에 걸쳐 아래로 처진다. 발은 화면 밖이므로 바닥 접촉은 확인할 수 없지만, 몸이 공중에 떠 있다는 증거는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 등을 보인 채 오른쪽 벽화의 작업 지점을 바라보는 방향으로 머리를 돌리고 있다. 뻗은 오른손의 짧은 크레파스 끝은 옅게 그려진 인물의 머리 옆 벽면에 닿는다. 그 접점이 몸통 바깥으로 명확히 분리되어 그림을 이어 그리는 행동이 읽힌다.",
        "built_space": "대형 벽화가 오른쪽 벽을 덮고 왼쪽 뒤로 비스듬히 이어진다. 왼쪽에는 창 한 조, 천장에는 일부 잘린 조명 한 개와 어두운 개구부 한 개, 벽 아래에는 걸레받이가 보인다. 찰리는 벽 바로 앞 바닥에 서 있으며 전신과 벽화의 넓은 면적을 함께 담은 와이드 숏이다. 이전 사진의 책장은 이 방향에서 보이지 않아 같은 방의 배치를 확증하기 어렵지만, 드러난 구조적 모순은 없다.",
        "entities": "실제 등장 개체는 찰리 하나이며 벽화의 인물들은 모두 표면에 그린 그림이다. 찰리의 육중한 몸통, 긴 팔, 짧은 다리, 마모된 베이지 장갑, 안테나와 줄무늬 천은 참조에 대체로 맞는다. 밧줄의 등쪽 묶음은 이전 사진에서 확인되지 않는 세부다. 벽화에는 푸른 바다, 해양 생물, 야자수, 캠핑카와 성직자를 포함한 여섯 사람 그림이 보인다. 별도의 찰리 그림과 지명된 일곱 인물 전체의 대응은 명확하지 않으며, 캠핑카에 사람들이 탑승한 모습도 아니다. 왼쪽 리본에는 문자 같은 흔적이 있으나 확실하게 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 두 기계 발은 바닥에 닿아 몸을 지지하고, 다리를 벌린 자세에서 오른팔을 벽으로 내민다. 크레파스는 오른손 손가락 사이에 잡혀 있고 끝이 벽면에 닿아 있어 지지와 사용 방향이 자연스럽다. 천은 어깨와 몸통에 걸리고 밧줄에 고정되어 아래로 늘어진다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.4,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.15,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷의 장소 고정 위반 (도서관 책장 누락 및 전혀 다른 공간 묘사)",
     "[gemini-pro] 텍스트 포함 금지 위반 (벽화 내 깃발에 문자 표기)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1150,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1150,
    "verdict_ko": "찰리의 외형은 잘 구현되었으나, 이전 샷에서 고정된 장소(도서관 책장)를 무시하고 텍스트 제약을 위반하여 실격입니다.  ★위반: [gemini-pro] 이전 샷의 장소 고정 위반 (도서관 책장 누락 및 전혀 다른 공간 묘사) / [gemini-pro] 텍스트 포함 금지 위반 (벽화 내 깃발에 문자 표기)"
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "찰리의 뒷모습, 이전 샷의 도서관 배경, 벽화의 디테일(찰리 본인의 그림 포함)을 모두 지시사항에 맞게 완벽하게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh28_sel.png",
    "asset_id": "4493e6df-e443-4f98-b125-7120a0207757",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-d48f-791a-957e-c650c860dcf5",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S46sh28"
  }
 },
 "S47sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:08:36.714246+00:00",
  "fingerprint": "84f83be8b80cb67e80adf23416f037cc2ddd709c1783c85fcf8448e591fcae03",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S47sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S47sh3_sel.png",
  "source_sha256": "c1b45fa610be36bf408699b868feb3a2099c85f62c320575392821752258bb0c",
  "file": "S47sh3_cine.png",
  "staged_sha256": "a92bc96db6f778be709f0dbb795b845a6191eb3c1b38477f159ce2ac354665ad",
  "latency_ms": 12930
 },
 "S47sh9::signage": {
  "fp": "5c3081ff9a2933bd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S47sh9::bgfirst_bg": {
  "input_fingerprint": "fa83a79386c6066b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning.\n\nTIME OF DAY (lock): early morning.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning.\n\nTIME OF DAY (lock): early morning.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9__bgfirst_bg.png",
  "asset_id": "85df77f5-d63d-4165-9bbc-cafbc3fedb4e",
  "input_asset_ids": [
   "198f725a-7e21-4a39-9d9c-8249fc9fea2c",
   "f4370321-fb89-4785-899e-192be57a9cc4"
  ]
 },
 "S47sh9": {
  "input_fingerprint": "34d324bd0c1ea9f5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The completed portions of the colorful Haenam mural and its written message remain on the library wall. The camper remains outside the abandoned library. 현우: He has removed one shoe and is extracting the intact contact card concealed inside it, with the disposable phone ready. His facial bruises and untreated leg wound remain.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The completed portions of the colorful Haenam mural and its written message remain on the library wall. The camper remains outside the abandoned library. 현우: He has removed one shoe and is extracting the intact contact card concealed inside it, with the disposable phone ready. His facial bruises and untreated leg wound remain.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The completed portions of the colorful Haenam mural and its written message remain on the library wall. The camper remains outside the abandoned library. 현우: He has removed one shoe and is extracting the intact contact card concealed inside it, with the disposable phone ready. His facial bruises and untreated leg wound remain.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9__bgfirst_bg.png",
     "asset_id": "85df77f5-d63d-4165-9bbc-cafbc3fedb4e",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S47sh9.png",
     "asset_id": "198f725a-7e21-4a39-9d9c-8249fc9fea2c",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B02.png",
     "asset_id": "f4370321-fb89-4785-899e-192be57a9cc4",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "오른손이 공중에 들린 신발에서 흰색 카드를 뽑고 있으며, 특정 사물을 향한 시선이나 방향성은 나타나지 않음.",
    "built_space": "레퍼런스와 동일한 야외 구석으로, 배관, 환기구, 콘크리트 바닥 및 벽면이 올바른 위치에 묘사됨.",
    "entities": "회색 셔츠 소매와 카고 바지를 입은 인물의 손, 더러운 운동화, 흰색 카드가 식별되나 지시된 대포폰은 없음.",
    "hard_violations": [
     "[gemini-pro] 지시된 필수 소품(대포폰) 누락"
    ],
    "physics": "왼손이 신발을 허공에 들고 있고 오른손이 카드를 뽑는 자세로, 물리적으로 지탱되고 있으나 텍스트의 바닥 배치 설정과 불일치함."
   },
   {
    "label": "B",
    "direction": "오른손이 바닥에 놓인 신발에서 어두운 색상의 카드를 뽑아 올리고 있으며, 왼손은 휴대전화를 쥐고 대기 중임.",
    "built_space": "레퍼런스의 도서관 야외 구석 배경(배관, 벽면, 바닥)이 인물의 뒤와 아래에 정확하게 배치됨.",
    "entities": "상처 난 거친 손, 회색 셔츠 소매, 올리브색 바지, 바닥에 놓인 운동화, 카드, 대포폰이 모두 명확히 식별됨.",
    "hard_violations": [],
    "physics": "신발은 바닥에 안정적으로 놓여 있고, 무릎에 기댄 왼손과 카드를 잡은 오른손 모두 자연스럽게 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "대포폰을 쥔 왼손과 바닥에 놓인 신발 등 프롬프트의 구도와 소품 요구사항을 정확히 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "준비된 대포폰이 누락되었고 신발을 공중에 들고 있어 바닥을 배경으로 앵커링하라는 지시를 어겼습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른손이 공중에 들린 신발에서 흰색 카드를 뽑고 있으며, 특정 사물을 향한 시선이나 방향성은 나타나지 않음.",
        "built_space": "레퍼런스와 동일한 야외 구석으로, 배관, 환기구, 콘크리트 바닥 및 벽면이 올바른 위치에 묘사됨.",
        "entities": "회색 셔츠 소매와 카고 바지를 입은 인물의 손, 더러운 운동화, 흰색 카드가 식별되나 지시된 대포폰은 없음.",
        "hard_violations": [
         "지시된 필수 소품(대포폰) 누락"
        ],
        "physics": "왼손이 신발을 허공에 들고 있고 오른손이 카드를 뽑는 자세로, 물리적으로 지탱되고 있으나 텍스트의 바닥 배치 설정과 불일치함."
       },
       {
        "label": "B",
        "direction": "오른손이 바닥에 놓인 신발에서 어두운 색상의 카드를 뽑아 올리고 있으며, 왼손은 휴대전화를 쥐고 대기 중임.",
        "built_space": "레퍼런스의 도서관 야외 구석 배경(배관, 벽면, 바닥)이 인물의 뒤와 아래에 정확하게 배치됨.",
        "entities": "상처 난 거친 손, 회색 셔츠 소매, 올리브색 바지, 바닥에 놓인 운동화, 카드, 대포폰이 모두 명확히 식별됨.",
        "hard_violations": [],
        "physics": "신발은 바닥에 안정적으로 놓여 있고, 무릎에 기댄 왼손과 카드를 잡은 오른손 모두 자연스럽게 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "대포폰을 쥔 왼손과 바닥에 놓인 신발 등 프롬프트의 구도와 소품 요구사항을 정확히 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "준비된 대포폰이 누락되었고 신발을 공중에 들고 있어 바닥을 배경으로 앵커링하라는 지시를 어겼습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "오른손이 공중에 들린 신발에서 흰색 카드를 뽑고 있으며, 특정 사물을 향한 시선이나 방향성은 나타나지 않음.",
        "built_space": "레퍼런스와 동일한 야외 구석으로, 배관, 환기구, 콘크리트 바닥 및 벽면이 올바른 위치에 묘사됨.",
        "entities": "회색 셔츠 소매와 카고 바지를 입은 인물의 손, 더러운 운동화, 흰색 카드가 식별되나 지시된 대포폰은 없음.",
        "hard_violations": [
         "지시된 필수 소품(대포폰) 누락"
        ],
        "physics": "왼손이 신발을 허공에 들고 있고 오른손이 카드를 뽑는 자세로, 물리적으로 지탱되고 있으나 텍스트의 바닥 배치 설정과 불일치함."
       },
       {
        "label": "B",
        "direction": "오른손이 바닥에 놓인 신발에서 어두운 색상의 카드를 뽑아 올리고 있으며, 왼손은 휴대전화를 쥐고 대기 중임.",
        "built_space": "레퍼런스의 도서관 야외 구석 배경(배관, 벽면, 바닥)이 인물의 뒤와 아래에 정확하게 배치됨.",
        "entities": "상처 난 거친 손, 회색 셔츠 소매, 올리브색 바지, 바닥에 놓인 운동화, 카드, 대포폰이 모두 명확히 식별됨.",
        "hard_violations": [],
        "physics": "신발은 바닥에 안정적으로 놓여 있고, 무릎에 기댄 왼손과 카드를 잡은 오른손 모두 자연스럽게 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "거친 손과 준비된 휴대전화는 잘 구현했지만, 신발이 전경을 크게 차지하고 카드의 넓은 면이 대부분 드러나 손·무릎 중심의 클로즈업과 반쯤 인출된 순간에서 멀어진다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "거친 양손과 굽힌 무릎을 중심으로 신발 입구에서 카드를 꺼내는 관계를 더 충실히 보여주며, 비스듬한 카드 면과 의상·장소도 잘 맞지만 카드가 정확히 절반만 나온 상태는 불명확하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른손이 카드 윗부분을 잡고 신발 입구 위로 들어 올리는 방향이다. 신발 앞코는 화면 오른쪽, 입구는 카메라와 손 쪽으로 비스듬히 위를 향한다. 카드의 넓은 무지 면이 카메라에 노출되어, 연락처 면이 좁게 보인다는 지시와는 차이가 있다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "얼룩진 콘크리트 벽 두 면, 모서리 배수관 한 개, 뒤쪽 상단 창 구획, 오른쪽 루버 설비 한 곳과 젖은 바닥이 보인다. 위치 관계와 재질은 장소 참조에 부합한다. 현우는 왼쪽에서 웅크리고 신발은 오른쪽 전경 바닥에 놓여 있다. 낮은 촬영 위치는 허용되지만, 신발이 크게 부각되고 무릎 맥락은 상대적으로 약하다.",
        "entities": "한 사람의 손 두 개와 회색 셔츠 소매, 짙은 올리브색 바지, 착용 중인 신발 일부가 보인다. 얼굴이 없어 정확한 나이와 한국계 미국인 정체성은 확인할 수 없지만, 손의 체격과 피부에 뚜렷한 충돌은 없다. 손에는 때와 작은 상처가 있다. 벗은 신발 한 짝, 빳빳하고 찢어지지 않은 카드 한 장, 다른 손에 쥔 소형 휴대전화 한 대가 보인다. 벗은 신발은 참조의 갈색 계열 신발보다 검은 운동화에 가깝다. 카드에는 읽을 수 있는 연락처가 없고 노출 면적이 크다. 얼굴 멍과 다리 상처, 벽화와 캠핑카는 이 크롭에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "웅크린 몸 아래 착용한 신발이 바닥에 닿아 지지한다. 벗은 신발도 옆으로 기울어진 채 바닥에 접촉한다. 카드는 손가락으로 잡혀 있고 휴대전화는 다른 손에 받쳐져 있어 무지지 물체는 없다. 인출 동작은 가능하지만 카드 아래쪽까지 넓게 보여 절반이 신발 안에 남아 있다는 상태는 명확하지 않다."
       },
       {
        "label": "B",
        "direction": "위쪽 손이 카드 끝을 집어 신발 입구에서 위로 꺼내고, 아래쪽 손은 신발 뒤축을 잡는다. 신발 앞코는 오른쪽 위로 향하고 입구는 손과 카메라 쪽으로 열린다. 카드 면은 비스듬히 좁게 보이며 작은 인쇄 흔적은 읽히지 않는다. 얼굴이 잘려 시선 방향은 확인할 수 없다.",
        "built_space": "참조와 같은 거친 콘크리트 벽 두 면, 모서리 배수관 한 개, 뒤쪽 상단의 분할 창, 오른쪽 루버 설비와 그 아래 턱, 바닥 배수구 한 개가 보인다. 고정 구조의 중복은 없다. 왼쪽의 굽힌 무릎과 전경 양손이 웅크린 자세의 맥락을 만들고, 신발은 그 사이에서 들려 있다. 바닥과 벽 하단의 습기·오염도 참조 장소와 잘 맞는다.",
        "entities": "한 사람의 거칠고 더러워진 양손, 팔뚝, 회색 셔츠와 올리브색 바지 무릎이 보인다. 노출된 체격과 복장은 현우 참조와 대체로 맞으며, 얼굴이 없어 정확한 나이·민족적 정체성은 확인할 수 없다. 벗은 신발 한 짝은 진흙 묻은 갈색 계열로 참조 신발의 인상에 가깝다. 카드 한 장은 빳빳하고 온전하며, 연락처 인쇄로 볼 수 있는 작은 흔적은 판독 불가능하다. 휴대전화는 보이지 않으므로 준비 상태를 확인할 수 없다. 얼굴 멍과 다리 상처, 벽화와 캠핑카도 프레임 밖이다.",
        "hard_violations": [],
        "physics": "들린 신발은 아래쪽 손이 뒤축과 옆면을 확실히 잡아 지지하고, 카드는 위쪽 손의 엄지와 손가락 사이에 고정되어 있다. 굽힌 무릎과 팔의 배치는 웅크려 물건을 꺼내는 동작으로 가능하다. 발과 엉덩이의 지지점은 크롭 밖이며, 몸이 떠 있다는 징후는 없다. 카드가 입구와 겹치지만 정확히 절반이 내부에 남아 있는지는 분명하지 않다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "거친 손과 준비된 휴대전화는 잘 구현했지만, 신발이 전경을 크게 차지하고 카드의 넓은 면이 대부분 드러나 손·무릎 중심의 클로즈업과 반쯤 인출된 순간에서 멀어진다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "거친 양손과 굽힌 무릎을 중심으로 신발 입구에서 카드를 꺼내는 관계를 더 충실히 보여주며, 비스듬한 카드 면과 의상·장소도 잘 맞지만 카드가 정확히 절반만 나온 상태는 불명확하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른손이 카드 윗부분을 잡고 신발 입구 위로 들어 올리는 방향이다. 신발 앞코는 화면 오른쪽, 입구는 카메라와 손 쪽으로 비스듬히 위를 향한다. 카드의 넓은 무지 면이 카메라에 노출되어, 연락처 면이 좁게 보인다는 지시와는 차이가 있다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "얼룩진 콘크리트 벽 두 면, 모서리 배수관 한 개, 뒤쪽 상단 창 구획, 오른쪽 루버 설비 한 곳과 젖은 바닥이 보인다. 위치 관계와 재질은 장소 참조에 부합한다. 현우는 왼쪽에서 웅크리고 신발은 오른쪽 전경 바닥에 놓여 있다. 낮은 촬영 위치는 허용되지만, 신발이 크게 부각되고 무릎 맥락은 상대적으로 약하다.",
        "entities": "한 사람의 손 두 개와 회색 셔츠 소매, 짙은 올리브색 바지, 착용 중인 신발 일부가 보인다. 얼굴이 없어 정확한 나이와 한국계 미국인 정체성은 확인할 수 없지만, 손의 체격과 피부에 뚜렷한 충돌은 없다. 손에는 때와 작은 상처가 있다. 벗은 신발 한 짝, 빳빳하고 찢어지지 않은 카드 한 장, 다른 손에 쥔 소형 휴대전화 한 대가 보인다. 벗은 신발은 참조의 갈색 계열 신발보다 검은 운동화에 가깝다. 카드에는 읽을 수 있는 연락처가 없고 노출 면적이 크다. 얼굴 멍과 다리 상처, 벽화와 캠핑카는 이 크롭에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "웅크린 몸 아래 착용한 신발이 바닥에 닿아 지지한다. 벗은 신발도 옆으로 기울어진 채 바닥에 접촉한다. 카드는 손가락으로 잡혀 있고 휴대전화는 다른 손에 받쳐져 있어 무지지 물체는 없다. 인출 동작은 가능하지만 카드 아래쪽까지 넓게 보여 절반이 신발 안에 남아 있다는 상태는 명확하지 않다."
       },
       {
        "label": "A",
        "direction": "위쪽 손이 카드 끝을 집어 신발 입구에서 위로 꺼내고, 아래쪽 손은 신발 뒤축을 잡는다. 신발 앞코는 오른쪽 위로 향하고 입구는 손과 카메라 쪽으로 열린다. 카드 면은 비스듬히 좁게 보이며 작은 인쇄 흔적은 읽히지 않는다. 얼굴이 잘려 시선 방향은 확인할 수 없다.",
        "built_space": "참조와 같은 거친 콘크리트 벽 두 면, 모서리 배수관 한 개, 뒤쪽 상단의 분할 창, 오른쪽 루버 설비와 그 아래 턱, 바닥 배수구 한 개가 보인다. 고정 구조의 중복은 없다. 왼쪽의 굽힌 무릎과 전경 양손이 웅크린 자세의 맥락을 만들고, 신발은 그 사이에서 들려 있다. 바닥과 벽 하단의 습기·오염도 참조 장소와 잘 맞는다.",
        "entities": "한 사람의 거칠고 더러워진 양손, 팔뚝, 회색 셔츠와 올리브색 바지 무릎이 보인다. 노출된 체격과 복장은 현우 참조와 대체로 맞으며, 얼굴이 없어 정확한 나이·민족적 정체성은 확인할 수 없다. 벗은 신발 한 짝은 진흙 묻은 갈색 계열로 참조 신발의 인상에 가깝다. 카드 한 장은 빳빳하고 온전하며, 연락처 인쇄로 볼 수 있는 작은 흔적은 판독 불가능하다. 휴대전화는 보이지 않으므로 준비 상태를 확인할 수 없다. 얼굴 멍과 다리 상처, 벽화와 캠핑카도 프레임 밖이다.",
        "hard_violations": [],
        "physics": "들린 신발은 아래쪽 손이 뒤축과 옆면을 확실히 잡아 지지하고, 카드는 위쪽 손의 엄지와 손가락 사이에 고정되어 있다. 굽힌 무릎과 팔의 배치는 웅크려 물건을 꺼내는 동작으로 가능하다. 발과 엉덩이의 지지점은 크롭 밖이며, 몸이 떠 있다는 징후는 없다. 카드가 입구와 겹치지만 정확히 절반이 내부에 남아 있는지는 분명하지 않다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.875
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시된 필수 소품(대포폰) 누락"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "대포폰을 쥔 왼손과 바닥에 놓인 신발 등 프롬프트의 구도와 소품 요구사항을 정확히 구현했습니다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "준비된 대포폰이 누락되었고 신발을 공중에 들고 있어 바닥을 배경으로 앵커링하라는 지시를 어겼습니다.  ★위반: [gemini-pro] 지시된 필수 소품(대포폰) 누락"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B02.png",
    "asset_id": "f4370321-fb89-4785-899e-192be57a9cc4",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-d645-778e-8384-6d2f63108c7d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9__bgfirst_bg.png",
   "bg_asset_id": "85df77f5-d63d-4165-9bbc-cafbc3fedb4e",
   "bg_record_key": "S47sh9::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S47sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:09:48.222306+00:00",
  "fingerprint": "4f7f6093c4e48bf359ab403bdd1aeadcde3a1377895fbb22b5faa099227ba96d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S47sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S47sh9_sel.png",
  "source_sha256": "1ae8b474ddba68ba4a5f9dd1b9c62436ecad10fcfec1d292b5f9ca4197318bb0",
  "file": "S47sh9_cine.png",
  "staged_sha256": "4c851fcb9520af2bbc1098cae4a43f34fe9535848c98810ecb640195e87f9da1",
  "latency_ms": 10310
 },
 "S47sh12::signage": {
  "fp": "bc58d9dd83132e37",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S47sh12": {
  "input_fingerprint": "68609b4720b54fde",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 뒤돌아선 현우의 등 바로 뒤에 소리 없이 묵묵히 서 있는 찰리의 육중한 전신.\n\nLOCATION (lock): In the same secluded exterior corner of the abandoned library, where the private phone call has just ended. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Continuous corner floor between foreground 현우 and 찰리's fully visible feet in the lower-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Secluded library corner (Shared by 현우 and the silently arrived 찰리) — The floor continues visibly from 현우's foreground position to 찰리's feet, with the corner behind them; used as Proves their immediate physical proximity and prevents the reveal from reading as a separate space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the corner's restrained early-morning ambient illumination, using gentle tonal separation to make the silent full-body reveal legible without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Haenam crayon mural and its written message remain intact inside the dusty library. The camper remains parked outside. 현우: He has finished the call and still possesses the disposable phone and the intact contact card taken from his shoe. His facial bruises and untreated leg wound remain. 찰리: He now stands close by in the secluded library corner, with his worn metal body and retained disguise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 뒤돌아선 현우의 등 바로 뒤에 소리 없이 묵묵히 서 있는 찰리의 육중한 전신.\n\nLOCATION (lock): In the same secluded exterior corner of the abandoned library, where the private phone call has just ended. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Continuous corner floor between foreground 현우 and 찰리's fully visible feet in the lower-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Secluded library corner (Shared by 현우 and the silently arrived 찰리) — The floor continues visibly from 현우's foreground position to 찰리's feet, with the corner behind them; used as Proves their immediate physical proximity and prevents the reveal from reading as a separate space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the corner's restrained early-morning ambient illumination, using gentle tonal separation to make the silent full-body reveal legible without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Haenam crayon mural and its written message remain intact inside the dusty library. The camper remains parked outside. 현우: He has finished the call and still possesses the disposable phone and the intact contact card taken from his shoe. His facial bruises and untreated leg wound remain. 찰리: He now stands close by in the secluded library corner, with his worn metal body and retained disguise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 뒤돌아선 현우의 등 바로 뒤에 소리 없이 묵묵히 서 있는 찰리의 육중한 전신.\n\nLOCATION (lock): In the same secluded exterior corner of the abandoned library, where the private phone call has just ended. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Continuous corner floor between foreground 현우 and 찰리's fully visible feet in the lower-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Secluded library corner (Shared by 현우 and the silently arrived 찰리) — The floor continues visibly from 현우's foreground position to 찰리's feet, with the corner behind them; used as Proves their immediate physical proximity and prevents the reveal from reading as a separate space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the corner's restrained early-morning ambient illumination, using gentle tonal separation to make the silent full-body reveal legible without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Haenam crayon mural and its written message remain intact inside the dusty library. The camper remains parked outside. 현우: He has finished the call and still possesses the disposable phone and the intact contact card taken from his shoe. His facial bruises and untreated leg wound remain. 찰리: He now stands close by in the secluded library corner, with his worn metal body and retained disguise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우는 왼쪽을 바라보며 서 있고, 찰리는 정면을 향해 서 있습니다.",
    "built_space": "야외가 아닌, 타일 바닥과 어린이들의 크레용 그림이 그려진 실내 공간으로 묘사되었습니다. 기준 이미지의 환경과 전혀 일치하지 않습니다.",
    "entities": "현우는 뒷모습으로 서서 휴대전화를 들고 있으며 다리에 핏자국이 보입니다. 찰리는 로봇 형태는 갖추었으나 프롬프트에서 지시한 위장용 망토를 착용하지 않은 상태입니다.",
    "hard_violations": [
     "[gemini-pro] 지정된 장소(야외 콘크리트 구석)를 무시하고 완전히 다른 실내 공간을 생성함",
     "[gemini-pro] 찰리가 '현우의 등 바로 뒤'가 아닌 현우의 우측에 나란히 배치됨",
     "[gpt-high] 두 인물을 잠금된 도서관 외부 모퉁이가 아니라 천장과 걸레받이가 있는 벽화 실내에 배치했다.",
     "[gpt-high] 현우를 왼쪽 뒤편에, 찰리를 오른쪽 전경에 놓아 명시된 전경 현우·하단 중앙 중경 찰리의 배치를 뒤집었다."
    ],
    "physics": "두 인물 모두 바닥에 정상적으로 서 있으며, 지지 상태에 물리적인 이상은 없습니다."
   },
   {
    "label": "B",
    "direction": "현우는 뒤돌아서서 아래의 휴대전화를 내려다보고 있으며, 찰리는 그 바로 뒤에서 앞을 향해 묵묵히 서 있습니다.",
    "built_space": "이전 샷 기준 이미지의 콘크리트 벽면, 배수관, 우측 상단의 창문, 오염되고 이끼 낀 바닥 등 야외 구석의 디테일을 매우 정확하게 재현했습니다.",
    "entities": "현우는 지정된 복장으로 등 돌려 서 있으며, 왼팔 아래로 휴대전화가 보입니다. 찰리는 현우 바로 뒤에 위치해 있으나, 지시된 위장용 망토(disguise)를 두르지 않고 금속 몸체를 그대로 드러내고 있습니다.",
    "hard_violations": [
     "[gpt-high] 찰리가 현우의 등 바로 뒤가 아니라 현우와 마주 보는 앞쪽에 서 있어, 핵심 인물 배치가 반대로 구현되었다."
    ],
    "physics": "두 캐릭터 모두 콘크리트 바닥 위에 무게 중심을 두고 안정적으로 서 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "기준 이미지의 야외 배경 요소와 현우의 등 바로 뒤에 찰리가 서 있는 구도를 훌륭하게 구현했으나, 찰리의 위장용 망토가 누락되어 일부 감점되었습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지정된 야외 환경을 완전히 무시하고 실내 공간으로 묘사하였으며, 찰리가 현우의 등 뒤가 아닌 옆에 서 있어 핵심 구도와 배경을 모두 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 왼쪽을 바라보며 서 있고, 찰리는 정면을 향해 서 있습니다.",
        "built_space": "야외가 아닌, 타일 바닥과 어린이들의 크레용 그림이 그려진 실내 공간으로 묘사되었습니다. 기준 이미지의 환경과 전혀 일치하지 않습니다.",
        "entities": "현우는 뒷모습으로 서서 휴대전화를 들고 있으며 다리에 핏자국이 보입니다. 찰리는 로봇 형태는 갖추었으나 프롬프트에서 지시한 위장용 망토를 착용하지 않은 상태입니다.",
        "hard_violations": [
         "지정된 장소(야외 콘크리트 구석)를 무시하고 완전히 다른 실내 공간을 생성함",
         "찰리가 '현우의 등 바로 뒤'가 아닌 현우의 우측에 나란히 배치됨"
        ],
        "physics": "두 인물 모두 바닥에 정상적으로 서 있으며, 지지 상태에 물리적인 이상은 없습니다."
       },
       {
        "label": "B",
        "direction": "현우는 뒤돌아서서 아래의 휴대전화를 내려다보고 있으며, 찰리는 그 바로 뒤에서 앞을 향해 묵묵히 서 있습니다.",
        "built_space": "이전 샷 기준 이미지의 콘크리트 벽면, 배수관, 우측 상단의 창문, 오염되고 이끼 낀 바닥 등 야외 구석의 디테일을 매우 정확하게 재현했습니다.",
        "entities": "현우는 지정된 복장으로 등 돌려 서 있으며, 왼팔 아래로 휴대전화가 보입니다. 찰리는 현우 바로 뒤에 위치해 있으나, 지시된 위장용 망토(disguise)를 두르지 않고 금속 몸체를 그대로 드러내고 있습니다.",
        "hard_violations": [],
        "physics": "두 캐릭터 모두 콘크리트 바닥 위에 무게 중심을 두고 안정적으로 서 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "기준 이미지의 야외 배경 요소와 현우의 등 바로 뒤에 찰리가 서 있는 구도를 훌륭하게 구현했으나, 찰리의 위장용 망토가 누락되어 일부 감점되었습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지정된 야외 환경을 완전히 무시하고 실내 공간으로 묘사하였으며, 찰리가 현우의 등 뒤가 아닌 옆에 서 있어 핵심 구도와 배경을 모두 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 왼쪽을 바라보며 서 있고, 찰리는 정면을 향해 서 있습니다.",
        "built_space": "야외가 아닌, 타일 바닥과 어린이들의 크레용 그림이 그려진 실내 공간으로 묘사되었습니다. 기준 이미지의 환경과 전혀 일치하지 않습니다.",
        "entities": "현우는 뒷모습으로 서서 휴대전화를 들고 있으며 다리에 핏자국이 보입니다. 찰리는 로봇 형태는 갖추었으나 프롬프트에서 지시한 위장용 망토를 착용하지 않은 상태입니다.",
        "hard_violations": [
         "지정된 장소(야외 콘크리트 구석)를 무시하고 완전히 다른 실내 공간을 생성함",
         "찰리가 '현우의 등 바로 뒤'가 아닌 현우의 우측에 나란히 배치됨"
        ],
        "physics": "두 인물 모두 바닥에 정상적으로 서 있으며, 지지 상태에 물리적인 이상은 없습니다."
       },
       {
        "label": "B",
        "direction": "현우는 뒤돌아서서 아래의 휴대전화를 내려다보고 있으며, 찰리는 그 바로 뒤에서 앞을 향해 묵묵히 서 있습니다.",
        "built_space": "이전 샷 기준 이미지의 콘크리트 벽면, 배수관, 우측 상단의 창문, 오염되고 이끼 낀 바닥 등 야외 구석의 디테일을 매우 정확하게 재현했습니다.",
        "entities": "현우는 지정된 복장으로 등 돌려 서 있으며, 왼팔 아래로 휴대전화가 보입니다. 찰리는 현우 바로 뒤에 위치해 있으나, 지시된 위장용 망토(disguise)를 두르지 않고 금속 몸체를 그대로 드러내고 있습니다.",
        "hard_violations": [],
        "physics": "두 캐릭터 모두 콘크리트 바닥 위에 무게 중심을 두고 안정적으로 서 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "외부 모퉁이와 전경 현우에서 찰리의 발까지 이어지는 바닥은 맞지만, 찰리가 현우의 등 뒤가 아니라 앞에 서 있고 줄무늬 위장 천도 없다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "외부 모퉁이를 벽화가 있는 실내로 바꾸고 찰리를 오른쪽 전경에 배치하여 장소 잠금과 전경 현우·중경 찰리 구도를 모두 어겼다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 카메라에 등을 보인 채 오른손의 휴대전화 쪽으로 고개를 숙인다. 현우의 몸 앞쪽에 찰리가 마주 서 있으므로, 화면상 뒤에 있다는 사실이 현우의 등 바로 뒤라는 신체적 관계를 충족하지는 않는다. 찰리의 얼굴은 현우의 머리 너머 카메라 쪽을 향한다. 무기나 이동 동작은 없다.",
        "built_space": "오염된 콘크리트 외벽 두 면과 연속된 젖은 바닥이 보인다. 왼쪽 벽에 굵은 배수관 하나와 가는 배관 하나, 뒤쪽 모서리에 배수관 하나, 오른쪽 위에 창문 한 구획이 있다. 이전 사진의 콘크리트·이끼·창문·모서리 배수관과 대체로 이어지는 외부 공간이다. 현우는 중앙 전경, 찰리는 중앙 중경에 있으며 두 사람 사이 바닥과 찰리의 양발이 보인다. 이전 사진의 오른쪽 환기 루버는 이 화면에서 확인되지 않는다.",
        "entities": "등을 보인 젊은 동아시아계 남성 한 명과 찰리 한 대가 있다. 현우의 검은 머리, 회색 셔츠, 올리브색 화물 바지와 다리의 혈흔은 설정에 부합하지만 얼굴 동일성과 멍은 확인하기 어렵다. 손에는 작은 은색·검은색 휴대전화가 있고 온전한 연락 카드는 식별되지 않는다. 찰리는 닳은 샌드 베이지 장갑과 흰 마스크형 얼굴, 긴 팔을 갖췄으나 참고보다 다리가 길고, 참고의 줄무늬 위장 천과 매듭이 없다. 실내 벽화와 외부 캠핑카는 프레임 밖이라 평가하지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "찰리가 현우의 등 바로 뒤가 아니라 현우와 마주 보는 앞쪽에 서 있어, 핵심 인물 배치가 반대로 구현되었다."
        ],
        "physics": "현우는 두 신발을 바닥에 딛고 서 있고 휴대전화는 오른손으로 잡고 있다. 화면에 드러난 전화 면은 현우 쪽으로 기울어져 사용 방향에 뚜렷한 모순이 없다. 찰리도 양발로 바닥을 지지하고 양팔은 어깨 관절에서 자연스럽게 내려온다. 떠 있거나 지지 없이 놓인 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 몸과 고개를 왼쪽 벽 방향으로 돌리고 전화는 오른손에 내려 들고 있다. 찰리는 현우를 향하기보다 카메라를 정면으로 바라본다. 찰리가 현우의 오른쪽 뒤편에 걸치기는 하지만 두 인물이 좌우로 분리되어, 등 바로 뒤에서 드러나는 관계가 명확하지 않다. 무기나 이동 동작은 없다.",
        "built_space": "천장, 두 면의 실내 벽, 검은 걸레받이와 마른 바닥이 보이며 벽에는 여러 어린이 그림이 이어진다. 이전 사진의 외부 창문, 모서리 배수관, 환기 루버는 보이지 않고 건축 마감 자체가 실내로 바뀌었다. 현우는 왼쪽 뒤편, 찰리는 오른쪽 전경을 크게 차지한다. 바닥은 연결되지만 요구된 전경 현우와 하단 중앙 중경의 찰리라는 배치가 뒤집혔다.",
        "entities": "젊은 동아시아계 남성 현우와 베이지색 기계 몸체의 찰리만 보인다. 현우의 검은 머리, 회색 셔츠, 올리브색 바지와 다리 혈흔은 대체로 맞지만 셔츠를 넣어 입은 모습은 이전 사진의 느슨한 착장과 다르다. 얼굴은 일부만 보여 정확한 나이와 멍은 확정하기 어렵다. 오른손에 작은 전화가 있고 연락 카드는 식별되지 않는다. 찰리의 긴 팔, 짧은 다리, 흰 얼굴과 마모된 장갑은 참고에 가깝지만 줄무늬 위장 천과 매듭은 없다. 벽화는 등장하지만 이 장면의 외부 장소에 있어야 할 배경은 아니다. 캠핑카와 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "두 인물을 잠금된 도서관 외부 모퉁이가 아니라 천장과 걸레받이가 있는 벽화 실내에 배치했다.",
         "현우를 왼쪽 뒤편에, 찰리를 오른쪽 전경에 놓아 명시된 전경 현우·하단 중앙 중경 찰리의 배치를 뒤집었다."
        ],
        "physics": "현우의 양발은 바닥에 닿고 전화는 오른손에 잡혀 있다. 찰리의 넓게 벌린 양발도 바닥에 접촉하며 무게를 지지한다. 긴 팔과 손은 몸체 관절에 연결되어 내려와 있다. 두 인물 모두 지지 없는 부유나 불가능한 접촉은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "외부 모퉁이와 전경 현우에서 찰리의 발까지 이어지는 바닥은 맞지만, 찰리가 현우의 등 뒤가 아니라 앞에 서 있고 줄무늬 위장 천도 없다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "외부 모퉁이를 벽화가 있는 실내로 바꾸고 찰리를 오른쪽 전경에 배치하여 장소 잠금과 전경 현우·중경 찰리 구도를 모두 어겼다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 카메라에 등을 보인 채 오른손의 휴대전화 쪽으로 고개를 숙인다. 현우의 몸 앞쪽에 찰리가 마주 서 있으므로, 화면상 뒤에 있다는 사실이 현우의 등 바로 뒤라는 신체적 관계를 충족하지는 않는다. 찰리의 얼굴은 현우의 머리 너머 카메라 쪽을 향한다. 무기나 이동 동작은 없다.",
        "built_space": "오염된 콘크리트 외벽 두 면과 연속된 젖은 바닥이 보인다. 왼쪽 벽에 굵은 배수관 하나와 가는 배관 하나, 뒤쪽 모서리에 배수관 하나, 오른쪽 위에 창문 한 구획이 있다. 이전 사진의 콘크리트·이끼·창문·모서리 배수관과 대체로 이어지는 외부 공간이다. 현우는 중앙 전경, 찰리는 중앙 중경에 있으며 두 사람 사이 바닥과 찰리의 양발이 보인다. 이전 사진의 오른쪽 환기 루버는 이 화면에서 확인되지 않는다.",
        "entities": "등을 보인 젊은 동아시아계 남성 한 명과 찰리 한 대가 있다. 현우의 검은 머리, 회색 셔츠, 올리브색 화물 바지와 다리의 혈흔은 설정에 부합하지만 얼굴 동일성과 멍은 확인하기 어렵다. 손에는 작은 은색·검은색 휴대전화가 있고 온전한 연락 카드는 식별되지 않는다. 찰리는 닳은 샌드 베이지 장갑과 흰 마스크형 얼굴, 긴 팔을 갖췄으나 참고보다 다리가 길고, 참고의 줄무늬 위장 천과 매듭이 없다. 실내 벽화와 외부 캠핑카는 프레임 밖이라 평가하지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "찰리가 현우의 등 바로 뒤가 아니라 현우와 마주 보는 앞쪽에 서 있어, 핵심 인물 배치가 반대로 구현되었다."
        ],
        "physics": "현우는 두 신발을 바닥에 딛고 서 있고 휴대전화는 오른손으로 잡고 있다. 화면에 드러난 전화 면은 현우 쪽으로 기울어져 사용 방향에 뚜렷한 모순이 없다. 찰리도 양발로 바닥을 지지하고 양팔은 어깨 관절에서 자연스럽게 내려온다. 떠 있거나 지지 없이 놓인 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 몸과 고개를 왼쪽 벽 방향으로 돌리고 전화는 오른손에 내려 들고 있다. 찰리는 현우를 향하기보다 카메라를 정면으로 바라본다. 찰리가 현우의 오른쪽 뒤편에 걸치기는 하지만 두 인물이 좌우로 분리되어, 등 바로 뒤에서 드러나는 관계가 명확하지 않다. 무기나 이동 동작은 없다.",
        "built_space": "천장, 두 면의 실내 벽, 검은 걸레받이와 마른 바닥이 보이며 벽에는 여러 어린이 그림이 이어진다. 이전 사진의 외부 창문, 모서리 배수관, 환기 루버는 보이지 않고 건축 마감 자체가 실내로 바뀌었다. 현우는 왼쪽 뒤편, 찰리는 오른쪽 전경을 크게 차지한다. 바닥은 연결되지만 요구된 전경 현우와 하단 중앙 중경의 찰리라는 배치가 뒤집혔다.",
        "entities": "젊은 동아시아계 남성 현우와 베이지색 기계 몸체의 찰리만 보인다. 현우의 검은 머리, 회색 셔츠, 올리브색 바지와 다리 혈흔은 대체로 맞지만 셔츠를 넣어 입은 모습은 이전 사진의 느슨한 착장과 다르다. 얼굴은 일부만 보여 정확한 나이와 멍은 확정하기 어렵다. 오른손에 작은 전화가 있고 연락 카드는 식별되지 않는다. 찰리의 긴 팔, 짧은 다리, 흰 얼굴과 마모된 장갑은 참고에 가깝지만 줄무늬 위장 천과 매듭은 없다. 벽화는 등장하지만 이 장면의 외부 장소에 있어야 할 배경은 아니다. 캠핑카와 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "두 인물을 잠금된 도서관 외부 모퉁이가 아니라 천장과 걸레받이가 있는 벽화 실내에 배치했다.",
         "현우를 왼쪽 뒤편에, 찰리를 오른쪽 전경에 놓아 명시된 전경 현우·하단 중앙 중경 찰리의 배치를 뒤집었다."
        ],
        "physics": "현우의 양발은 바닥에 닿고 전화는 오른손에 잡혀 있다. 찰리의 넓게 벌린 양발도 바닥에 접촉하며 무게를 지지한다. 긴 팔과 손은 몸체 관절에 연결되어 내려와 있다. 두 인물 모두 지지 없는 부유나 불가능한 접촉은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.75,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.5,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 장소(야외 콘크리트 구석)를 무시하고 완전히 다른 실내 공간을 생성함",
     "[gemini-pro] 찰리가 '현우의 등 바로 뒤'가 아닌 현우의 우측에 나란히 배치됨",
     "[gpt-high] 두 인물을 잠금된 도서관 외부 모퉁이가 아니라 천장과 걸레받이가 있는 벽화 실내에 배치했다.",
     "[gpt-high] 현우를 왼쪽 뒤편에, 찰리를 오른쪽 전경에 놓아 명시된 전경 현우·하단 중앙 중경 찰리의 배치를 뒤집었다."
    ],
    "B": [
     "[gpt-high] 찰리가 현우의 등 바로 뒤가 아니라 현우와 마주 보는 앞쪽에 서 있어, 핵심 인물 배치가 반대로 구현되었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "기준 이미지의 야외 배경 요소와 현우의 등 바로 뒤에 찰리가 서 있는 구도를 훌륭하게 구현했으나, 찰리의 위장용 망토가 누락되어 일부 감점되었습니다.  ★위반: [gpt-high] 찰리가 현우의 등 바로 뒤가 아니라 현우와 마주 보는 앞쪽에 서 있어, 핵심 인물 배치가 반대로 구현되었다."
   },
   {
    "label": "A",
    "score": 500,
    "verdict_ko": "지정된 야외 환경을 완전히 무시하고 실내 공간으로 묘사하였으며, 찰리가 현우의 등 뒤가 아닌 옆에 서 있어 핵심 구도와 배경을 모두 위반했습니다.  ★위반: [gemini-pro] 지정된 장소(야외 콘크리트 구석)를 무시하고 완전히 다른 실내 공간을 생성함 / [gemini-pro] 찰리가 '현우의 등 바로 뒤'가 아닌 현우의 우측에 나란히 배치됨 / [gpt-high] 두 인물을 잠금된 도서관 외부 모퉁이가 아니라 천장과 걸레받이가 있는 벽화 실내에 배치했다. / [gpt-high] 현우를 왼쪽 뒤편에, 찰리를 오른쪽 전경에 놓아 명시된 전경 현우·하단 중앙 중경 찰리의 배치를 뒤집었다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9_sel.png",
    "asset_id": "f612becd-13ee-4dbb-8fe9-2954f0cc65f6",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-d9a5-7237-83b7-c12f5a8e35f1",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S47sh9"
  }
 },
 "S47sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:10:58.859266+00:00",
  "fingerprint": "fd318ddda67e91538567f7aa26f05df6cb864532f9fe1214e5beabea01ff3864",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S47sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S47sh12_sel.png",
  "source_sha256": "b2a4838a613f33d332e7a9069c32f669eee31791253db1309dadb2f2845530d5",
  "file": "S47sh12_cine.png",
  "staged_sha256": "c11dbe4c39b6aebe745d31e603d1cde190cbc7b9d8f41d724fe51c63d4caa0b9",
  "latency_ms": 9785
 },
 "S48sh5::confined_fp_apt": {
  "applies": true,
  "reason_ko": "이 샷은 캠핑카 내부 조수석을 배경으로 하며, 앰버가 조수석에 앉아 앞유리를 통해 전방의 경찰 검문소를 바라보는 상황입니다. 차량 내부의 정확한 좌석 배치와 인물의 시선 방향이 어긋나면 화면의 일관성과 몰입을 해칠 수 있으므로 평면도 형태의 레이아웃 가이드가 필요합니다.",
  "input_fingerprint": "2d90f3866d2fd521"
 },
 "S48sh5::signage": {
  "fp": "3f8891c62006ae26",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::0261ae55cee7": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_0261ae55cee7.png",
  "place_text": "Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight.",
  "input_fingerprint": "a8d5a55357967c68"
 },
 "S48sh16::signage": {
  "fp": "cebdca82f1f62a9c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S48sh16": {
  "input_fingerprint": "613a884cea18e021",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우가 탄 캠핑카가 도로변을 향해 차체가 기울어진 채 거친 흙먼지를 뒤로 뿜어내며 멀어지는 중간 순간의 뒷모습.\n\nLOCATION (lock): On a rough mountain access track descending toward the roadside, where the camper throws up dust. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Camper (Descending toward the roadside with its body tilted) — Rear and passenger-side flank face the elevated camera; used as Provides the receding subject and a readable measure of the uneven descent; Mountain track (Used by the departing camper) — Runs from the lower foreground toward the upper-right distance; used as Carries the departure direction through the wide composition; Trailing earth dust (Thrown up behind the moving camper); used as Marks the vehicle's recent path without obscuring its rear silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled tonal separation keep the departing vehicle legible against the mountain track.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper's wheel fasteners have just been tightened, but its tire remains punctured and its fuel is nearly exhausted; the earlier roof and window leaks have not been repaired. Charlie retains his blanket covering and worn metal exterior on the mountain path.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우가 탄 캠핑카가 도로변을 향해 차체가 기울어진 채 거친 흙먼지를 뒤로 뿜어내며 멀어지는 중간 순간의 뒷모습.\n\nLOCATION (lock): On a rough mountain access track descending toward the roadside, where the camper throws up dust. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Camper (Descending toward the roadside with its body tilted) — Rear and passenger-side flank face the elevated camera; used as Provides the receding subject and a readable measure of the uneven descent; Mountain track (Used by the departing camper) — Runs from the lower foreground toward the upper-right distance; used as Carries the departure direction through the wide composition; Trailing earth dust (Thrown up behind the moving camper); used as Marks the vehicle's recent path without obscuring its rear silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled tonal separation keep the departing vehicle legible against the mountain track.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper's wheel fasteners have just been tightened, but its tire remains punctured and its fuel is nearly exhausted; the earlier roof and window leaks have not been repaired. Charlie retains his blanket covering and worn metal exterior on the mountain path.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우가 탄 캠핑카가 도로변을 향해 차체가 기울어진 채 거친 흙먼지를 뒤로 뿜어내며 멀어지는 중간 순간의 뒷모습.\n\nLOCATION (lock): On a rough mountain access track descending toward the roadside, where the camper throws up dust. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Camper (Descending toward the roadside with its body tilted) — Rear and passenger-side flank face the elevated camera; used as Provides the receding subject and a readable measure of the uneven descent; Mountain track (Used by the departing camper) — Runs from the lower foreground toward the upper-right distance; used as Carries the departure direction through the wide composition; Trailing earth dust (Thrown up behind the moving camper); used as Marks the vehicle's recent path without obscuring its rear silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled tonal separation keep the departing vehicle legible against the mountain track.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper's wheel fasteners have just been tightened, but its tire remains punctured and its fuel is nearly exhausted; the earlier roof and window leaks have not been repaired. Charlie retains his blanket covering and worn metal exterior on the mountain path.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "캠핑카가 우측 상단 원경의 포장도로를 향해 주행 중.",
    "built_space": "흙길이 전경에서 우측 상단의 도로를 향해 완만한 내리막 혹은 평행을 이루며 이어짐.",
    "entities": "캠핑카, 뒷면에 담요가 덮인 구조물, 흙먼지가 프롬프트와 일치함.",
    "hard_violations": [],
    "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 일으키며 이동 중."
   },
   {
    "label": "B",
    "direction": "캠핑카가 하단의 포장도로를 등지고 우측 상단의 산 위를 향해 주행 중.",
    "built_space": "포장도로가 하단 전경에 위치하고 흙길이 우측 상단으로 오르막을 형성함.",
    "entities": "캠핑카와 흙먼지는 존재하나, 담요가 덮인 구조물이 없고 원경에 정체불명의 사람/오토바이가 있음.",
    "hard_violations": [
     "[gemini-pro] 프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
     "[gemini-pro] 수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함",
     "[gpt-high] 도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
     "[gpt-high] 사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
    ],
    "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 내며 주행 중."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도, 내리막길 설정, 뒷면의 담요 덮인 물체 등 프롬프트의 요구사항을 충실히 반영했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "원경에 사람이 등장하여 인물 배제 규칙을 위반했으며, 도로를 향하는 내리막이 아닌 산을 향하는 오르막을 주행하여 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 우측 상단 원경의 포장도로를 향해 주행 중.",
        "built_space": "흙길이 전경에서 우측 상단의 도로를 향해 완만한 내리막 혹은 평행을 이루며 이어짐.",
        "entities": "캠핑카, 뒷면에 담요가 덮인 구조물, 흙먼지가 프롬프트와 일치함.",
        "hard_violations": [],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 일으키며 이동 중."
       },
       {
        "label": "B",
        "direction": "캠핑카가 하단의 포장도로를 등지고 우측 상단의 산 위를 향해 주행 중.",
        "built_space": "포장도로가 하단 전경에 위치하고 흙길이 우측 상단으로 오르막을 형성함.",
        "entities": "캠핑카와 흙먼지는 존재하나, 담요가 덮인 구조물이 없고 원경에 정체불명의 사람/오토바이가 있음.",
        "hard_violations": [
         "프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
         "수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함"
        ],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 내며 주행 중."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도, 내리막길 설정, 뒷면의 담요 덮인 물체 등 프롬프트의 요구사항을 충실히 반영했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "원경에 사람이 등장하여 인물 배제 규칙을 위반했으며, 도로를 향하는 내리막이 아닌 산을 향하는 오르막을 주행하여 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 우측 상단 원경의 포장도로를 향해 주행 중.",
        "built_space": "흙길이 전경에서 우측 상단의 도로를 향해 완만한 내리막 혹은 평행을 이루며 이어짐.",
        "entities": "캠핑카, 뒷면에 담요가 덮인 구조물, 흙먼지가 프롬프트와 일치함.",
        "hard_violations": [],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 일으키며 이동 중."
       },
       {
        "label": "B",
        "direction": "캠핑카가 하단의 포장도로를 등지고 우측 상단의 산 위를 향해 주행 중.",
        "built_space": "포장도로가 하단 전경에 위치하고 흙길이 우측 상단으로 오르막을 형성함.",
        "entities": "캠핑카와 흙먼지는 존재하나, 담요가 덮인 구조물이 없고 원경에 정체불명의 사람/오토바이가 있음.",
        "hard_violations": [
         "프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
         "수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함"
        ],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 내며 주행 중."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "차량이 도로를 향해 내려가는 대신 도로에서 산길 위로 올라가며, 원경 인물과 낮고 가까운 시점도 지정된 무인 와이드 구도에 어긋납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "높은 시점에서 후면과 조수석 측면을 보이며 우상단 도로로 내려가는 와이드 구도를 충실히 구현하지만, 차체 기울기는 약하고 찰리의 정체는 불명확합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 화면 우상단의 오르막 산길을 향하고, 먼지는 뒤쪽인 좌하단으로 뻗습니다. 포장도로가 전경과 좌측 아래에 있어 차량은 도로변으로 내려가는 것이 아니라 도로에서 멀어져 산으로 올라갑니다. 우상단 길에는 차량에 등을 돌린 듯한 작은 인물이 보이며 시선은 판별되지 않습니다.",
        "built_space": "포장도로가 화면 하단을 가로지르고 흙길 한 갈래가 우상단 산비탈로 올라갑니다. 캠핑카에는 후면 창 하나와 조수석 쪽 출입문 하나가 보입니다. 차량이 화면을 크게 차지하며 지붕 윗면이 거의 보이지 않는 낮은 시점으로, 요구한 높은 카메라의 와이드 구도와 다릅니다.",
        "entities": "낡은 흰색 캠핑카, 돌이 많은 산길, 흙먼지와 낮의 산악 환경은 보입니다. 원경에 회색 옷 또는 덮개를 두른 사람 형태 하나가 있으며 나이·성별·민족은 판별할 수 없습니다. 찰리의 마모된 금속 외장은 확인되지 않습니다. 보이는 타이어는 뚜렷하게 주저앉지 않았고, 연료량·체결 상태·누수 지속 여부는 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
         "사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
        ],
        "physics": "차체는 왼쪽으로 크게 기울고 바퀴들은 요철이 있는 지면에 닿거나 바로 위를 통과하는 모습입니다. 바퀴 주변에서 뒤로 튀는 흙과 먼지는 거친 길을 달리는 동작으로 설명됩니다. 기울기 자체를 불가능한 자세로 볼 근거는 없으며, 원경 인물도 길 위에 서 있습니다."
       },
       {
        "label": "B",
        "direction": "캠핑카는 화면 우상단의 포장도로 접속부를 향해 멀어지고 있습니다. 후면과 조수석 측면이 카메라를 향하며, 흙먼지는 지나온 길인 좌하단으로 이어져 후면 윤곽을 크게 가리지 않습니다. 노출된 사람이나 시선은 없습니다.",
        "built_space": "거친 흙길 한 갈래가 하단 전경에서 우상단 도로 접속부까지 이어지고, 도착 지점의 포장도로와 가드레일이 보입니다. 높은 카메라에서 차량 지붕·후면·조수석 측면을 함께 봅니다. 후면 창 하나와 측면 출입문 하나가 있으며, 뒤쪽에는 금속 상자 형태의 적재물이 묶여 있습니다. 넓은 주변 지형이 포함되어 요구한 와이드 구도에 부합합니다.",
        "entities": "낡은 캠핑카, 산악 비포장길, 뒤로 퍼지는 흙먼지와 주간 환경이 모두 보입니다. 후면 적재물은 담요가 덮인 낡은 금속 외장이어서 찰리의 재질·덮개 조건에는 대응하지만, 상자 같은 형태라 찰리 자체인지는 확정하기 어렵습니다. 뒤 타이어는 접지부가 눌려 보이나 펑크 여부는 단정할 수 없습니다. 연료량·볼트 체결·누수 상태는 확인되지 않습니다. 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "차량 바퀴가 울퉁불퉁한 흙길에 접지하고 차체가 약간 기울어 있어 내려가는 주행으로 읽힙니다. 먼지는 타이어 뒤에서 발생해 지나온 경로로 퍼집니다. 후면 금속 적재물은 받침과 감싼 고정끈으로 지지되며, 담요는 그 위에 걸쳐져 있어 지지 없이 떠 있는 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "차량이 도로를 향해 내려가는 대신 도로에서 산길 위로 올라가며, 원경 인물과 낮고 가까운 시점도 지정된 무인 와이드 구도에 어긋납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "높은 시점에서 후면과 조수석 측면을 보이며 우상단 도로로 내려가는 와이드 구도를 충실히 구현하지만, 차체 기울기는 약하고 찰리의 정체는 불명확합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 화면 우상단의 오르막 산길을 향하고, 먼지는 뒤쪽인 좌하단으로 뻗습니다. 포장도로가 전경과 좌측 아래에 있어 차량은 도로변으로 내려가는 것이 아니라 도로에서 멀어져 산으로 올라갑니다. 우상단 길에는 차량에 등을 돌린 듯한 작은 인물이 보이며 시선은 판별되지 않습니다.",
        "built_space": "포장도로가 화면 하단을 가로지르고 흙길 한 갈래가 우상단 산비탈로 올라갑니다. 캠핑카에는 후면 창 하나와 조수석 쪽 출입문 하나가 보입니다. 차량이 화면을 크게 차지하며 지붕 윗면이 거의 보이지 않는 낮은 시점으로, 요구한 높은 카메라의 와이드 구도와 다릅니다.",
        "entities": "낡은 흰색 캠핑카, 돌이 많은 산길, 흙먼지와 낮의 산악 환경은 보입니다. 원경에 회색 옷 또는 덮개를 두른 사람 형태 하나가 있으며 나이·성별·민족은 판별할 수 없습니다. 찰리의 마모된 금속 외장은 확인되지 않습니다. 보이는 타이어는 뚜렷하게 주저앉지 않았고, 연료량·체결 상태·누수 지속 여부는 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
         "사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
        ],
        "physics": "차체는 왼쪽으로 크게 기울고 바퀴들은 요철이 있는 지면에 닿거나 바로 위를 통과하는 모습입니다. 바퀴 주변에서 뒤로 튀는 흙과 먼지는 거친 길을 달리는 동작으로 설명됩니다. 기울기 자체를 불가능한 자세로 볼 근거는 없으며, 원경 인물도 길 위에 서 있습니다."
       },
       {
        "label": "A",
        "direction": "캠핑카는 화면 우상단의 포장도로 접속부를 향해 멀어지고 있습니다. 후면과 조수석 측면이 카메라를 향하며, 흙먼지는 지나온 길인 좌하단으로 이어져 후면 윤곽을 크게 가리지 않습니다. 노출된 사람이나 시선은 없습니다.",
        "built_space": "거친 흙길 한 갈래가 하단 전경에서 우상단 도로 접속부까지 이어지고, 도착 지점의 포장도로와 가드레일이 보입니다. 높은 카메라에서 차량 지붕·후면·조수석 측면을 함께 봅니다. 후면 창 하나와 측면 출입문 하나가 있으며, 뒤쪽에는 금속 상자 형태의 적재물이 묶여 있습니다. 넓은 주변 지형이 포함되어 요구한 와이드 구도에 부합합니다.",
        "entities": "낡은 캠핑카, 산악 비포장길, 뒤로 퍼지는 흙먼지와 주간 환경이 모두 보입니다. 후면 적재물은 담요가 덮인 낡은 금속 외장이어서 찰리의 재질·덮개 조건에는 대응하지만, 상자 같은 형태라 찰리 자체인지는 확정하기 어렵습니다. 뒤 타이어는 접지부가 눌려 보이나 펑크 여부는 단정할 수 없습니다. 연료량·볼트 체결·누수 상태는 확인되지 않습니다. 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "차량 바퀴가 울퉁불퉁한 흙길에 접지하고 차체가 약간 기울어 있어 내려가는 주행으로 읽힙니다. 먼지는 타이어 뒤에서 발생해 지나온 경로로 퍼집니다. 후면 금속 적재물은 받침과 감싼 고정끈으로 지지되며, 담요는 그 위에 걸쳐져 있어 지지 없이 떠 있는 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.411
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.161
   },
   "violations": {
    "B": [
     "[gemini-pro] 프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
     "[gemini-pro] 수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함",
     "[gpt-high] 도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
     "[gpt-high] 사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 161
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "구도, 내리막길 설정, 뒷면의 담요 덮인 물체 등 프롬프트의 요구사항을 충실히 반영했습니다."
   },
   {
    "label": "B",
    "score": 161,
    "verdict_ko": "원경에 사람이 등장하여 인물 배제 규칙을 위반했으며, 도로를 향하는 내리막이 아닌 산을 향하는 오르막을 주행하여 실격입니다.  ★위반: [gemini-pro] 프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함 / [gemini-pro] 수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함 / [gpt-high] 도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다. / [gpt-high] 사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-de83-7361-afda-2d174525d4fc",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S48sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:51:20.122995+00:00",
  "fingerprint": "8fe80be0893e7edc7850aad071e6b9afa0927b1edcd0b211e9ba5c540e29b99d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S48sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S48sh16_sel.png",
  "source_sha256": "9847de5c4c836208861c95d23b0219ab5422adbb23b27f989b3873e17bdc769f",
  "file": "S48sh16_cine.png",
  "staged_sha256": "5d44b78e6e2b5b06c7581e2ed064f463f961fd8a4a2a82aed672f62a9f91881a",
  "latency_ms": 11956
 },
 "S48sh22::signage": {
  "fp": "7baa36487147b709",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S48sh22": {
  "input_fingerprint": "95375d3e74a6fbe1",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 손을 양손으로 꽉 감싸 쥔 채 따뜻하게 올려다보는 앰버의 다정한 상체.\n\nLOCATION (lock): On the mountain track descending toward the main road, where the children and robot walk behind the camper. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Mountain path (The group has been descending along it); used as Remains a subdued strip of spatial context behind the promise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Softly controlled daylight keeps the human hands and metal hand equally readable, allowing tenderness to come from their contact rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains compromised by a punctured tire and very low fuel after its wheel fasteners were tightened. Charlie stands on the mountain path with his worn metal body and blanket covering, extending a large hand for the promise. 앰버: She is standing on the mountain path with her hand extended for a promise, still wearing the replacement shoes.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 손을 양손으로 꽉 감싸 쥔 채 따뜻하게 올려다보는 앰버의 다정한 상체.\n\nLOCATION (lock): On the mountain track descending toward the main road, where the children and robot walk behind the camper. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Mountain path (The group has been descending along it); used as Remains a subdued strip of spatial context behind the promise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Softly controlled daylight keeps the human hands and metal hand equally readable, allowing tenderness to come from their contact rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains compromised by a punctured tire and very low fuel after its wheel fasteners were tightened. Charlie stands on the mountain path with his worn metal body and blanket covering, extending a large hand for the promise. 앰버: She is standing on the mountain path with her hand extended for a promise, still wearing the replacement shoes.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 손을 양손으로 꽉 감싸 쥔 채 따뜻하게 올려다보는 앰버의 다정한 상체.\n\nLOCATION (lock): On the mountain track descending toward the main road, where the children and robot walk behind the camper. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Mountain path (The group has been descending along it); used as Remains a subdued strip of spatial context behind the promise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Softly controlled daylight keeps the human hands and metal hand equally readable, allowing tenderness to come from their contact rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains compromised by a punctured tire and very low fuel after its wheel fasteners were tightened. Charlie stands on the mountain path with his worn metal body and blanket covering, extending a large hand for the promise. 앰버: She is standing on the mountain path with her hand extended for a promise, still wearing the replacement shoes.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 왼쪽 위에 있는 찰리의 얼굴을 올려다보고, 찰리는 앰버 쪽으로 고개를 숙입니다. 찰리의 큰 손은 앰버의 가슴 앞까지 뻗어 있으며, 앰버의 두 손은 금속 엄지와 손등에 각각 닿습니다. 시선의 대상은 맞지만 두 손으로 꽉 감싸기보다 위에 얹어 잡는 모습입니다.",
    "built_space": "돌과 바퀴 자국이 있는 비포장 산길 뒤로 가드레일이 있는 포장도로가 이어지고, 캠핑카 한 대가 멀리 있습니다. 오른쪽은 돌이 드러난 비탈이며 전봇대도 배경에 보입니다. 두 인물은 길의 전경에 자리해 이전 장면의 공간과 대체로 연결되지만, 산길과 차량이 단순한 배경 띠보다 넓게 드러납니다. 두 인물 사이 아래쪽 길바닥에는 별도의 부츠 두 짝이 있습니다.",
    "entities": "금발의 어린 여자아이 한 명과 대형 로봇 한 대가 보입니다. 앰버의 나이대, 둥근 얼굴, 금발, 남색 반팔과 갈색 작업복은 참조와 대체로 맞지만 머리 위 고글은 없습니다. 혼혈 배경 자체는 외모만으로 확정할 수 없습니다. 찰리의 흰 마스크형 얼굴, 마모된 샌드 베이지 장갑판, 육중한 팔과 회색 줄무늬 담요는 참조에 부합합니다. 신발 착용 상태는 화면 밖이라 확인할 수 없으며, 바닥의 여분 부츠 한 켤레는 지시에 없는 물건입니다. 읽을 수 있는 글자는 없습니다.",
    "hard_violations": [
     "앰버와 찰리 사이의 길바닥에 지시되지 않은 여분 부츠 한 켤레를 추가했습니다."
    ],
    "physics": "찰리의 손은 손목과 전완, 팔꿈치로 이어져 지지되고, 앰버의 양손도 각 팔에 자연스럽게 연결되어 금속 손과 접촉합니다. 담요는 어깨와 몸통에 걸쳐져 있습니다. 두 인물의 발은 프레임 밖이지만 상체에 부유를 시사하는 모습은 없습니다. 여분 부츠는 길바닥에 놓여 있어 물리적 지지는 있으나 소품 추가가 문제입니다."
   },
   {
    "label": "A",
    "direction": "앰버는 왼쪽 위 찰리의 얼굴을 따뜻하게 올려다보고, 찰리도 앰버 쪽으로 얼굴을 내립니다. 찰리의 전완과 손은 앰버 앞으로 뻗어 있습니다. 앰버는 양손으로 금속 손의 윗부분과 엄지 부근을 잡지만, 거대한 손 전체를 꽉 감싸는 정도는 약합니다.",
    "built_space": "전경의 두 인물 뒤로 돌과 깊은 바퀴 자국이 있는 산길이 내려가며, 멀리 캠핑카 한 대와 가드레일이 있는 포장도로가 보입니다. 오른쪽 비탈, 뒤편 산 능선과 전봇대가 이전 장면의 장소를 이어 줍니다. 구조물의 중복이나 불가능한 반사는 없고 캠핑카의 원근 크기도 자연스럽습니다. 다만 산 능선과 도로가 넓게 보여 배경을 억제된 띠로 쓰라는 요구에는 덜 정확합니다.",
    "entities": "앰버 한 명과 찰리 한 대만 보이며 추가 인물이나 불필요한 전경 소품은 없습니다. 앰버는 참조와 유사한 금발의 약 10세 여자아이로, 둥근 얼굴과 남색 반팔, 갈색 작업복을 유지하지만 머리 위 고글은 빠졌습니다. 혼혈 배경은 외모만으로 단정할 수 없습니다. 찰리의 거대한 팔, 각진 베이지 장갑판, 흰 얼굴과 회색 줄무늬 담요가 참조와 잘 맞습니다. 두 인물의 하체와 앰버의 교체 신발은 구도 밖이므로 평가할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
    "hard_violations": [],
    "physics": "찰리의 큰 금속 손은 손목과 수평으로 뻗은 전완에 연결되고 팔꿈치가 이를 받칩니다. 앰버의 두 손은 굽힌 팔 끝에서 금속 손에 실제로 닿아 있으며 손가락의 잡는 자세도 가능합니다. 담요는 어깨에 걸려 중력 방향으로 늘어집니다. 발은 화면 밖이지만 두 인물의 직립 상체와 팔 자세에는 부유나 불가능한 지지 관계가 보이지 않습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "올려다보는 시선과 상체 구도는 맞지만, 산길 바닥에 지시되지 않은 부츠 한 켤레를 추가한 중대 위반이 있습니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "앰버의 상체와 찰리를 향한 다정한 시선, 양손의 금속 손 접촉을 충실히 구현했으나, 꽉 감싸 쥐는 동작은 다소 약하고 배경의 비중이 큽니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 왼쪽 위에 있는 찰리의 얼굴을 올려다보고, 찰리는 앰버 쪽으로 고개를 숙입니다. 찰리의 큰 손은 앰버의 가슴 앞까지 뻗어 있으며, 앰버의 두 손은 금속 엄지와 손등에 각각 닿습니다. 시선의 대상은 맞지만 두 손으로 꽉 감싸기보다 위에 얹어 잡는 모습입니다.",
        "built_space": "돌과 바퀴 자국이 있는 비포장 산길 뒤로 가드레일이 있는 포장도로가 이어지고, 캠핑카 한 대가 멀리 있습니다. 오른쪽은 돌이 드러난 비탈이며 전봇대도 배경에 보입니다. 두 인물은 길의 전경에 자리해 이전 장면의 공간과 대체로 연결되지만, 산길과 차량이 단순한 배경 띠보다 넓게 드러납니다. 두 인물 사이 아래쪽 길바닥에는 별도의 부츠 두 짝이 있습니다.",
        "entities": "금발의 어린 여자아이 한 명과 대형 로봇 한 대가 보입니다. 앰버의 나이대, 둥근 얼굴, 금발, 남색 반팔과 갈색 작업복은 참조와 대체로 맞지만 머리 위 고글은 없습니다. 혼혈 배경 자체는 외모만으로 확정할 수 없습니다. 찰리의 흰 마스크형 얼굴, 마모된 샌드 베이지 장갑판, 육중한 팔과 회색 줄무늬 담요는 참조에 부합합니다. 신발 착용 상태는 화면 밖이라 확인할 수 없으며, 바닥의 여분 부츠 한 켤레는 지시에 없는 물건입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "앰버와 찰리 사이의 길바닥에 지시되지 않은 여분 부츠 한 켤레를 추가했습니다."
        ],
        "physics": "찰리의 손은 손목과 전완, 팔꿈치로 이어져 지지되고, 앰버의 양손도 각 팔에 자연스럽게 연결되어 금속 손과 접촉합니다. 담요는 어깨와 몸통에 걸쳐져 있습니다. 두 인물의 발은 프레임 밖이지만 상체에 부유를 시사하는 모습은 없습니다. 여분 부츠는 길바닥에 놓여 있어 물리적 지지는 있으나 소품 추가가 문제입니다."
       },
       {
        "label": "B",
        "direction": "앰버는 왼쪽 위 찰리의 얼굴을 따뜻하게 올려다보고, 찰리도 앰버 쪽으로 얼굴을 내립니다. 찰리의 전완과 손은 앰버 앞으로 뻗어 있습니다. 앰버는 양손으로 금속 손의 윗부분과 엄지 부근을 잡지만, 거대한 손 전체를 꽉 감싸는 정도는 약합니다.",
        "built_space": "전경의 두 인물 뒤로 돌과 깊은 바퀴 자국이 있는 산길이 내려가며, 멀리 캠핑카 한 대와 가드레일이 있는 포장도로가 보입니다. 오른쪽 비탈, 뒤편 산 능선과 전봇대가 이전 장면의 장소를 이어 줍니다. 구조물의 중복이나 불가능한 반사는 없고 캠핑카의 원근 크기도 자연스럽습니다. 다만 산 능선과 도로가 넓게 보여 배경을 억제된 띠로 쓰라는 요구에는 덜 정확합니다.",
        "entities": "앰버 한 명과 찰리 한 대만 보이며 추가 인물이나 불필요한 전경 소품은 없습니다. 앰버는 참조와 유사한 금발의 약 10세 여자아이로, 둥근 얼굴과 남색 반팔, 갈색 작업복을 유지하지만 머리 위 고글은 빠졌습니다. 혼혈 배경은 외모만으로 단정할 수 없습니다. 찰리의 거대한 팔, 각진 베이지 장갑판, 흰 얼굴과 회색 줄무늬 담요가 참조와 잘 맞습니다. 두 인물의 하체와 앰버의 교체 신발은 구도 밖이므로 평가할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "찰리의 큰 금속 손은 손목과 수평으로 뻗은 전완에 연결되고 팔꿈치가 이를 받칩니다. 앰버의 두 손은 굽힌 팔 끝에서 금속 손에 실제로 닿아 있으며 손가락의 잡는 자세도 가능합니다. 담요는 어깨에 걸려 중력 방향으로 늘어집니다. 발은 화면 밖이지만 두 인물의 직립 상체와 팔 자세에는 부유나 불가능한 지지 관계가 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "올려다보는 시선과 상체 구도는 맞지만, 산길 바닥에 지시되지 않은 부츠 한 켤레를 추가한 중대 위반이 있습니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "앰버의 상체와 찰리를 향한 다정한 시선, 양손의 금속 손 접촉을 충실히 구현했으나, 꽉 감싸 쥐는 동작은 다소 약하고 배경의 비중이 큽니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 왼쪽 위에 있는 찰리의 얼굴을 올려다보고, 찰리는 앰버 쪽으로 고개를 숙입니다. 찰리의 큰 손은 앰버의 가슴 앞까지 뻗어 있으며, 앰버의 두 손은 금속 엄지와 손등에 각각 닿습니다. 시선의 대상은 맞지만 두 손으로 꽉 감싸기보다 위에 얹어 잡는 모습입니다.",
        "built_space": "돌과 바퀴 자국이 있는 비포장 산길 뒤로 가드레일이 있는 포장도로가 이어지고, 캠핑카 한 대가 멀리 있습니다. 오른쪽은 돌이 드러난 비탈이며 전봇대도 배경에 보입니다. 두 인물은 길의 전경에 자리해 이전 장면의 공간과 대체로 연결되지만, 산길과 차량이 단순한 배경 띠보다 넓게 드러납니다. 두 인물 사이 아래쪽 길바닥에는 별도의 부츠 두 짝이 있습니다.",
        "entities": "금발의 어린 여자아이 한 명과 대형 로봇 한 대가 보입니다. 앰버의 나이대, 둥근 얼굴, 금발, 남색 반팔과 갈색 작업복은 참조와 대체로 맞지만 머리 위 고글은 없습니다. 혼혈 배경 자체는 외모만으로 확정할 수 없습니다. 찰리의 흰 마스크형 얼굴, 마모된 샌드 베이지 장갑판, 육중한 팔과 회색 줄무늬 담요는 참조에 부합합니다. 신발 착용 상태는 화면 밖이라 확인할 수 없으며, 바닥의 여분 부츠 한 켤레는 지시에 없는 물건입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "앰버와 찰리 사이의 길바닥에 지시되지 않은 여분 부츠 한 켤레를 추가했습니다."
        ],
        "physics": "찰리의 손은 손목과 전완, 팔꿈치로 이어져 지지되고, 앰버의 양손도 각 팔에 자연스럽게 연결되어 금속 손과 접촉합니다. 담요는 어깨와 몸통에 걸쳐져 있습니다. 두 인물의 발은 프레임 밖이지만 상체에 부유를 시사하는 모습은 없습니다. 여분 부츠는 길바닥에 놓여 있어 물리적 지지는 있으나 소품 추가가 문제입니다."
       },
       {
        "label": "A",
        "direction": "앰버는 왼쪽 위 찰리의 얼굴을 따뜻하게 올려다보고, 찰리도 앰버 쪽으로 얼굴을 내립니다. 찰리의 전완과 손은 앰버 앞으로 뻗어 있습니다. 앰버는 양손으로 금속 손의 윗부분과 엄지 부근을 잡지만, 거대한 손 전체를 꽉 감싸는 정도는 약합니다.",
        "built_space": "전경의 두 인물 뒤로 돌과 깊은 바퀴 자국이 있는 산길이 내려가며, 멀리 캠핑카 한 대와 가드레일이 있는 포장도로가 보입니다. 오른쪽 비탈, 뒤편 산 능선과 전봇대가 이전 장면의 장소를 이어 줍니다. 구조물의 중복이나 불가능한 반사는 없고 캠핑카의 원근 크기도 자연스럽습니다. 다만 산 능선과 도로가 넓게 보여 배경을 억제된 띠로 쓰라는 요구에는 덜 정확합니다.",
        "entities": "앰버 한 명과 찰리 한 대만 보이며 추가 인물이나 불필요한 전경 소품은 없습니다. 앰버는 참조와 유사한 금발의 약 10세 여자아이로, 둥근 얼굴과 남색 반팔, 갈색 작업복을 유지하지만 머리 위 고글은 빠졌습니다. 혼혈 배경은 외모만으로 단정할 수 없습니다. 찰리의 거대한 팔, 각진 베이지 장갑판, 흰 얼굴과 회색 줄무늬 담요가 참조와 잘 맞습니다. 두 인물의 하체와 앰버의 교체 신발은 구도 밖이므로 평가할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "찰리의 큰 금속 손은 손목과 수평으로 뻗은 전완에 연결되고 팔꿈치가 이를 받칩니다. 앰버의 두 손은 굽힌 팔 끝에서 금속 손에 실제로 닿아 있으며 손가락의 잡는 자세도 가능합니다. 담요는 어깨에 걸려 중력 방향으로 늘어집니다. 발은 화면 밖이지만 두 인물의 직립 상체와 팔 자세에는 부유나 불가능한 지지 관계가 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "올려다보는 시선과 상체 구도는 맞지만, 산길 바닥에 지시되지 않은 부츠 한 켤레를 추가한 중대 위반이 있습니다."
   },
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "앰버의 상체와 찰리를 향한 다정한 시선, 양손의 금속 손 접촉을 충실히 구현했으나, 꽉 감싸 쥐는 동작은 다소 약하고 배경의 비중이 큽니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh16_sel.png",
    "asset_id": "4d47e73a-4dc9-4293-bf85-04b84ecae483",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-e01e-7d8d-87e0-6fc49c84cbd8",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S48sh16"
  }
 },
 "S48sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:13:01.579479+00:00",
  "fingerprint": "33046c4936fec649dfd9a6e593cded3f30d791635b5d58dbe07ae5397ba61942",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S48sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S48sh22_sel.png",
  "source_sha256": "dba9eb59551c4b8d081118bacd8f910a9431d2f6d928bf368a1fb8c871a87796",
  "file": "S48sh22_cine.png",
  "staged_sha256": "601b51a6864f3ca00d891695ebcf701c8b5baadebf85ec3528c3f756f7f88b27",
  "latency_ms": 10354
 },
 "S49sh14::signage": {
  "fp": "31fab31b484392d3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::0e2a5a685247648c": {
  "subjects": [],
  "subject_text": "휴게소 주유장과 마트 앞\n주유기와 미터기가 설치된 도로변 주유 공간. 안쪽에 작은 마트의 전면과 출입문이 있고 바깥 도로와 바로 연결된다.",
  "identity": "canonical",
  "scope_id": "L205",
  "scope_role": "location_exterior",
  "scope_sha": "8fffeb10f99bc0d1"
 },
 "S49sh14::bgfirst_bg": {
  "input_fingerprint": "726f908a2a433162",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh14__bgfirst_bg.png",
  "asset_id": "837d1f24-2726-405b-897e-9d85c53d8fca",
  "input_asset_ids": [
   "4280069e-a896-4a04-9a87-3534c1d3220b",
   "64f67bba-2b52-476e-9e96-19519e3cc7c9"
  ]
 },
 "S49sh14": {
  "input_fingerprint": "13a557a0633f1dc4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is outside at the fuel pump with the luggage reloaded, and the shop television carries the wanted report. Charlie retains his worn metal body and blanket covering and holds the medicine bottle taken from the medicine section. 조일광: He wears a military uniform with medals and a Marine Corps cap. He holds a shotgun in an aiming position. 앰버: She is inside the mart near the medicine section, still wearing the replacement shoes. 라울: He is inside the mart near the medicine section.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 조일광 right now, so 조일광's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 조일광: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 조일광 (한국인 남성, 50대, 중년의 얼굴, 눈가 주름, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is outside at the fuel pump with the luggage reloaded, and the shop television carries the wanted report. Charlie retains his worn metal body and blanket covering and holds the medicine bottle taken from the medicine section. 조일광: He wears a military uniform with medals and a Marine Corps cap. He holds a shotgun in an aiming position. 앰버: She is inside the mart near the medicine section, still wearing the replacement shoes. 라울: He is inside the mart near the medicine section.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 조일광 right now, so 조일광's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 조일광: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 조일광 (한국인 남성, 50대, 중년의 얼굴, 눈가 주름, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is outside at the fuel pump with the luggage reloaded, and the shop television carries the wanted report. Charlie retains his worn metal body and blanket covering and holds the medicine bottle taken from the medicine section. 조일광: He wears a military uniform with medals and a Marine Corps cap. He holds a shotgun in an aiming position. 앰버: She is inside the mart near the medicine section, still wearing the replacement shoes. 라울: He is inside the mart near the medicine section.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 조일광 right now, so 조일광's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 조일광: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 조일광 (한국인 남성, 50대, 중년의 얼굴, 눈가 주름, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh14__bgfirst_bg.png",
     "asset_id": "837d1f24-2726-405b-897e-9d85c53d8fca",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S49sh14.png",
     "asset_id": "4280069e-a896-4a04-9a87-3534c1d3220b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 조일광: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1310373>",
     "asset_id": "5e8cf758-89bf-4628-a119-36601926db97",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B01.png",
     "asset_id": "64f67bba-2b52-476e-9e96-19519e3cc7c9",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 조일광: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1310373>",
     "asset_id": "5e8cf758-89bf-4628-a119-36601926db97",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 두 아이의 등 뒤에서 통로를 바라본다. 배경의 총잡이는 아이들을 향해 시선을 두고 있으나, 산탄총은 허리춤에 가로로 비스듬히 들고 있으며 총구는 바닥 우측을 향할 뿐 아이들을 겨누지 않고 있다.",
    "built_space": "양옆으로 진열대가 늘어선 편의점 통로이며 뒤편에 냉장고가 보인다. 제시된 레퍼런스 사진의 구체적인 약품 코너와 정확히 일치하는 디테일은 부족하나, 전반적인 편의점 내부 구조로는 성립한다.",
    "entities": "총잡이는 군복과 정모를 착용했으나 얼굴이 조일광 레퍼런스와 다르며, 레퍼런스의 산탄총 대신 일반 펌프액션 산탄총을 들고 있다. 왼쪽 아이(앰버)는 금발이지만 오른쪽 아이(라울)와 완전히 동일한 회색 티셔츠와 반바지를 입고 있어 의상 복제가 일어났다.",
    "hard_violations": [],
    "physics": "세 인물 모두 바닥에 안정적으로 서 있으며, 총잡이는 두 손으로 산탄총의 무게를 지탱하며 들고 있다."
   },
   {
    "label": "B",
    "direction": "카메라는 두 아이의 등 뒤에서 앞을 향한다. 중앙의 총잡이는 무기를 들어 왼쪽 아이(앰버)를 똑바로 겨누고 있다.",
    "built_space": "편의점 통로이며 왼쪽에 외부를 향한 통유리창이 보인다. 레퍼런스 사진의 일반의약품 코너는 매장 안쪽 벽면에 위치하므로, 창가 바로 옆에 구성된 이 공간은 레퍼런스와 어긋난다.",
    "entities": "총잡이의 복장은 군복이나 얼굴 묘사가 레퍼런스와 다르다. 가장 중요한 소품인 산탄총 대신 권총을 쥐고 있어 지문을 위반했다. 왼쪽 아이는 흰색 티셔츠를, 오른쪽 아이는 묶은 머리에 회색 티셔츠를 입고 있다.",
    "hard_violations": [],
    "physics": "인물들은 자연스럽게 서 있고, 총잡이는 한 손으로 권총을 지탱한 채 겨누는 자세를 유지하고 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "지문이 요구한 와이드 샷과 전신 구도, 산탄총이라는 핵심 요소는 포함했으나, 총을 겨누는 동작을 누락하고 두 아이에게 동일한 의상을 복제하여 입힌 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "총을 겨누는 액션은 구현했으나 프레이밍이 전신 샷이 아니며, 지문과 레퍼런스에서 명시한 산탄총 대신 권총을 들고 있는 치명적인 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 두 아이의 등 뒤에서 통로를 바라본다. 배경의 총잡이는 아이들을 향해 시선을 두고 있으나, 산탄총은 허리춤에 가로로 비스듬히 들고 있으며 총구는 바닥 우측을 향할 뿐 아이들을 겨누지 않고 있다.",
        "built_space": "양옆으로 진열대가 늘어선 편의점 통로이며 뒤편에 냉장고가 보인다. 제시된 레퍼런스 사진의 구체적인 약품 코너와 정확히 일치하는 디테일은 부족하나, 전반적인 편의점 내부 구조로는 성립한다.",
        "entities": "총잡이는 군복과 정모를 착용했으나 얼굴이 조일광 레퍼런스와 다르며, 레퍼런스의 산탄총 대신 일반 펌프액션 산탄총을 들고 있다. 왼쪽 아이(앰버)는 금발이지만 오른쪽 아이(라울)와 완전히 동일한 회색 티셔츠와 반바지를 입고 있어 의상 복제가 일어났다.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 안정적으로 서 있으며, 총잡이는 두 손으로 산탄총의 무게를 지탱하며 들고 있다."
       },
       {
        "label": "B",
        "direction": "카메라는 두 아이의 등 뒤에서 앞을 향한다. 중앙의 총잡이는 무기를 들어 왼쪽 아이(앰버)를 똑바로 겨누고 있다.",
        "built_space": "편의점 통로이며 왼쪽에 외부를 향한 통유리창이 보인다. 레퍼런스 사진의 일반의약품 코너는 매장 안쪽 벽면에 위치하므로, 창가 바로 옆에 구성된 이 공간은 레퍼런스와 어긋난다.",
        "entities": "총잡이의 복장은 군복이나 얼굴 묘사가 레퍼런스와 다르다. 가장 중요한 소품인 산탄총 대신 권총을 쥐고 있어 지문을 위반했다. 왼쪽 아이는 흰색 티셔츠를, 오른쪽 아이는 묶은 머리에 회색 티셔츠를 입고 있다.",
        "hard_violations": [],
        "physics": "인물들은 자연스럽게 서 있고, 총잡이는 한 손으로 권총을 지탱한 채 겨누는 자세를 유지하고 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "지문이 요구한 와이드 샷과 전신 구도, 산탄총이라는 핵심 요소는 포함했으나, 총을 겨누는 동작을 누락하고 두 아이에게 동일한 의상을 복제하여 입힌 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "총을 겨누는 액션은 구현했으나 프레이밍이 전신 샷이 아니며, 지문과 레퍼런스에서 명시한 산탄총 대신 권총을 들고 있는 치명적인 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 두 아이의 등 뒤에서 통로를 바라본다. 배경의 총잡이는 아이들을 향해 시선을 두고 있으나, 산탄총은 허리춤에 가로로 비스듬히 들고 있으며 총구는 바닥 우측을 향할 뿐 아이들을 겨누지 않고 있다.",
        "built_space": "양옆으로 진열대가 늘어선 편의점 통로이며 뒤편에 냉장고가 보인다. 제시된 레퍼런스 사진의 구체적인 약품 코너와 정확히 일치하는 디테일은 부족하나, 전반적인 편의점 내부 구조로는 성립한다.",
        "entities": "총잡이는 군복과 정모를 착용했으나 얼굴이 조일광 레퍼런스와 다르며, 레퍼런스의 산탄총 대신 일반 펌프액션 산탄총을 들고 있다. 왼쪽 아이(앰버)는 금발이지만 오른쪽 아이(라울)와 완전히 동일한 회색 티셔츠와 반바지를 입고 있어 의상 복제가 일어났다.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 안정적으로 서 있으며, 총잡이는 두 손으로 산탄총의 무게를 지탱하며 들고 있다."
       },
       {
        "label": "B",
        "direction": "카메라는 두 아이의 등 뒤에서 앞을 향한다. 중앙의 총잡이는 무기를 들어 왼쪽 아이(앰버)를 똑바로 겨누고 있다.",
        "built_space": "편의점 통로이며 왼쪽에 외부를 향한 통유리창이 보인다. 레퍼런스 사진의 일반의약품 코너는 매장 안쪽 벽면에 위치하므로, 창가 바로 옆에 구성된 이 공간은 레퍼런스와 어긋난다.",
        "entities": "총잡이의 복장은 군복이나 얼굴 묘사가 레퍼런스와 다르다. 가장 중요한 소품인 산탄총 대신 권총을 쥐고 있어 지문을 위반했다. 왼쪽 아이는 흰색 티셔츠를, 오른쪽 아이는 묶은 머리에 회색 티셔츠를 입고 있다.",
        "hard_violations": [],
        "physics": "인물들은 자연스럽게 서 있고, 총잡이는 한 손으로 권총을 지탱한 채 겨누는 자세를 유지하고 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "아이들의 전경 어깨는 구현했지만 조일광의 발을 잘랐고, 총구가 아이보다 카메라 쪽 빈틈을 향하며 장소와 앰버의 의상도 다릅니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "조일광의 전신과 총의 측면을 보여 우세하지만, 총을 들어 아이들을 겨누는 순간 대신 낮춰 들었으며 장소와 앰버의 의상이 맞지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "아이들은 등을 카메라에 보이고 조일광을 바라본다. 조일광의 시선은 전경 아이들 쪽이지만 총구는 두 아이 사이의 빈 공간, 카메라에 가까운 방향을 향한다. 어느 아이의 몸을 겨누는지 명확하지 않고, 총의 측면보다 총구 정면과 단축된 총신이 두드러진다.",
        "built_space": "통로 양쪽에 상품 진열대가 하나씩 있고, 뒤에는 유리문 수납부와 벽면 약품 선반이 있다. 왼쪽 유리창과 천장 형광등은 낮의 편의점이라는 조건에 맞는다. 그러나 참고 장소의 중앙 독립 진열대 옆 열린 통로가 대칭적인 약품 통로로 바뀌었고 바닥도 다른 타일이다. 참고의 오른쪽 책상, 문, 벽걸이 텔레비전, 볼록거울은 보이지 않는다. 조일광은 통로 중앙에 서지만 다리 아래와 발이 잘려 전신 공개가 아니다.",
        "entities": "성인 남성 한 명과 아이 두 명만 보인다. 조일광은 중년 동아시아계 남성으로 보이며 훈장 달린 군복을 입었지만, 모자는 해병대 모자보다는 정복용 정모로 읽힌다. 앰버는 금발 여자아이지만 참고의 풀어 내린 머리, 고글, 남색 상의와 갈색 작업복 대신 묶은 머리와 밝은 티셔츠를 착용했다. 라울의 어두운 피부, 뒤로 묶은 검은 머리와 회색 티셔츠는 대체로 맞으며 두 아이의 얼굴 정체성은 뒷모습이라 확인하기 어렵다. 총은 참고처럼 짧은 금속제 총에 가깝지만 개머리판 형태는 확인되지 않는다. 상품 포장에 문자 모양이 많이 남아 있으나 특정 문구의 판독 여부는 확정하기 어렵다.",
        "hard_violations": [],
        "physics": "조일광은 오른손으로 총의 손잡이를 잡고 팔을 뻗고 있어 총 자체는 손으로 지지된다. 다른 팔은 아래로 내려가 있으며 어깨에 개머리판을 대는 조준 자세는 아니다. 발은 화면 밖이지만 몸이 공중에 뜬 흔적은 없고 서 있는 하체로 이어진다. 아이들의 하체도 화면 밖이며 부유나 불가능한 접촉은 보이지 않는다. 상품은 선반 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "두 아이는 조일광 쪽을 바라보고 조일광도 아이들 쪽으로 시선을 둔다. 총은 카메라 정면이 아니라 화면 오른쪽 아래로 향하며, 총신의 연장선은 라울의 왼팔 아래쪽과 골반 부근으로 이어진다. 다만 총을 가슴이나 어깨 높이로 올려 겨누지 않고 허리에서 비스듬히 낮춰 들었으므로 지시된 조준 순간과 다르다.",
        "built_space": "상품이 그대로 놓인 긴 진열대 두 줄이 좌우에서 통로를 감싸고, 뒤쪽에는 여러 유리문으로 나뉜 냉장 진열장이 있다. 천장에는 긴 조명 줄과 환기구가, 뒤 벽에는 녹색 띠와 시계 하나가 보인다. 세 사람은 통로 바닥에 서 있으며 조일광의 모자부터 신발까지 모두 들어온다. 그러나 참고의 창가, 중앙 독립 진열대, 후면 약품장과 오른쪽 업무 공간으로 구성된 장소와는 다르다. 아이들까지 전신으로 보여 지시된 전경 어깨 중심의 위협 연결은 약해졌다.",
        "entities": "성인 남성 한 명과 아이 두 명만 있다. 조일광은 중년 동아시아계 남성으로 보이고 훈장이 달린 군복을 입었으나 정모의 해병대 정체성은 불명확하다. 앰버는 금발 여자아이지만 참고의 고글과 갈색 작업복 대신 라울과 비슷한 회색 티셔츠와 반바지를 입고 있다. 라울의 묶은 검은 머리, 피부색, 회색 티셔츠와 반바지는 참고와 대체로 맞는다. 아이들은 뒤를 돌아 얼굴의 세부 일치 여부를 확인할 수 없다. 총은 긴 산탄총 형태로 읽히지만 참고의 짧은 권총형 몸체와 신축식 개머리판 형태와는 다르다. 텔레비전의 수배 보도는 보이지 않으며 상품의 작은 글자는 대체로 판독하기 어렵다.",
        "hard_violations": [],
        "physics": "세 사람 모두 두 발이 바닥에 닿아 체중을 지탱한다. 조일광의 오른손은 총 손잡이를 잡고 왼손은 총의 앞부분을 받쳐 총의 무게를 지지한다. 낮춰 든 자세 자체는 물리적으로 가능하지만 올려 조준하는 동작은 아니다. 아이들의 내려놓은 팔과 벌린 발도 가능한 자세이며, 선반 상품을 포함해 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "아이들의 전경 어깨는 구현했지만 조일광의 발을 잘랐고, 총구가 아이보다 카메라 쪽 빈틈을 향하며 장소와 앰버의 의상도 다릅니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "조일광의 전신과 총의 측면을 보여 우세하지만, 총을 들어 아이들을 겨누는 순간 대신 낮춰 들었으며 장소와 앰버의 의상이 맞지 않습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "아이들은 등을 카메라에 보이고 조일광을 바라본다. 조일광의 시선은 전경 아이들 쪽이지만 총구는 두 아이 사이의 빈 공간, 카메라에 가까운 방향을 향한다. 어느 아이의 몸을 겨누는지 명확하지 않고, 총의 측면보다 총구 정면과 단축된 총신이 두드러진다.",
        "built_space": "통로 양쪽에 상품 진열대가 하나씩 있고, 뒤에는 유리문 수납부와 벽면 약품 선반이 있다. 왼쪽 유리창과 천장 형광등은 낮의 편의점이라는 조건에 맞는다. 그러나 참고 장소의 중앙 독립 진열대 옆 열린 통로가 대칭적인 약품 통로로 바뀌었고 바닥도 다른 타일이다. 참고의 오른쪽 책상, 문, 벽걸이 텔레비전, 볼록거울은 보이지 않는다. 조일광은 통로 중앙에 서지만 다리 아래와 발이 잘려 전신 공개가 아니다.",
        "entities": "성인 남성 한 명과 아이 두 명만 보인다. 조일광은 중년 동아시아계 남성으로 보이며 훈장 달린 군복을 입었지만, 모자는 해병대 모자보다는 정복용 정모로 읽힌다. 앰버는 금발 여자아이지만 참고의 풀어 내린 머리, 고글, 남색 상의와 갈색 작업복 대신 묶은 머리와 밝은 티셔츠를 착용했다. 라울의 어두운 피부, 뒤로 묶은 검은 머리와 회색 티셔츠는 대체로 맞으며 두 아이의 얼굴 정체성은 뒷모습이라 확인하기 어렵다. 총은 참고처럼 짧은 금속제 총에 가깝지만 개머리판 형태는 확인되지 않는다. 상품 포장에 문자 모양이 많이 남아 있으나 특정 문구의 판독 여부는 확정하기 어렵다.",
        "hard_violations": [],
        "physics": "조일광은 오른손으로 총의 손잡이를 잡고 팔을 뻗고 있어 총 자체는 손으로 지지된다. 다른 팔은 아래로 내려가 있으며 어깨에 개머리판을 대는 조준 자세는 아니다. 발은 화면 밖이지만 몸이 공중에 뜬 흔적은 없고 서 있는 하체로 이어진다. 아이들의 하체도 화면 밖이며 부유나 불가능한 접촉은 보이지 않는다. 상품은 선반 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "두 아이는 조일광 쪽을 바라보고 조일광도 아이들 쪽으로 시선을 둔다. 총은 카메라 정면이 아니라 화면 오른쪽 아래로 향하며, 총신의 연장선은 라울의 왼팔 아래쪽과 골반 부근으로 이어진다. 다만 총을 가슴이나 어깨 높이로 올려 겨누지 않고 허리에서 비스듬히 낮춰 들었으므로 지시된 조준 순간과 다르다.",
        "built_space": "상품이 그대로 놓인 긴 진열대 두 줄이 좌우에서 통로를 감싸고, 뒤쪽에는 여러 유리문으로 나뉜 냉장 진열장이 있다. 천장에는 긴 조명 줄과 환기구가, 뒤 벽에는 녹색 띠와 시계 하나가 보인다. 세 사람은 통로 바닥에 서 있으며 조일광의 모자부터 신발까지 모두 들어온다. 그러나 참고의 창가, 중앙 독립 진열대, 후면 약품장과 오른쪽 업무 공간으로 구성된 장소와는 다르다. 아이들까지 전신으로 보여 지시된 전경 어깨 중심의 위협 연결은 약해졌다.",
        "entities": "성인 남성 한 명과 아이 두 명만 있다. 조일광은 중년 동아시아계 남성으로 보이고 훈장이 달린 군복을 입었으나 정모의 해병대 정체성은 불명확하다. 앰버는 금발 여자아이지만 참고의 고글과 갈색 작업복 대신 라울과 비슷한 회색 티셔츠와 반바지를 입고 있다. 라울의 묶은 검은 머리, 피부색, 회색 티셔츠와 반바지는 참고와 대체로 맞는다. 아이들은 뒤를 돌아 얼굴의 세부 일치 여부를 확인할 수 없다. 총은 긴 산탄총 형태로 읽히지만 참고의 짧은 권총형 몸체와 신축식 개머리판 형태와는 다르다. 텔레비전의 수배 보도는 보이지 않으며 상품의 작은 글자는 대체로 판독하기 어렵다.",
        "hard_violations": [],
        "physics": "세 사람 모두 두 발이 바닥에 닿아 체중을 지탱한다. 조일광의 오른손은 총 손잡이를 잡고 왼손은 총의 앞부분을 받쳐 총의 무게를 지지한다. 낮춰 든 자세 자체는 물리적으로 가능하지만 올려 조준하는 동작은 아니다. 아이들의 내려놓은 팔과 벌린 발도 가능한 자세이며, 선반 상품을 포함해 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.35
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.35
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1350
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지문이 요구한 와이드 샷과 전신 구도, 산탄총이라는 핵심 요소는 포함했으나, 총을 겨누는 동작을 누락하고 두 아이에게 동일한 의상을 복제하여 입힌 점이 감점 요인입니다."
   },
   {
    "label": "B",
    "score": 1350,
    "verdict_ko": "총을 겨누는 액션은 구현했으나 프레이밍이 전신 샷이 아니며, 지문과 레퍼런스에서 명시한 산탄총 대신 권총을 들고 있는 치명적인 오류가 있습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B01.png",
    "asset_id": "64f67bba-2b52-476e-9e96-19519e3cc7c9",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 조일광: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1310373>",
    "asset_id": "5e8cf758-89bf-4628-a119-36601926db97",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-e1d5-7f56-b52c-538d7e5af8c2",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh14__bgfirst_bg.png",
   "bg_asset_id": "837d1f24-2726-405b-897e-9d85c53d8fca",
   "bg_record_key": "S49sh14::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S49sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:15:20.387340+00:00",
  "fingerprint": "8b8e92405ff8849c2315c2f9bc59c7ca1512295265876bfb4e33cbdf83e85f6b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S49sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S49sh14_sel.png",
  "source_sha256": "c4a0f8a58faee33c920dc690ce2159dc43ce7537fc98595d511edf392184cd61",
  "file": "S49sh14_cine.png",
  "staged_sha256": "9da5ac3c58f300dde8954c4523f0ab54ec9b86eb3cecb0725954b7e3cc22e148",
  "latency_ms": 10737
 },
 "S49sh47::signage": {
  "fp": "43280a2465ad2a5b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S49sh47": {
  "input_fingerprint": "4d7c63f44b030f4c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기 둔덕과 강하게 충돌한 반동으로 공중으로 솟구쳐 거꾸로 뒤집히기 시작한 캠핑카의 폭발적인 순간.\n\nLOCATION (lock): At a rubbish embankment beside the road beyond the service station, where the pursued camper crashes and overturns. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Trash mound beneath the rebounding camper in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Camper (Airborne after impact and beginning to overturn) — Rear and passenger-side surfaces rotate away from their normal upright relationship to the road; used as Its rotation against the level horizon carries the impact; Trash mound (Struck by the camper); used as Remains below the airborne vehicle as the visible cause of the rebound; Road (Beside the collision site) — The road recedes diagonally while its horizon remains level; used as Provides a stable spatial reference for the overturning vehicle.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled contrast makes the airborne body and ground contact readable without adding an explosion or impact glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper strikes the rubbish embankment with a shotgun-shattered window and a separately smashed passenger-side window; its interior bulbs have flared brightly. Charlie has a bullet-grazed, sparking shoulder, and the pursuing hunting drone has exploded.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기 둔덕과 강하게 충돌한 반동으로 공중으로 솟구쳐 거꾸로 뒤집히기 시작한 캠핑카의 폭발적인 순간.\n\nLOCATION (lock): At a rubbish embankment beside the road beyond the service station, where the pursued camper crashes and overturns. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Trash mound beneath the rebounding camper in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Camper (Airborne after impact and beginning to overturn) — Rear and passenger-side surfaces rotate away from their normal upright relationship to the road; used as Its rotation against the level horizon carries the impact; Trash mound (Struck by the camper); used as Remains below the airborne vehicle as the visible cause of the rebound; Road (Beside the collision site) — The road recedes diagonally while its horizon remains level; used as Provides a stable spatial reference for the overturning vehicle.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled contrast makes the airborne body and ground contact readable without adding an explosion or impact glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper strikes the rubbish embankment with a shotgun-shattered window and a separately smashed passenger-side window; its interior bulbs have flared brightly. Charlie has a bullet-grazed, sparking shoulder, and the pursuing hunting drone has exploded.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기 둔덕과 강하게 충돌한 반동으로 공중으로 솟구쳐 거꾸로 뒤집히기 시작한 캠핑카의 폭발적인 순간.\n\nLOCATION (lock): At a rubbish embankment beside the road beyond the service station, where the pursued camper crashes and overturns. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Trash mound beneath the rebounding camper in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Camper (Airborne after impact and beginning to overturn) — Rear and passenger-side surfaces rotate away from their normal upright relationship to the road; used as Its rotation against the level horizon carries the impact; Trash mound (Struck by the camper); used as Remains below the airborne vehicle as the visible cause of the rebound; Road (Beside the collision site) — The road recedes diagonally while its horizon remains level; used as Provides a stable spatial reference for the overturning vehicle.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled contrast makes the airborne body and ground contact readable without adding an explosion or impact glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper strikes the rubbish embankment with a shotgun-shattered window and a separately smashed passenger-side window; its interior bulbs have flared brightly. Charlie has a bullet-grazed, sparking shoulder, and the pursuing hunting drone has exploded.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "캠핑카가 쓰레기 더미 위로 앞부분이 들린 채 전진함.",
    "built_space": "위치 사진과 유사한 주유소 캐노피와 주유기가 배경에 배치됨.",
    "entities": "캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 파손된 창문.",
    "hard_violations": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gemini-pro] 읽을 수 있는 텍스트('S-OIL') 포함",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
     "[gpt-high] 캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
    ],
    "physics": "앞바퀴가 들려 도약하는 형태이나, 텍스트가 요구한 '거꾸로 뒤집히는' 회전이나 물리적 반동은 나타나지 않음."
   },
   {
    "label": "B",
    "direction": "캠핑카가 공중에서 심하게 앞으로 쏠리며 측면으로 회전 중임.",
    "built_space": "배경에 주유소 캐노피와 주유기가 위치하며 텍스트가 생략됨.",
    "entities": "전복 중인 캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 깨진 창문과 내부 불빛.",
    "hard_violations": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
    ],
    "physics": "쓰레기 더미와 충돌한 반동으로 완전히 허공에 뜬 채 거꾸로 뒤집히기 시작하는 궤적을 제대로 받쳐줌."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캠핑카가 전복되지 않고 도약만 하여 핵심 액션에 실패했으며, 금지된 인물과 텍스트가 모두 포함되어 위반이 심함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캠핑카가 공중에서 뒤집히는 역동적인 순간은 정확히 구현했으나, 배경에 금지된 인물이 그대로 복사되어 하드 위반임."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 쓰레기 더미 위로 앞부분이 들린 채 전진함.",
        "built_space": "위치 사진과 유사한 주유소 캐노피와 주유기가 배경에 배치됨.",
        "entities": "캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 파손된 창문.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장",
         "읽을 수 있는 텍스트('S-OIL') 포함"
        ],
        "physics": "앞바퀴가 들려 도약하는 형태이나, 텍스트가 요구한 '거꾸로 뒤집히는' 회전이나 물리적 반동은 나타나지 않음."
       },
       {
        "label": "B",
        "direction": "캠핑카가 공중에서 심하게 앞으로 쏠리며 측면으로 회전 중임.",
        "built_space": "배경에 주유소 캐노피와 주유기가 위치하며 텍스트가 생략됨.",
        "entities": "전복 중인 캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 깨진 창문과 내부 불빛.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장"
        ],
        "physics": "쓰레기 더미와 충돌한 반동으로 완전히 허공에 뜬 채 거꾸로 뒤집히기 시작하는 궤적을 제대로 받쳐줌."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캠핑카가 전복되지 않고 도약만 하여 핵심 액션에 실패했으며, 금지된 인물과 텍스트가 모두 포함되어 위반이 심함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캠핑카가 공중에서 뒤집히는 역동적인 순간은 정확히 구현했으나, 배경에 금지된 인물이 그대로 복사되어 하드 위반임."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 쓰레기 더미 위로 앞부분이 들린 채 전진함.",
        "built_space": "위치 사진과 유사한 주유소 캐노피와 주유기가 배경에 배치됨.",
        "entities": "캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 파손된 창문.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장",
         "읽을 수 있는 텍스트('S-OIL') 포함"
        ],
        "physics": "앞바퀴가 들려 도약하는 형태이나, 텍스트가 요구한 '거꾸로 뒤집히는' 회전이나 물리적 반동은 나타나지 않음."
       },
       {
        "label": "B",
        "direction": "캠핑카가 공중에서 심하게 앞으로 쏠리며 측면으로 회전 중임.",
        "built_space": "배경에 주유소 캐노피와 주유기가 위치하며 텍스트가 생략됨.",
        "entities": "전복 중인 캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 깨진 창문과 내부 불빛.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장"
        ],
        "physics": "쓰레기 더미와 충돌한 반동으로 완전히 허공에 뜬 채 거꾸로 뒤집히기 시작하는 궤적을 제대로 받쳐줌."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "쓰레기 둔덕 위로 튀어 올라 크게 뒤집히는 차량의 후면과 하부는 요청한 순간에 더 가깝지만, 배경 인물 등장과 구조 참조 불일치, 과도한 충돌 섬광 때문에 사용할 수 없다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "뒤집히기보다 둔덕을 타고 앞머리가 들리는 모습이며, 배경 인물과 읽히는 주유소 로고까지 보여 핵심 동작과 명시적 금지 조건을 모두 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 후면이 카메라를 향하고 차체가 화면 오른쪽으로 크게 기울어 하부가 드러난다. 아래의 쓰레기 둔덕에서 위로 튕겨 나온 뒤 옆으로 전복되기 시작하는 방향이 읽힌다. 조수석 쪽 면은 좁게 보여 방향 확인이 제한적이다. 도로 경계는 대각선으로 이어지고 배경의 수평 기준은 유지된다. 인물의 시선은 식별되지 않는다.",
        "built_space": "왼쪽에 대형 캐노피 하나, 사각 지지 기둥 세 개와 주유기 세 대가 보이고, 뒤에는 파란 띠와 유리 전면이 있는 낮은 매장이 있다. 오른쪽에는 전신주 하나, 큰 금속 쓰레기통 하나, 낮은 상자 하나와 쓰레기 둔덕이 있다. 둔덕은 차량 바로 아래 오른쪽에 있으나 차량은 화면 상단에 거의 닿을 만큼 커서 중경 와이드 배치보다 가까워 보인다. 주유소는 노란 지붕 테두리만 구조 참조에 가까우며, 참조의 원형 금속 기둥·노란 주유 설비·현대적인 매장 전면을 재현하지 않았다.",
        "entities": "캠핑카 한 대, 쓰레기 둔덕, 도로, 주유소가 있다. 후면 유리에는 큰 파손 구멍이 있고 왼쪽으로 드러난 다른 창도 깨져 있지만, 별도로 파손된 조수석 창인지는 확정하기 어렵다. 내부의 밝은 전구는 보인다. 차량 바깥 왼쪽에도 강한 섬광이 있어 충돌 발광 금지와 맞지 않는다. 중앙 기둥 오른쪽 매장 앞에 검은 옷을 입은 사람이 한 명 보인다. 성별·연령·민족성은 판별할 수 없다. 찰리나 드론은 식별되지 않으며, 뚜렷하게 읽히는 문구는 보이지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
        ],
        "physics": "차량 바퀴는 지면에서 떨어져 있지만 바로 아래 둔덕에서 솟는 먼지와 파편, 둔덕 가까이에 남은 차체 하단이 충돌에 의한 이륙 원인을 보여 준다. 기울어진 차체는 반동과 회전으로 설명 가능하며, 근처 도로와 둔덕이 낙하할 지면이다. 공중 파편도 같은 충격에서 튀어나온 것으로 읽힌다. 근거 없이 정지해 떠 있는 차량은 아니다."
       },
       {
        "label": "B",
        "direction": "캠핑카 앞머리는 화면 오른쪽 쓰레기 둔덕을 향하고 전면과 조수석 쪽 측면이 카메라에 보인다. 앞바퀴가 올라가고 뒤쪽이 낮아 둔덕을 타고 오르는 방향은 분명하지만, 후면을 보이며 거꾸로 뒤집히기 시작하는 회전은 약하다. 도로는 왼쪽으로 대각선 후퇴하며 배경은 수평을 유지한다. 배경 인물의 시선은 확인되지 않는다.",
        "built_space": "대형 캐노피 하나 아래 사각 기둥 세 개와 주유기 세 대가 보인다. 뒤에는 파란 띠의 낮은 매장과 회색 상부 매스가 있다. 오른쪽에는 줄무늬 전신주 하나, 금속 쓰레기통 하나, 낮은 상자 하나가 있고 둔덕은 캠핑카 아래 오른쪽에 놓인다. 차량 전체를 담은 와이드 구도이지만 위치 참조의 시점을 상당히 그대로 따른다. 구조 참조의 원형 금속 기둥, 노란 주유기와 보조 차양, 유리 매장 구성은 재현되지 않았다.",
        "entities": "캠핑카 한 대와 도로 옆 쓰레기 둔덕이 있다. 측면 뒤 창에 파손 구멍이 있고 앞 유리와 조수석 창에도 균열이 보이며, 여러 창 너머로 내부 조명이 밝게 빛난다. 산탄총으로 깨진 창과 별도로 부서진 조수석 창이라는 손상 구분은 충분히 선명하지 않다. 주유소 중앙 기둥 오른쪽에 검은 옷을 입은 사람이 한 명 있다. 작아서 성별·연령·민족성은 확정할 수 없다. 캐노피에는 주유소 영문 상호와 로고가 선명하게 읽힌다. 찰리와 드론은 식별되지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
         "캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
        ],
        "physics": "앞바퀴는 둔덕 위 공중에 있고 뒷바퀴는 노면에 매우 가깝다. 앞부분 아래의 쓰레기와 충돌 먼지, 튀는 파편이 차량을 들어 올린 원인을 제공한다. 따라서 무근거한 부유는 아니지만, 차체가 아직 대체로 바로 서 있어 충돌 반동으로 전복되는 순간보다는 둔덕을 타고 들리는 단계로 보인다. 파편은 충돌 지점에서 튀어나오며 지면으로 떨어질 수 있는 배치다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "쓰레기 둔덕 위로 튀어 올라 크게 뒤집히는 차량의 후면과 하부는 요청한 순간에 더 가깝지만, 배경 인물 등장과 구조 참조 불일치, 과도한 충돌 섬광 때문에 사용할 수 없다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "뒤집히기보다 둔덕을 타고 앞머리가 들리는 모습이며, 배경 인물과 읽히는 주유소 로고까지 보여 핵심 동작과 명시적 금지 조건을 모두 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 후면이 카메라를 향하고 차체가 화면 오른쪽으로 크게 기울어 하부가 드러난다. 아래의 쓰레기 둔덕에서 위로 튕겨 나온 뒤 옆으로 전복되기 시작하는 방향이 읽힌다. 조수석 쪽 면은 좁게 보여 방향 확인이 제한적이다. 도로 경계는 대각선으로 이어지고 배경의 수평 기준은 유지된다. 인물의 시선은 식별되지 않는다.",
        "built_space": "왼쪽에 대형 캐노피 하나, 사각 지지 기둥 세 개와 주유기 세 대가 보이고, 뒤에는 파란 띠와 유리 전면이 있는 낮은 매장이 있다. 오른쪽에는 전신주 하나, 큰 금속 쓰레기통 하나, 낮은 상자 하나와 쓰레기 둔덕이 있다. 둔덕은 차량 바로 아래 오른쪽에 있으나 차량은 화면 상단에 거의 닿을 만큼 커서 중경 와이드 배치보다 가까워 보인다. 주유소는 노란 지붕 테두리만 구조 참조에 가까우며, 참조의 원형 금속 기둥·노란 주유 설비·현대적인 매장 전면을 재현하지 않았다.",
        "entities": "캠핑카 한 대, 쓰레기 둔덕, 도로, 주유소가 있다. 후면 유리에는 큰 파손 구멍이 있고 왼쪽으로 드러난 다른 창도 깨져 있지만, 별도로 파손된 조수석 창인지는 확정하기 어렵다. 내부의 밝은 전구는 보인다. 차량 바깥 왼쪽에도 강한 섬광이 있어 충돌 발광 금지와 맞지 않는다. 중앙 기둥 오른쪽 매장 앞에 검은 옷을 입은 사람이 한 명 보인다. 성별·연령·민족성은 판별할 수 없다. 찰리나 드론은 식별되지 않으며, 뚜렷하게 읽히는 문구는 보이지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
        ],
        "physics": "차량 바퀴는 지면에서 떨어져 있지만 바로 아래 둔덕에서 솟는 먼지와 파편, 둔덕 가까이에 남은 차체 하단이 충돌에 의한 이륙 원인을 보여 준다. 기울어진 차체는 반동과 회전으로 설명 가능하며, 근처 도로와 둔덕이 낙하할 지면이다. 공중 파편도 같은 충격에서 튀어나온 것으로 읽힌다. 근거 없이 정지해 떠 있는 차량은 아니다."
       },
       {
        "label": "A",
        "direction": "캠핑카 앞머리는 화면 오른쪽 쓰레기 둔덕을 향하고 전면과 조수석 쪽 측면이 카메라에 보인다. 앞바퀴가 올라가고 뒤쪽이 낮아 둔덕을 타고 오르는 방향은 분명하지만, 후면을 보이며 거꾸로 뒤집히기 시작하는 회전은 약하다. 도로는 왼쪽으로 대각선 후퇴하며 배경은 수평을 유지한다. 배경 인물의 시선은 확인되지 않는다.",
        "built_space": "대형 캐노피 하나 아래 사각 기둥 세 개와 주유기 세 대가 보인다. 뒤에는 파란 띠의 낮은 매장과 회색 상부 매스가 있다. 오른쪽에는 줄무늬 전신주 하나, 금속 쓰레기통 하나, 낮은 상자 하나가 있고 둔덕은 캠핑카 아래 오른쪽에 놓인다. 차량 전체를 담은 와이드 구도이지만 위치 참조의 시점을 상당히 그대로 따른다. 구조 참조의 원형 금속 기둥, 노란 주유기와 보조 차양, 유리 매장 구성은 재현되지 않았다.",
        "entities": "캠핑카 한 대와 도로 옆 쓰레기 둔덕이 있다. 측면 뒤 창에 파손 구멍이 있고 앞 유리와 조수석 창에도 균열이 보이며, 여러 창 너머로 내부 조명이 밝게 빛난다. 산탄총으로 깨진 창과 별도로 부서진 조수석 창이라는 손상 구분은 충분히 선명하지 않다. 주유소 중앙 기둥 오른쪽에 검은 옷을 입은 사람이 한 명 있다. 작아서 성별·연령·민족성은 확정할 수 없다. 캐노피에는 주유소 영문 상호와 로고가 선명하게 읽힌다. 찰리와 드론은 식별되지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
         "캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
        ],
        "physics": "앞바퀴는 둔덕 위 공중에 있고 뒷바퀴는 노면에 매우 가깝다. 앞부분 아래의 쓰레기와 충돌 먼지, 튀는 파편이 차량을 들어 올린 원인을 제공한다. 따라서 무근거한 부유는 아니지만, 차체가 아직 대체로 바로 서 있어 충돌 반동으로 전복되는 순간보다는 둔덕을 타고 들리는 단계로 보인다. 파편은 충돌 지점에서 튀어나오며 지면으로 떨어질 수 있는 배치다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.083,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.833,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gemini-pro] 읽을 수 있는 텍스트('S-OIL') 포함",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
     "[gpt-high] 캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 833,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 833,
    "verdict_ko": "캠핑카가 전복되지 않고 도약만 하여 핵심 액션에 실패했으며, 금지된 인물과 텍스트가 모두 포함되어 위반이 심함.  ★위반: [gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장 / [gemini-pro] 읽을 수 있는 텍스트('S-OIL') 포함 / [gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다. / [gpt-high] 캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "캠핑카가 공중에서 뒤집히는 역동적인 순간은 정확히 구현했으나, 배경에 금지된 인물이 그대로 복사되어 하드 위반임.  ★위반: [gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장 / [gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B02.png",
    "asset_id": "d012fbdf-5de4-474f-8f94-751d1a131761",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_roadside_service_station_sel.png",
    "asset_id": "ab322096-550b-4b37-8cda-d1efdc3c511b",
    "role": "structure_seed_look"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-e541-757e-a39e-a3c2e96e850a",
  "ref_mode": "플레이트+seed만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_bypass:bg_only"
 },
 "S49sh47::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:56:46.705788+00:00",
  "fingerprint": "d30f7a91812bc58c672ea89fbe54598fc630ca77d0415cc124648601b9a251c2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S49sh47_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S49sh47_sel.png",
  "source_sha256": "6f9a7ba0b7af615ffc41fc50bfb71a1699a15629c4c7fc4e6acf537190644102",
  "file": "S49sh47_cine.png",
  "staged_sha256": "64123d32276da34aad45604efcfdbb4eccda811ad9118321de3b71e5e4f0bc79",
  "latency_ms": 13632
 },
 "S49sh51::signage": {
  "fp": "532422eeffb2d15a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S49sh51::bgfirst_bg": {
  "input_fingerprint": "d1c5362b7fa4f95d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh51__bgfirst_bg.png",
  "asset_id": "c25c9337-c13b-46af-8135-de9acaf2fcf0",
  "input_asset_ids": [
   "f4934d4b-9116-40dd-a1b2-2a834fade0e4",
   "d012fbdf-5de4-474f-8f94-751d1a131761",
   "ab322096-550b-4b37-8cda-d1efdc3c511b"
  ]
 },
 "S49sh51": {
  "input_fingerprint": "4c9c8a10b6af19da",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper lies completely overturned with shattered windows, smoke inside and an interior bulb flickering faintly. Charlie is unconscious inside with the damaged shoulder sustained during the shooting. 태진: She stands outside the wreck looking into it and has short bobbed hair. 은영: She stands outside the wreck looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발); 은영 (한국인 여성, 31세, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper lies completely overturned with shattered windows, smoke inside and an interior bulb flickering faintly. Charlie is unconscious inside with the damaged shoulder sustained during the shooting. 태진: She stands outside the wreck looking into it and has short bobbed hair. 은영: She stands outside the wreck looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발); 은영 (한국인 여성, 31세, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper lies completely overturned with shattered windows, smoke inside and an interior bulb flickering faintly. Charlie is unconscious inside with the damaged shoulder sustained during the shooting. 태진: She stands outside the wreck looking into it and has short bobbed hair. 은영: She stands outside the wreck looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발); 은영 (한국인 여성, 31세, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh51__bgfirst_bg.png",
     "asset_id": "c25c9337-c13b-46af-8135-de9acaf2fcf0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S49sh51.png",
     "asset_id": "f4934d4b-9116-40dd-a1b2-2a834fade0e4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:922276>",
     "asset_id": "d985233d-df72-452a-9cf8-93dc78e15a29",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 은영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1258866>",
     "asset_id": "d98e5c15-3084-4367-9745-556e40a4a1c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B02.png",
     "asset_id": "d012fbdf-5de4-474f-8f94-751d1a131761",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_roadside_service_station_sel.png",
     "asset_id": "ab322096-550b-4b37-8cda-d1efdc3c511b",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:922276>",
     "asset_id": "d985233d-df72-452a-9cf8-93dc78e15a29",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 은영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1258866>",
     "asset_id": "d98e5c15-3084-4367-9745-556e40a4a1c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 인물의 시선이 깨진 창문을 통해 차량 내부(카메라 방향)를 향해 아래로 내려다보고 있음.",
    "built_space": "차량 내부에서 부서진 창틀 너머로 밖을 내다보는 시점. 프레임 가장자리에 차량 파편이 위치하며 배경에 주유소와 쓰레기 더미가 적절한 거리감으로 배치됨.",
    "entities": "은영(왼쪽, 얼굴은 일치하나 머리를 뒤로 묶음), 태진(오른쪽, 단발머리와 작업복 일치). 전경에 의식을 잃은 찰리로 추정되는 인물의 일부가 보임.",
    "hard_violations": [
     "[gpt-high] 하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 샷에 보이는 신체는 은영과 태진에게만 속해야 한다는 명시적 제한을 위반한다."
    ],
    "physics": "두 인물 모두 차량 밖 지면에 안정적으로 서서 안을 들여다보는 자세를 취하고 있음."
   },
   {
    "label": "B",
    "direction": "두 인물의 시선이 깨진 창문을 통해 캠퍼 내부 정면을 향하고 있음.",
    "built_space": "캠퍼 내부에서 밖을 내다보는 시점이나, 내부의 싱크대와 상단 수납장이 뒤집히지 않고 정상적인 위아래 방향을 유지하고 있음.",
    "entities": "은영(왼쪽, 단발머리와 얼굴 일치), 태진(오른쪽, 단발머리와 작업복 일치).",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시된 '거꾸로 뒤집힌 차(Inverted after the crash)' 지시를 어기고 캠퍼 내부(싱크대, 수납장 등)를 정상적인 수직 방향으로 렌더링함.",
     "[gemini-pro] 배경 건물 간판에 읽을 수 있는 텍스트(S-OIL mart)가 노출됨.",
     "[gpt-high] 주유소 캐노피와 상점에 읽을 수 있는 영문 상호·로고가 노출되어 문자와 로고 금지 조건을 위반한다.",
     "[gpt-high] 싱크대와 상부 수납장 등이 정상 상하관계로 놓여 있어, 완전히 전복된 캠퍼 내부라는 필수 공간 상태를 구현하지 않았다."
    ],
    "physics": "두 인물 모두 지면에 안정적으로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 미디엄 샷 프레이밍을 잘 준수하고 뒤집힌 차량 내부의 시점과 전경의 인물을 효과적으로 연출하였으나, 은영의 헤어스타일이 레퍼런스와 다소 다름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물들의 외형은 레퍼런스와 잘 일치하나, 거꾸로 뒤집혀야 할 차량 내부가 똑바로 서 있는 상태로 묘사되는 치명적인 구조적 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물의 시선이 깨진 창문을 통해 차량 내부(카메라 방향)를 향해 아래로 내려다보고 있음.",
        "built_space": "차량 내부에서 부서진 창틀 너머로 밖을 내다보는 시점. 프레임 가장자리에 차량 파편이 위치하며 배경에 주유소와 쓰레기 더미가 적절한 거리감으로 배치됨.",
        "entities": "은영(왼쪽, 얼굴은 일치하나 머리를 뒤로 묶음), 태진(오른쪽, 단발머리와 작업복 일치). 전경에 의식을 잃은 찰리로 추정되는 인물의 일부가 보임.",
        "hard_violations": [],
        "physics": "두 인물 모두 차량 밖 지면에 안정적으로 서서 안을 들여다보는 자세를 취하고 있음."
       },
       {
        "label": "B",
        "direction": "두 인물의 시선이 깨진 창문을 통해 캠퍼 내부 정면을 향하고 있음.",
        "built_space": "캠퍼 내부에서 밖을 내다보는 시점이나, 내부의 싱크대와 상단 수납장이 뒤집히지 않고 정상적인 위아래 방향을 유지하고 있음.",
        "entities": "은영(왼쪽, 단발머리와 얼굴 일치), 태진(오른쪽, 단발머리와 작업복 일치).",
        "hard_violations": [
         "프롬프트에 명시된 '거꾸로 뒤집힌 차(Inverted after the crash)' 지시를 어기고 캠퍼 내부(싱크대, 수납장 등)를 정상적인 수직 방향으로 렌더링함.",
         "배경 건물 간판에 읽을 수 있는 텍스트(S-OIL mart)가 노출됨."
        ],
        "physics": "두 인물 모두 지면에 안정적으로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 미디엄 샷 프레이밍을 잘 준수하고 뒤집힌 차량 내부의 시점과 전경의 인물을 효과적으로 연출하였으나, 은영의 헤어스타일이 레퍼런스와 다소 다름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물들의 외형은 레퍼런스와 잘 일치하나, 거꾸로 뒤집혀야 할 차량 내부가 똑바로 서 있는 상태로 묘사되는 치명적인 구조적 오류가 있음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물의 시선이 깨진 창문을 통해 차량 내부(카메라 방향)를 향해 아래로 내려다보고 있음.",
        "built_space": "차량 내부에서 부서진 창틀 너머로 밖을 내다보는 시점. 프레임 가장자리에 차량 파편이 위치하며 배경에 주유소와 쓰레기 더미가 적절한 거리감으로 배치됨.",
        "entities": "은영(왼쪽, 얼굴은 일치하나 머리를 뒤로 묶음), 태진(오른쪽, 단발머리와 작업복 일치). 전경에 의식을 잃은 찰리로 추정되는 인물의 일부가 보임.",
        "hard_violations": [],
        "physics": "두 인물 모두 차량 밖 지면에 안정적으로 서서 안을 들여다보는 자세를 취하고 있음."
       },
       {
        "label": "B",
        "direction": "두 인물의 시선이 깨진 창문을 통해 캠퍼 내부 정면을 향하고 있음.",
        "built_space": "캠퍼 내부에서 밖을 내다보는 시점이나, 내부의 싱크대와 상단 수납장이 뒤집히지 않고 정상적인 위아래 방향을 유지하고 있음.",
        "entities": "은영(왼쪽, 단발머리와 얼굴 일치), 태진(오른쪽, 단발머리와 작업복 일치).",
        "hard_violations": [
         "프롬프트에 명시된 '거꾸로 뒤집힌 차(Inverted after the crash)' 지시를 어기고 캠퍼 내부(싱크대, 수납장 등)를 정상적인 수직 방향으로 렌더링함.",
         "배경 건물 간판에 읽을 수 있는 텍스트(S-OIL mart)가 노출됨."
        ],
        "physics": "두 인물 모두 지면에 안정적으로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "인물 외형은 더 가깝지만, 상체 투숏보다 실내를 넓게 보여주며 정방향 가구 배치와 읽히는 주유소 간판이 전복 상태·문자 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "깨진 창을 둘러싼 상체 투숏과 아래쪽 실내를 살피는 행동은 더 정확하지만, 전경의 제삼자 신체가 명시적 인물 제한을 위반하고 은영의 머리·의상도 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 은영과 오른쪽 태진 모두 창 너머 카메라가 있는 실내 쪽을 바라본다. 시선은 거의 정면이며 특정 부상자를 내려다보는 방향은 뚜렷하지 않지만, 차 안을 들여다본다는 기본 목표에는 맞는다. 무기나 손에 든 지시성 물체는 보이지 않는다.",
        "built_space": "중앙의 큰 파손 창 하나와 좌우에 일부 보이는 창 두 개가 있다. 왼쪽 아래에는 싱크대와 조리대, 오른쪽 위에는 목재 수납장, 중앙 창 아래에는 켜진 전구 하나가 보인다. 두 여성은 중앙 창 밖에 나란히 있다. 실내가 화면 대부분을 차지하여 인물 상체가 작다. 싱크대는 위로 열리고 수납장은 위쪽에 있어, 완전히 뒤집힌 차량보다 정상 실내를 조금 기울여 찍은 배치로 읽힌다. 창은 반사가 아닌 직접 보이는 개구부다. 밖의 노란 테두리 캐노피와 유리 상점은 구조 참고에 비교적 가깝다.",
        "entities": "등장인물은 두 여성뿐이다. 왼쪽은 은영 참고의 앞머리 있는 검은 중단발과 남색 상의, 오른쪽은 태진 참고의 짧은 검은 단발과 짙은 작업복·회색 안쪽 상의를 대체로 따른다. 두 사람 모두 지정된 연령대의 한국인 여성 설정과 모순되는 뚜렷한 외형은 없다. 깨진 유리, 쓰레기 비탈, 낮의 주유소, 옅은 실내 연기와 전구가 보인다. 정지 이미지로 전구의 깜빡임은 확인할 수 없다. 배경 간판의 영문 상호는 읽을 수 있다.",
        "hard_violations": [
         "주유소 캐노피와 상점에 읽을 수 있는 영문 상호·로고가 노출되어 문자와 로고 금지 조건을 위반한다.",
         "싱크대와 상부 수납장 등이 정상 상하관계로 놓여 있어, 완전히 전복된 캠퍼 내부라는 필수 공간 상태를 구현하지 않았다."
        ],
        "physics": "두 여성의 하체와 발은 창 아래에 가려져 있으나 몸통은 외부에 서 있는 자세로 자연스럽게 이어지며 공중에 뜬 징후는 없다. 창턱의 천과 파편은 표면에 놓여 있고 전구는 벽 쪽 고정부에 붙어 있다. 지지 없이 떠 있는 물체는 보이지 않는다. 문제는 개별 물체의 부유가 아니라 차량 전복과 맞지 않는 실내 전체의 방향이다."
       },
       {
        "label": "B",
        "direction": "왼쪽 여성은 허리를 앞으로 숙여 실내 아래쪽을 보고, 오른쪽 태진도 눈을 내려 전경의 누운 사람 쪽을 살핀다. 두 시선 모두 차 안이라는 목표에 도달하며 심각한 표정도 분명하다. 다만 시선의 구체적 대상으로 추가된 사람은 이 샷에서 허용되지 않는다.",
        "built_space": "큰 파손 창 하나가 화면 가장자리를 둘러싸며 두 여성의 상체를 직접 보여준다. 어두운 내장재와 비스듬한 창 아래 패널이 하단·측면에만 들어와, 내부에서 밖을 보는 미디엄 투숏에 가깝다. 반사 장면은 아니다. 전복 방향을 확정할 가구는 적지만 A처럼 명백한 정방향 싱크대 배치는 없다. 배경에는 노란 테두리 캐노피 하나, 유리 상점, 오른쪽 줄무늬 전주와 쓰레기 비탈이 있다. 구조 참고보다 주유소 지지부와 입면은 단순화되어 있다.",
        "entities": "창 밖에는 성인 여성 두 명이 있고, 오른쪽 태진은 검은 단발과 짙은 작업복·회색 안쪽 상의가 참고에 가깝다. 왼쪽 은영은 뒤로 묶은 듯한 머리와 회색 겉옷·밝은 안쪽 상의를 입어 참고의 앞머리 있는 중단발·남색 상의와 다르다. 전경에는 별도 인물의 검은 머리와 옷 입은 몸 일부가 보인다. 이는 찰리로 의도했더라도 은영과 태진만 보이도록 한 제한에 어긋난다. 깨진 유리와 옅은 연기, 낮의 외부는 확인되며 전구는 보이지 않는다. 읽히는 문자는 없다.",
        "hard_violations": [
         "하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 샷에 보이는 신체는 은영과 태진에게만 속해야 한다는 명시적 제한을 위반한다."
        ],
        "physics": "두 여성의 상체는 창 밖 아래로 이어지며 왼쪽의 전방 굽힘도 서서 들여다보는 행동으로 가능하다. 발은 화면 밖이므로 직접적인 지면 접촉은 확인되지 않지만 부유로 보이지 않는다. 전경 인물은 하단 실내 표면에 누운 형태로 보이고, 공중에 매달린 자세는 아니다. 유리 조각은 창틀에 붙어 있거나 아래 테두리에 걸쳐 있다. 지지 없는 물체나 불가능한 관절 형태는 뚜렷하지 않다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "인물 외형은 더 가깝지만, 상체 투숏보다 실내를 넓게 보여주며 정방향 가구 배치와 읽히는 주유소 간판이 전복 상태·문자 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "깨진 창을 둘러싼 상체 투숏과 아래쪽 실내를 살피는 행동은 더 정확하지만, 전경의 제삼자 신체가 명시적 인물 제한을 위반하고 은영의 머리·의상도 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 은영과 오른쪽 태진 모두 창 너머 카메라가 있는 실내 쪽을 바라본다. 시선은 거의 정면이며 특정 부상자를 내려다보는 방향은 뚜렷하지 않지만, 차 안을 들여다본다는 기본 목표에는 맞는다. 무기나 손에 든 지시성 물체는 보이지 않는다.",
        "built_space": "중앙의 큰 파손 창 하나와 좌우에 일부 보이는 창 두 개가 있다. 왼쪽 아래에는 싱크대와 조리대, 오른쪽 위에는 목재 수납장, 중앙 창 아래에는 켜진 전구 하나가 보인다. 두 여성은 중앙 창 밖에 나란히 있다. 실내가 화면 대부분을 차지하여 인물 상체가 작다. 싱크대는 위로 열리고 수납장은 위쪽에 있어, 완전히 뒤집힌 차량보다 정상 실내를 조금 기울여 찍은 배치로 읽힌다. 창은 반사가 아닌 직접 보이는 개구부다. 밖의 노란 테두리 캐노피와 유리 상점은 구조 참고에 비교적 가깝다.",
        "entities": "등장인물은 두 여성뿐이다. 왼쪽은 은영 참고의 앞머리 있는 검은 중단발과 남색 상의, 오른쪽은 태진 참고의 짧은 검은 단발과 짙은 작업복·회색 안쪽 상의를 대체로 따른다. 두 사람 모두 지정된 연령대의 한국인 여성 설정과 모순되는 뚜렷한 외형은 없다. 깨진 유리, 쓰레기 비탈, 낮의 주유소, 옅은 실내 연기와 전구가 보인다. 정지 이미지로 전구의 깜빡임은 확인할 수 없다. 배경 간판의 영문 상호는 읽을 수 있다.",
        "hard_violations": [
         "주유소 캐노피와 상점에 읽을 수 있는 영문 상호·로고가 노출되어 문자와 로고 금지 조건을 위반한다.",
         "싱크대와 상부 수납장 등이 정상 상하관계로 놓여 있어, 완전히 전복된 캠퍼 내부라는 필수 공간 상태를 구현하지 않았다."
        ],
        "physics": "두 여성의 하체와 발은 창 아래에 가려져 있으나 몸통은 외부에 서 있는 자세로 자연스럽게 이어지며 공중에 뜬 징후는 없다. 창턱의 천과 파편은 표면에 놓여 있고 전구는 벽 쪽 고정부에 붙어 있다. 지지 없이 떠 있는 물체는 보이지 않는다. 문제는 개별 물체의 부유가 아니라 차량 전복과 맞지 않는 실내 전체의 방향이다."
       },
       {
        "label": "A",
        "direction": "왼쪽 여성은 허리를 앞으로 숙여 실내 아래쪽을 보고, 오른쪽 태진도 눈을 내려 전경의 누운 사람 쪽을 살핀다. 두 시선 모두 차 안이라는 목표에 도달하며 심각한 표정도 분명하다. 다만 시선의 구체적 대상으로 추가된 사람은 이 샷에서 허용되지 않는다.",
        "built_space": "큰 파손 창 하나가 화면 가장자리를 둘러싸며 두 여성의 상체를 직접 보여준다. 어두운 내장재와 비스듬한 창 아래 패널이 하단·측면에만 들어와, 내부에서 밖을 보는 미디엄 투숏에 가깝다. 반사 장면은 아니다. 전복 방향을 확정할 가구는 적지만 A처럼 명백한 정방향 싱크대 배치는 없다. 배경에는 노란 테두리 캐노피 하나, 유리 상점, 오른쪽 줄무늬 전주와 쓰레기 비탈이 있다. 구조 참고보다 주유소 지지부와 입면은 단순화되어 있다.",
        "entities": "창 밖에는 성인 여성 두 명이 있고, 오른쪽 태진은 검은 단발과 짙은 작업복·회색 안쪽 상의가 참고에 가깝다. 왼쪽 은영은 뒤로 묶은 듯한 머리와 회색 겉옷·밝은 안쪽 상의를 입어 참고의 앞머리 있는 중단발·남색 상의와 다르다. 전경에는 별도 인물의 검은 머리와 옷 입은 몸 일부가 보인다. 이는 찰리로 의도했더라도 은영과 태진만 보이도록 한 제한에 어긋난다. 깨진 유리와 옅은 연기, 낮의 외부는 확인되며 전구는 보이지 않는다. 읽히는 문자는 없다.",
        "hard_violations": [
         "하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 샷에 보이는 신체는 은영과 태진에게만 속해야 한다는 명시적 제한을 위반한다."
        ],
        "physics": "두 여성의 상체는 창 밖 아래로 이어지며 왼쪽의 전방 굽힘도 서서 들여다보는 행동으로 가능하다. 발은 화면 밖이므로 직접적인 지면 접촉은 확인되지 않지만 부유로 보이지 않는다. 전경 인물은 하단 실내 표면에 누운 형태로 보이고, 공중에 매달린 자세는 아니다. 유리 조각은 창틀에 붙어 있거나 아래 테두리에 걸쳐 있다. 지지 없는 물체나 불가능한 관절 형태는 뚜렷하지 않다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.095
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.845
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 명시된 '거꾸로 뒤집힌 차(Inverted after the crash)' 지시를 어기고 캠퍼 내부(싱크대, 수납장 등)를 정상적인 수직 방향으로 렌더링함.",
     "[gemini-pro] 배경 건물 간판에 읽을 수 있는 텍스트(S-OIL mart)가 노출됨.",
     "[gpt-high] 주유소 캐노피와 상점에 읽을 수 있는 영문 상호·로고가 노출되어 문자와 로고 금지 조건을 위반한다.",
     "[gpt-high] 싱크대와 상부 수납장 등이 정상 상하관계로 놓여 있어, 완전히 전복된 캠퍼 내부라는 필수 공간 상태를 구현하지 않았다."
    ],
    "A": [
     "[gpt-high] 하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 샷에 보이는 신체는 은영과 태진에게만 속해야 한다는 명시적 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 845
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 미디엄 샷 프레이밍을 잘 준수하고 뒤집힌 차량 내부의 시점과 전경의 인물을 효과적으로 연출하였으나, 은영의 헤어스타일이 레퍼런스와 다소 다름.  ★위반: [gpt-high] 하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 샷에 보이는 신체는 은영과 태진에게만 속해야 한다는 명시적 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 845,
    "verdict_ko": "인물들의 외형은 레퍼런스와 잘 일치하나, 거꾸로 뒤집혀야 할 차량 내부가 똑바로 서 있는 상태로 묘사되는 치명적인 구조적 오류가 있음.  ★위반: [gemini-pro] 프롬프트에 명시된 '거꾸로 뒤집힌 차(Inverted after the crash)' 지시를 어기고 캠퍼 내부(싱크대, 수납장 등)를 정상적인 수직 방향으로 렌더링함. / [gemini-pro] 배경 건물 간판에 읽을 수 있는 텍스트(S-OIL mart)가 노출됨. / [gpt-high] 주유소 캐노피와 상점에 읽을 수 있는 영문 상호·로고가 노출되어 문자와 로고 금지 조건을 위반한다. / [gpt-high] 싱크대와 상부 수납장 등이 정상 상하관계로 놓여 있어, 완전히 전복된 캠퍼 내부라는 필수 공간 상태를 구현하지 않았다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B02.png",
    "asset_id": "d012fbdf-5de4-474f-8f94-751d1a131761",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_roadside_service_station_sel.png",
    "asset_id": "ab322096-550b-4b37-8cda-d1efdc3c511b",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:922276>",
    "asset_id": "d985233d-df72-452a-9cf8-93dc78e15a29",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 은영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1258866>",
    "asset_id": "d98e5c15-3084-4367-9745-556e40a4a1c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-e701-74e7-966d-372735ac1af4",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh51__bgfirst_bg.png",
   "bg_asset_id": "c25c9337-c13b-46af-8135-de9acaf2fcf0",
   "bg_record_key": "S49sh51::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S49sh51::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:16:36.601965+00:00",
  "fingerprint": "f3008bdd9d83b66f83d3d18da52c36569a2538459db8501f929dc9c69eaa12c3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S49sh51_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S49sh51_sel.png",
  "source_sha256": "197dc07a2bb4f88dbf2a045515898319eca6d83511665cfbde18486656dcbe13",
  "file": "S49sh51_cine.png",
  "staged_sha256": "2947dbe58d8d658e21fda0c4c0467f754141fbb786638b482f75376e74abcc27",
  "latency_ms": 10432
 },
 "S50sh4::signage": {
  "fp": "e9cc6e53e13eeb1e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5ab1d05fce833b16": {
  "subjects": [],
  "subject_text": "태진과 은영의 카센터 창고\n벽에 여러 록밴드 포스터가 붙은 정비 창고. 거대한 스피커와 수리 공구, 펜치, 정비용 용액 병들이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L209",
  "scope_role": "location_interior",
  "scope_sha": "fd8d31a02c054e5b"
 },
 "S50sh4::bgfirst_bg": {
  "input_fingerprint": "35650877010a2ee4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4__bgfirst_bg.png",
  "asset_id": "7a6d7074-e7b2-4b88-bab5-afa565d34ff1",
  "input_asset_ids": [
   "36a68918-70ef-4f62-a881-1c4458c88531",
   "f0f3a2cb-dd51-4ebd-a380-2b0c867d50dd"
  ]
 },
 "S50sh4": {
  "input_fingerprint": "8b3e2b281cb04858",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rock-band posters cover the room, with a large pair of pliers within reach and a large speaker in the repair area. The camper is under repair and has extensive damage to its engine and brakes. 현우: He has awakened with his injured leg bandaged and visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe. 태진: She has short bobbed hair and is stationed at the camper for repair work.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rock-band posters cover the room, with a large pair of pliers within reach and a large speaker in the repair area. The camper is under repair and has extensive damage to its engine and brakes. 현우: He has awakened with his injured leg bandaged and visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe. 태진: She has short bobbed hair and is stationed at the camper for repair work.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rock-band posters cover the room, with a large pair of pliers within reach and a large speaker in the repair area. The camper is under repair and has extensive damage to its engine and brakes. 현우: He has awakened with his injured leg bandaged and visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe. 태진: She has short bobbed hair and is stationed at the camper for repair work.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4__bgfirst_bg.png",
     "asset_id": "7a6d7074-e7b2-4b88-bab5-afa565d34ff1",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S50sh4.png",
     "asset_id": "36a68918-70ef-4f62-a881-1c4458c88531",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1170167>",
     "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L209B01.png",
     "asset_id": "f0f3a2cb-dd51-4ebd-a380-2b0c867d50dd",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1170167>",
     "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 배경을 향하고 오른손은 바닥의 펜치로 뻗어 있다. 태진은 배경 작업대 쪽에서 등을 돌리고 있다.",
    "built_space": "창고 내부. 침대, 포스터, 공구, 캠핑카가 있으나 태진의 위치가 캠핑카 옆이 아닌 벽면 작업대로 잘못 배치되었다.",
    "entities": "현우와 태진 모두 실사가 아닌 2D 애니메이션 캐릭터로 렌더링되었다. 현우의 다리에 붕대가 묘사되었다. 바닥에 펜치가 있다.",
    "hard_violations": [
     "[gemini-pro] 전체 이미지가 실사가 아닌 2D 일러스트로 생성되어 'photorealistic live-action' 및 질감 현실성(Material Realism) 조건을 완전히 위반함."
    ],
    "physics": "현우는 간이 침대에 걸터앉아 손을 뻗고 있으며, 태진은 바닥에 서서 체중을 자연스럽게 지탱하고 있다."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 화면 왼쪽 밖을 향해 있어, 우측 배경에 있는 태진의 뒷모습을 노려보라는 연출 지시를 실패했다. 현우의 손은 뻗기만 한 것이 아니라 바닥의 펜치에 완전히 닿아 있다.",
    "built_space": "실내 창고. 공간 구조와 가구, 조명이 사진 레퍼런스와 정확히 일치하며 사실적이다. 태진은 캠핑카 옆에서 작업 중이다.",
    "entities": "현우의 얼굴과 의상은 레퍼런스와 일치하나 다리의 붕대가 보이지 않는다(긴 바지에 가려지거나 누락됨). 배경의 태진, 밴드 포스터, 펜치, 캠핑카 모두 실사로 정확히 구현되었다.",
    "hard_violations": [
     "[gpt-high] 왼쪽 벽 포스터의 한글 문구가 판독 가능하여, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "현우는 쭈그리고 앉아 한 손을 펜치 위에 얹은 채 균형을 잡고 있으며, 물리적으로 지지와 자세가 자연스럽다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "지정된 실사 영화 스틸컷이 아닌 완전한 2D 애니메이션 스타일로 렌더링되어 지시를 근본적으로 위반했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "공간과 인물의 실사 구현은 뛰어나지만, 현우가 태진을 보지 않고 딴 곳을 향하며 다리의 붕대가 누락되어 핵심 연출이 빗나갔습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 배경을 향하고 오른손은 바닥의 펜치로 뻗어 있다. 태진은 배경 작업대 쪽에서 등을 돌리고 있다.",
        "built_space": "창고 내부. 침대, 포스터, 공구, 캠핑카가 있으나 태진의 위치가 캠핑카 옆이 아닌 벽면 작업대로 잘못 배치되었다.",
        "entities": "현우와 태진 모두 실사가 아닌 2D 애니메이션 캐릭터로 렌더링되었다. 현우의 다리에 붕대가 묘사되었다. 바닥에 펜치가 있다.",
        "hard_violations": [
         "전체 이미지가 실사가 아닌 2D 일러스트로 생성되어 'photorealistic live-action' 및 질감 현실성(Material Realism) 조건을 완전히 위반함."
        ],
        "physics": "현우는 간이 침대에 걸터앉아 손을 뻗고 있으며, 태진은 바닥에 서서 체중을 자연스럽게 지탱하고 있다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 화면 왼쪽 밖을 향해 있어, 우측 배경에 있는 태진의 뒷모습을 노려보라는 연출 지시를 실패했다. 현우의 손은 뻗기만 한 것이 아니라 바닥의 펜치에 완전히 닿아 있다.",
        "built_space": "실내 창고. 공간 구조와 가구, 조명이 사진 레퍼런스와 정확히 일치하며 사실적이다. 태진은 캠핑카 옆에서 작업 중이다.",
        "entities": "현우의 얼굴과 의상은 레퍼런스와 일치하나 다리의 붕대가 보이지 않는다(긴 바지에 가려지거나 누락됨). 배경의 태진, 밴드 포스터, 펜치, 캠핑카 모두 실사로 정확히 구현되었다.",
        "hard_violations": [],
        "physics": "현우는 쭈그리고 앉아 한 손을 펜치 위에 얹은 채 균형을 잡고 있으며, 물리적으로 지지와 자세가 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "지정된 실사 영화 스틸컷이 아닌 완전한 2D 애니메이션 스타일로 렌더링되어 지시를 근본적으로 위반했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "공간과 인물의 실사 구현은 뛰어나지만, 현우가 태진을 보지 않고 딴 곳을 향하며 다리의 붕대가 누락되어 핵심 연출이 빗나갔습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 배경을 향하고 오른손은 바닥의 펜치로 뻗어 있다. 태진은 배경 작업대 쪽에서 등을 돌리고 있다.",
        "built_space": "창고 내부. 침대, 포스터, 공구, 캠핑카가 있으나 태진의 위치가 캠핑카 옆이 아닌 벽면 작업대로 잘못 배치되었다.",
        "entities": "현우와 태진 모두 실사가 아닌 2D 애니메이션 캐릭터로 렌더링되었다. 현우의 다리에 붕대가 묘사되었다. 바닥에 펜치가 있다.",
        "hard_violations": [
         "전체 이미지가 실사가 아닌 2D 일러스트로 생성되어 'photorealistic live-action' 및 질감 현실성(Material Realism) 조건을 완전히 위반함."
        ],
        "physics": "현우는 간이 침대에 걸터앉아 손을 뻗고 있으며, 태진은 바닥에 서서 체중을 자연스럽게 지탱하고 있다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 화면 왼쪽 밖을 향해 있어, 우측 배경에 있는 태진의 뒷모습을 노려보라는 연출 지시를 실패했다. 현우의 손은 뻗기만 한 것이 아니라 바닥의 펜치에 완전히 닿아 있다.",
        "built_space": "실내 창고. 공간 구조와 가구, 조명이 사진 레퍼런스와 정확히 일치하며 사실적이다. 태진은 캠핑카 옆에서 작업 중이다.",
        "entities": "현우의 얼굴과 의상은 레퍼런스와 일치하나 다리의 붕대가 보이지 않는다(긴 바지에 가려지거나 누락됨). 배경의 태진, 밴드 포스터, 펜치, 캠핑카 모두 실사로 정확히 구현되었다.",
        "hard_violations": [],
        "physics": "현우는 쭈그리고 앉아 한 손을 펜치 위에 얹은 채 균형을 잡고 있으며, 물리적으로 지지와 자세가 자연스럽다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우가 태진의 반대쪽을 보고 있고 전신에 가까운 구도로 상체 중심 미디엄 숏을 벗어나며, 벽의 읽을 수 있는 한글이 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "태진을 향한 시선과 하단 중앙 펜치 위로 뻗은 손, 긴장한 상체가 요구된 순간에 가장 가까우나, 태진의 작업 위치는 캠퍼 바로 앞보다 벽 쪽 작업대에 치우쳐 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽의 출입문 쪽을 향하며, 오른쪽 뒤 캠퍼 앞에 있는 태진의 등을 겨냥하지 않는다. 아래로 뻗은 손은 하단 중앙 펜치 손잡이에 이미 닿아 있지만 움켜쥐지는 않았다. 태진은 등을 보인 채 캠퍼의 열린 정비 부위를 향한다.",
        "built_space": "왼쪽 간이침대 한 개와 열린 출입문 한 개, 뒤쪽 창문 한 개, 긴 작업대 한 개, 공구판 한 개, 높은 선반 위 대형 스피커 한 개, 오른쪽 캠퍼 한 대가 보인다. 콘크리트 벽과 바닥, 노출 천장 구조는 장소 참조와 대체로 맞는다. 현우는 침대 옆 바닥에 쪼그려 있고 태진은 캠퍼 앞에 서 있다. 현우의 발과 넓은 천장까지 포함하여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "인물은 두 명이다. 현우는 젊은 동아시아계 남성으로 검은 머리, 회색 셔츠와 올리브색 바지가 참조에 대체로 맞지만 셔츠를 열어 속옷을 드러냈다. 다리 붕대나 뚜렷한 치료 흔적은 보이지 않는다. 태진은 뒷모습이라 얼굴과 연령을 확인하기 어렵고, 짙은 남색 상의 대신 회색 작업복을 입었다. 바닥의 붉은 손잡이 대형 펜치와 녹슬고 손상된 캠퍼, 대형 스피커가 있다. 벽 포스터에는 '다시, 앞으로' 등 읽을 수 있는 한글이 남아 있다.",
        "hard_violations": [
         "왼쪽 벽 포스터의 한글 문구가 판독 가능하여, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "현우는 굽힌 두 다리와 바닥에 닿은 부츠로 낮춘 몸을 지탱하며 손을 내린다. 펜치는 바닥에 놓여 있고 손가락이 손잡이에 접촉한다. 태진은 두 발로 서서 차량 쪽으로 상체를 기울인다. 공중에 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 뒤 작업 구역의 태진을 향해 고개와 눈을 돌리고 있다. 뻗은 손 아래 하단 중앙에 펜치가 비스듬히 놓여 있으며 손과 공구 사이에는 아직 간격이 있다. 태진은 현우에게 등과 옆모습을 보인 채 작업대 위 정비 물품을 내려다본다.",
        "built_space": "왼쪽 간이침대 한 개와 열린 출입문 한 개, 뒤쪽 창문 한 개, 긴 작업대 한 개, 공구판 한 개, 오른쪽 선반 위 스피커 한 개, 오른쪽 캠퍼 한 대가 보인다. 벽 포스터, 콘크리트 재질과 낮빛이 들어오는 구조는 참조 장소에 가깝다. 현우는 침대 가장자리에 앉아 전경을 크게 차지하며, 태진은 캠퍼 뒤쪽의 작은 작업대 앞에 서 있다. 캠퍼 측면은 오른쪽에 부분적으로 보이지만 태진이 직접 캠퍼를 수리하는 관계는 다소 약하다. 현우의 상체 중심이 유지되지만 붕대 감긴 무릎까지 포함되어 미디엄 숏으로는 조금 넓다.",
        "entities": "두 인물만 보인다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 회색 셔츠와 올리브색 바지가 참조에 대체로 부합한다. 드러난 다리에는 붕대가 있으며 다른 치료 흔적은 뚜렷하지 않다. 태진은 검은 단발과 남색 상의를 갖춘 여성으로 보이며 얼굴은 돌아서 있어 세부 동일성을 확인할 수 없다. 현실적인 크기의 붉은 손잡이 펜치, 손상된 캠퍼, 대형 스피커와 밴드 연주 이미지 포스터가 있다. 판독 가능한 제목이나 문구는 보이지 않는다. 엔진과 브레이크의 구체적인 손상은 이 구도에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "현우의 엉덩이는 침대 좌면에 놓여 있고 반대쪽 손은 침대에 짚혀 있어 앞으로 기울인 상체를 지탱한다. 뻗은 팔은 어깨에 연결된 자연스러운 동작이며 아직 펜치에 닿지 않았다. 펜치는 콘크리트 바닥에 안정적으로 놓여 있다. 태진은 바닥에 딛은 발로 몸을 지탱하며 작업대 쪽으로 숙이고 있다. 지지 없는 부유나 불가능한 동작은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우가 태진의 반대쪽을 보고 있고 전신에 가까운 구도로 상체 중심 미디엄 숏을 벗어나며, 벽의 읽을 수 있는 한글이 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "태진을 향한 시선과 하단 중앙 펜치 위로 뻗은 손, 긴장한 상체가 요구된 순간에 가장 가까우나, 태진의 작업 위치는 캠퍼 바로 앞보다 벽 쪽 작업대에 치우쳐 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽의 출입문 쪽을 향하며, 오른쪽 뒤 캠퍼 앞에 있는 태진의 등을 겨냥하지 않는다. 아래로 뻗은 손은 하단 중앙 펜치 손잡이에 이미 닿아 있지만 움켜쥐지는 않았다. 태진은 등을 보인 채 캠퍼의 열린 정비 부위를 향한다.",
        "built_space": "왼쪽 간이침대 한 개와 열린 출입문 한 개, 뒤쪽 창문 한 개, 긴 작업대 한 개, 공구판 한 개, 높은 선반 위 대형 스피커 한 개, 오른쪽 캠퍼 한 대가 보인다. 콘크리트 벽과 바닥, 노출 천장 구조는 장소 참조와 대체로 맞는다. 현우는 침대 옆 바닥에 쪼그려 있고 태진은 캠퍼 앞에 서 있다. 현우의 발과 넓은 천장까지 포함하여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "인물은 두 명이다. 현우는 젊은 동아시아계 남성으로 검은 머리, 회색 셔츠와 올리브색 바지가 참조에 대체로 맞지만 셔츠를 열어 속옷을 드러냈다. 다리 붕대나 뚜렷한 치료 흔적은 보이지 않는다. 태진은 뒷모습이라 얼굴과 연령을 확인하기 어렵고, 짙은 남색 상의 대신 회색 작업복을 입었다. 바닥의 붉은 손잡이 대형 펜치와 녹슬고 손상된 캠퍼, 대형 스피커가 있다. 벽 포스터에는 '다시, 앞으로' 등 읽을 수 있는 한글이 남아 있다.",
        "hard_violations": [
         "왼쪽 벽 포스터의 한글 문구가 판독 가능하여, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "현우는 굽힌 두 다리와 바닥에 닿은 부츠로 낮춘 몸을 지탱하며 손을 내린다. 펜치는 바닥에 놓여 있고 손가락이 손잡이에 접촉한다. 태진은 두 발로 서서 차량 쪽으로 상체를 기울인다. 공중에 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 뒤 작업 구역의 태진을 향해 고개와 눈을 돌리고 있다. 뻗은 손 아래 하단 중앙에 펜치가 비스듬히 놓여 있으며 손과 공구 사이에는 아직 간격이 있다. 태진은 현우에게 등과 옆모습을 보인 채 작업대 위 정비 물품을 내려다본다.",
        "built_space": "왼쪽 간이침대 한 개와 열린 출입문 한 개, 뒤쪽 창문 한 개, 긴 작업대 한 개, 공구판 한 개, 오른쪽 선반 위 스피커 한 개, 오른쪽 캠퍼 한 대가 보인다. 벽 포스터, 콘크리트 재질과 낮빛이 들어오는 구조는 참조 장소에 가깝다. 현우는 침대 가장자리에 앉아 전경을 크게 차지하며, 태진은 캠퍼 뒤쪽의 작은 작업대 앞에 서 있다. 캠퍼 측면은 오른쪽에 부분적으로 보이지만 태진이 직접 캠퍼를 수리하는 관계는 다소 약하다. 현우의 상체 중심이 유지되지만 붕대 감긴 무릎까지 포함되어 미디엄 숏으로는 조금 넓다.",
        "entities": "두 인물만 보인다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 회색 셔츠와 올리브색 바지가 참조에 대체로 부합한다. 드러난 다리에는 붕대가 있으며 다른 치료 흔적은 뚜렷하지 않다. 태진은 검은 단발과 남색 상의를 갖춘 여성으로 보이며 얼굴은 돌아서 있어 세부 동일성을 확인할 수 없다. 현실적인 크기의 붉은 손잡이 펜치, 손상된 캠퍼, 대형 스피커와 밴드 연주 이미지 포스터가 있다. 판독 가능한 제목이나 문구는 보이지 않는다. 엔진과 브레이크의 구체적인 손상은 이 구도에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "현우의 엉덩이는 침대 좌면에 놓여 있고 반대쪽 손은 침대에 짚혀 있어 앞으로 기울인 상체를 지탱한다. 뻗은 팔은 어깨에 연결된 자연스러운 동작이며 아직 펜치에 닿지 않았다. 펜치는 콘크리트 바닥에 안정적으로 놓여 있다. 태진은 바닥에 딛은 발로 몸을 지탱하며 작업대 쪽으로 숙이고 있다. 지지 없는 부유나 불가능한 동작은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.25,
    "B": 1.25
   },
   "adjusted": {
    "A": 1.0,
    "B": 1.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 전체 이미지가 실사가 아닌 2D 일러스트로 생성되어 'photorealistic live-action' 및 질감 현실성(Material Realism) 조건을 완전히 위반함."
    ],
    "B": [
     "[gpt-high] 왼쪽 벽 포스터의 한글 문구가 판독 가능하여, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1000,
   "B": 1000
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1000,
    "verdict_ko": "지정된 실사 영화 스틸컷이 아닌 완전한 2D 애니메이션 스타일로 렌더링되어 지시를 근본적으로 위반했습니다.  ★위반: [gemini-pro] 전체 이미지가 실사가 아닌 2D 일러스트로 생성되어 'photorealistic live-action' 및 질감 현실성(Material Realism) 조건을 완전히 위반함."
   },
   {
    "label": "B",
    "score": 1000,
    "verdict_ko": "공간과 인물의 실사 구현은 뛰어나지만, 현우가 태진을 보지 않고 딴 곳을 향하며 다리의 붕대가 누락되어 핵심 연출이 빗나갔습니다.  ★위반: [gpt-high] 왼쪽 벽 포스터의 한글 문구가 판독 가능하여, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L209B01.png",
    "asset_id": "f0f3a2cb-dd51-4ebd-a380-2b0c867d50dd",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1170167>",
    "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-ea5f-7299-b6f1-f294da76a36e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4__bgfirst_bg.png",
   "bg_asset_id": "7a6d7074-e7b2-4b88-bab5-afa565d34ff1",
   "bg_record_key": "S50sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C23"
  ]
 },
 "S50sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:18:07.052426+00:00",
  "fingerprint": "30e477f74a92b16c3749fd8526ee278c1c297a483d064f441f5c038c1715003e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S50sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S50sh4_sel.png",
  "source_sha256": "69be78f4ee28f1fc77c8a161c5e74015df44e6ea50060ee75b2162d60b9a4570",
  "file": "S50sh4_cine.png",
  "staged_sha256": "eedae219af4d5c81e459b7fb65e72dabe14dc3fa2f278d8836d3bcac65a552f4",
  "latency_ms": 10807
 },
 "S50sh9::signage": {
  "fp": "1a15b76936124854",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S50sh9": {
  "input_fingerprint": "eb3800f8f66f0aed",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 완전히 망가진 캠핑카를 가리키며 미간을 찌푸린 태진의 상체.\n\nLOCATION (lock): Beside the damaged camper inside the auto-repair workshop bay, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Damaged camper beside 태진, receiving her pointing gesture in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Damaged camper (Under repair, with major engine and brake damage described in dialogue) — Only the side section beside 태진 is visible; hidden mechanical damage is not illustrated as a cutaway; used as Receives her pointing gesture and supplies evidence of the subject under discussion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's neutral ambient illumination and controlled contrast across 태진 and the adjacent camper.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 태진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage, and repair fluid has just been poured into its filler opening. Rock-band posters, the large speaker and the large pliers remain in the workshop space. 태진: She has short bobbed hair and holds the repair-fluid bottle after pouring from it at the camper.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 완전히 망가진 캠핑카를 가리키며 미간을 찌푸린 태진의 상체.\n\nLOCATION (lock): Beside the damaged camper inside the auto-repair workshop bay, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Damaged camper beside 태진, receiving her pointing gesture in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Damaged camper (Under repair, with major engine and brake damage described in dialogue) — Only the side section beside 태진 is visible; hidden mechanical damage is not illustrated as a cutaway; used as Receives her pointing gesture and supplies evidence of the subject under discussion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's neutral ambient illumination and controlled contrast across 태진 and the adjacent camper.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 태진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage, and repair fluid has just been poured into its filler opening. Rock-band posters, the large speaker and the large pliers remain in the workshop space. 태진: She has short bobbed hair and holds the repair-fluid bottle after pouring from it at the camper.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 완전히 망가진 캠핑카를 가리키며 미간을 찌푸린 태진의 상체.\n\nLOCATION (lock): Beside the damaged camper inside the auto-repair workshop bay, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Damaged camper beside 태진, receiving her pointing gesture in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Damaged camper (Under repair, with major engine and brake damage described in dialogue) — Only the side section beside 태진 is visible; hidden mechanical damage is not illustrated as a cutaway; used as Receives her pointing gesture and supplies evidence of the subject under discussion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's neutral ambient illumination and controlled contrast across 태진 and the adjacent camper.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 태진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage, and repair fluid has just been poured into its filler opening. Rock-band posters, the large speaker and the large pliers remain in the workshop space. 태진: She has short bobbed hair and holds the repair-fluid bottle after pouring from it at the camper.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "태진의 시선과 왼손 검지가 캠핑카 측면에 비정상적으로 뚫린 엔진 노출부를 향하고 있습니다.",
    "built_space": "정비소 내부. 왼쪽 벽에 포스터, 정면 벽에 창문과 대형 스피커가 올려진 선반이 있으며, 바닥에 절단기가 배치되어 있습니다.",
    "entities": "태진의 인상착의는 참조와 일치하나, 필수 소품인 수리 용액 병이 없습니다. 캠핑카 측면에 거대한 엔진 내부가 부자연스럽게 노출되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트가 명시적으로 금지한 숨겨진 기계적 손상의 단면(cutaway) 묘사가 캠핑카 측면에 적용됨",
     "[gemini-pro] 수리 용액 병을 들고 있어야 한다는 지시 누락",
     "[gpt-high] 참조의 닫힌 운전석 측면을 대규모로 절개하고 그 자리에 거대한 노출 엔진을 만들어 넣었습니다. 이는 단순한 외판 손상 차이가 아니라, 명시적으로 금지된 숨은 기계부의 절개 전시를 장면에 추가한 것입니다."
    ],
    "physics": "두 발로 바닥에 서 있으며, 가리키는 왼손 외에 오른팔은 아무것도 쥐지 않은 채 아래로 향해 있습니다."
   },
   {
    "label": "B",
    "direction": "태진의 시선과 왼손 검지가 캠핑카의 손상된 측면 수납부를 정확히 향하고 있습니다.",
    "built_space": "정비소 내부. 창문, 공구 게시판, 대형 스피커가 올려진 선반, 록 밴드 포스터가 벽면에 배치되어 있으며, 바닥에는 붉은 손잡이의 대형 절단기가 놓여 있습니다.",
    "entities": "태진의 인상착의, 의상 및 헤어스타일이 참조 이미지와 일치하며 오른손에 수리 용액 병을 들고 있습니다. 캠핑카는 측면 줄무늬와 녹슨 자국이 잘 묘사되어 있습니다.",
    "hard_violations": [],
    "physics": "태진은 바닥에 안정적으로 서 있으며, 오른손은 병을 단단히 쥐고 있고 왼손은 자연스럽게 뻗어 대상을 가리키고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미간을 찌푸린 표정과 가리키는 동작, 수리 용액 병을 든 상태 및 배경의 연속성 요소를 모두 정확하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "수리 용액 병을 들고 있어야 한다는 유지(carried state) 지시를 누락했으며, 명시적으로 금지된 차량 내부 투시(cutaway)를 묘사해 치명적인 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "태진의 시선과 왼손 검지가 캠핑카의 손상된 측면 수납부를 정확히 향하고 있습니다.",
        "built_space": "정비소 내부. 창문, 공구 게시판, 대형 스피커가 올려진 선반, 록 밴드 포스터가 벽면에 배치되어 있으며, 바닥에는 붉은 손잡이의 대형 절단기가 놓여 있습니다.",
        "entities": "태진의 인상착의, 의상 및 헤어스타일이 참조 이미지와 일치하며 오른손에 수리 용액 병을 들고 있습니다. 캠핑카는 측면 줄무늬와 녹슨 자국이 잘 묘사되어 있습니다.",
        "hard_violations": [],
        "physics": "태진은 바닥에 안정적으로 서 있으며, 오른손은 병을 단단히 쥐고 있고 왼손은 자연스럽게 뻗어 대상을 가리키고 있습니다."
       },
       {
        "label": "A",
        "direction": "태진의 시선과 왼손 검지가 캠핑카 측면에 비정상적으로 뚫린 엔진 노출부를 향하고 있습니다.",
        "built_space": "정비소 내부. 왼쪽 벽에 포스터, 정면 벽에 창문과 대형 스피커가 올려진 선반이 있으며, 바닥에 절단기가 배치되어 있습니다.",
        "entities": "태진의 인상착의는 참조와 일치하나, 필수 소품인 수리 용액 병이 없습니다. 캠핑카 측면에 거대한 엔진 내부가 부자연스럽게 노출되어 있습니다.",
        "hard_violations": [
         "프롬프트가 명시적으로 금지한 숨겨진 기계적 손상의 단면(cutaway) 묘사가 캠핑카 측면에 적용됨",
         "수리 용액 병을 들고 있어야 한다는 지시 누락"
        ],
        "physics": "두 발로 바닥에 서 있으며, 가리키는 왼손 외에 오른팔은 아무것도 쥐지 않은 채 아래로 향해 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미간을 찌푸린 표정과 가리키는 동작, 수리 용액 병을 든 상태 및 배경의 연속성 요소를 모두 정확하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "수리 용액 병을 들고 있어야 한다는 유지(carried state) 지시를 누락했으며, 명시적으로 금지된 차량 내부 투시(cutaway)를 묘사해 치명적인 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "태진의 시선과 왼손 검지가 캠핑카의 손상된 측면 수납부를 정확히 향하고 있습니다.",
        "built_space": "정비소 내부. 창문, 공구 게시판, 대형 스피커가 올려진 선반, 록 밴드 포스터가 벽면에 배치되어 있으며, 바닥에는 붉은 손잡이의 대형 절단기가 놓여 있습니다.",
        "entities": "태진의 인상착의, 의상 및 헤어스타일이 참조 이미지와 일치하며 오른손에 수리 용액 병을 들고 있습니다. 캠핑카는 측면 줄무늬와 녹슨 자국이 잘 묘사되어 있습니다.",
        "hard_violations": [],
        "physics": "태진은 바닥에 안정적으로 서 있으며, 오른손은 병을 단단히 쥐고 있고 왼손은 자연스럽게 뻗어 대상을 가리키고 있습니다."
       },
       {
        "label": "A",
        "direction": "태진의 시선과 왼손 검지가 캠핑카 측면에 비정상적으로 뚫린 엔진 노출부를 향하고 있습니다.",
        "built_space": "정비소 내부. 왼쪽 벽에 포스터, 정면 벽에 창문과 대형 스피커가 올려진 선반이 있으며, 바닥에 절단기가 배치되어 있습니다.",
        "entities": "태진의 인상착의는 참조와 일치하나, 필수 소품인 수리 용액 병이 없습니다. 캠핑카 측면에 거대한 엔진 내부가 부자연스럽게 노출되어 있습니다.",
        "hard_violations": [
         "프롬프트가 명시적으로 금지한 숨겨진 기계적 손상의 단면(cutaway) 묘사가 캠핑카 측면에 적용됨",
         "수리 용액 병을 들고 있어야 한다는 지시 누락"
        ],
        "physics": "두 발로 바닥에 서 있으며, 가리키는 왼손 외에 오른팔은 아무것도 쥐지 않은 채 아래로 향해 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찌푸린 미간과 캠핑카 측면을 향한 손짓, 손에 든 수리액 병을 구현했지만, 상체 중심 미디엄 숏보다 넓고 차량 점검구가 참조와 다릅니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상체의 크기와 손짓은 더 적절하지만, 운전석 측면을 절개해 거대한 엔진을 노출한 것은 숨은 기계 손상을 도해처럼 보여주지 말라는 핵심 조건을 위반합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "태진의 뻗은 검지는 화면 중앙 오른쪽 캠핑카의 열린 점검구 아래 녹슨 외판을 향합니다. 시선은 손끝보다 높은 차량 쪽을 향하고, 미간은 뚜렷하게 찌푸려져 있습니다. 반대 손의 병은 아래로 내려 들고 있어 붓는 중이 아닌 상태입니다.",
        "built_space": "오른쪽에 캠핑카 한 대의 측면 일부가 있고, 왼쪽 뒤에 창 하나, 작업대 하나, 공구 걸이판 하나, 선반 위 대형 스피커 하나가 보입니다. 밴드 포스터와 바닥의 대형 펜치 한 개, 왼쪽 가장자리의 의자 일부도 있습니다. 콘크리트 벽과 작업장 재질은 이어지지만, 차량에는 참조에서 확인되지 않는 위아래 점검구가 생겼습니다. 인물을 허벅지까지 보여 주어 요청한 상체 중심 구도보다 넓습니다.",
        "entities": "실제 인물은 태진 한 명뿐이며, 약 30세 한국인 여성이라는 설정에 부합하는 외모, 검은 단발, 남색 작업복과 회색 속옷, 손목시계가 보입니다. 얼굴과 머리는 인물 참조에 가깝습니다. 다른 손에는 회색 수리액 병이 있습니다. 흰색 차체의 회색 줄무늬와 심한 부식은 참조 캠핑카의 특징을 따릅니다. 포스터의 인쇄 인물 외에 추가 사람은 없고, 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "인물의 하체는 프레임 아래로 이어지고 몸통과 팔의 연결은 자연스럽습니다. 병은 아래로 내린 손이 직접 감싸 쥐고 있습니다. 캠핑카는 보이는 바퀴로 바닥에 놓여 있고, 열린 위쪽 점검구 덮개는 경첩과 지지대로 연결되어 있습니다. 펜치는 바닥에, 스피커는 선반에 놓여 있어 지지 없는 부유 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "태진의 검지는 중앙 오른쪽 캠핑카 측면의 찢어지고 녹슨 외판을 정확히 향하며, 시선도 그 손끝과 손상 부위 쪽으로 내려갑니다. 미간을 찌푸린 표정이 보입니다. 내려 둔 반대 손은 화면 아래에서 잘려 병의 방향은 확인할 수 없습니다.",
        "built_space": "오른쪽 캠핑카 한 대와 뒤쪽 창 하나, 작업대 하나, 공구판 일부, 선반 위 스피커 하나, 왼쪽 밴드 포스터들이 보입니다. 대형 펜치 한 개도 바닥에 있습니다. 작업장의 콘크리트 벽과 낮 조명은 대체로 이어지고 상체 중심 크기도 A보다 가깝습니다. 그러나 참조에서 닫혀 있던 운전석 문 아래와 옆 차체가 크게 제거되고, 그 자리를 거대한 엔진과 배관이 차지합니다.",
        "entities": "태진 한 명의 얼굴, 검은 단발, 남색 작업복과 회색 속옷, 손목시계는 인물 참조에 대체로 부합합니다. 캠핑카의 흰색 외판과 회색 줄무늬는 유지되지만, 새로 노출된 대형 엔진이 차량의 중요한 구성 요소처럼 추가되었습니다. 수리액 병은 프레임 밖이므로 소지 여부를 판단할 수 없습니다. 읽을 수 있는 글자나 추가 실제 인물은 보이지 않습니다.",
        "hard_violations": [
         "참조의 닫힌 운전석 측면을 대규모로 절개하고 그 자리에 거대한 노출 엔진을 만들어 넣었습니다. 이는 단순한 외판 손상 차이가 아니라, 명시적으로 금지된 숨은 기계부의 절개 전시를 장면에 추가한 것입니다."
        ],
        "physics": "태진은 서 있는 몸통에서 팔을 자연스럽게 뻗고 있으며, 보이는 관절에 명백한 불가능성은 없습니다. 발과 반대 손의 끝은 화면 밖입니다. 차량은 바퀴로 지지되고 펜치는 바닥에 놓여 있습니다. 노출된 기계는 차량 내부에 연결된 것으로 보이므로 부유한다고 단정할 수는 없지만, 운전석 측면과 바퀴 주변을 차지하는 배치는 참조 차량의 닫힌 구조를 크게 바꿉니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찌푸린 미간과 캠핑카 측면을 향한 손짓, 손에 든 수리액 병을 구현했지만, 상체 중심 미디엄 숏보다 넓고 차량 점검구가 참조와 다릅니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "상체의 크기와 손짓은 더 적절하지만, 운전석 측면을 절개해 거대한 엔진을 노출한 것은 숨은 기계 손상을 도해처럼 보여주지 말라는 핵심 조건을 위반합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "태진의 뻗은 검지는 화면 중앙 오른쪽 캠핑카의 열린 점검구 아래 녹슨 외판을 향합니다. 시선은 손끝보다 높은 차량 쪽을 향하고, 미간은 뚜렷하게 찌푸려져 있습니다. 반대 손의 병은 아래로 내려 들고 있어 붓는 중이 아닌 상태입니다.",
        "built_space": "오른쪽에 캠핑카 한 대의 측면 일부가 있고, 왼쪽 뒤에 창 하나, 작업대 하나, 공구 걸이판 하나, 선반 위 대형 스피커 하나가 보입니다. 밴드 포스터와 바닥의 대형 펜치 한 개, 왼쪽 가장자리의 의자 일부도 있습니다. 콘크리트 벽과 작업장 재질은 이어지지만, 차량에는 참조에서 확인되지 않는 위아래 점검구가 생겼습니다. 인물을 허벅지까지 보여 주어 요청한 상체 중심 구도보다 넓습니다.",
        "entities": "실제 인물은 태진 한 명뿐이며, 약 30세 한국인 여성이라는 설정에 부합하는 외모, 검은 단발, 남색 작업복과 회색 속옷, 손목시계가 보입니다. 얼굴과 머리는 인물 참조에 가깝습니다. 다른 손에는 회색 수리액 병이 있습니다. 흰색 차체의 회색 줄무늬와 심한 부식은 참조 캠핑카의 특징을 따릅니다. 포스터의 인쇄 인물 외에 추가 사람은 없고, 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "인물의 하체는 프레임 아래로 이어지고 몸통과 팔의 연결은 자연스럽습니다. 병은 아래로 내린 손이 직접 감싸 쥐고 있습니다. 캠핑카는 보이는 바퀴로 바닥에 놓여 있고, 열린 위쪽 점검구 덮개는 경첩과 지지대로 연결되어 있습니다. 펜치는 바닥에, 스피커는 선반에 놓여 있어 지지 없는 부유 물체는 없습니다."
       },
       {
        "label": "A",
        "direction": "태진의 검지는 중앙 오른쪽 캠핑카 측면의 찢어지고 녹슨 외판을 정확히 향하며, 시선도 그 손끝과 손상 부위 쪽으로 내려갑니다. 미간을 찌푸린 표정이 보입니다. 내려 둔 반대 손은 화면 아래에서 잘려 병의 방향은 확인할 수 없습니다.",
        "built_space": "오른쪽 캠핑카 한 대와 뒤쪽 창 하나, 작업대 하나, 공구판 일부, 선반 위 스피커 하나, 왼쪽 밴드 포스터들이 보입니다. 대형 펜치 한 개도 바닥에 있습니다. 작업장의 콘크리트 벽과 낮 조명은 대체로 이어지고 상체 중심 크기도 A보다 가깝습니다. 그러나 참조에서 닫혀 있던 운전석 문 아래와 옆 차체가 크게 제거되고, 그 자리를 거대한 엔진과 배관이 차지합니다.",
        "entities": "태진 한 명의 얼굴, 검은 단발, 남색 작업복과 회색 속옷, 손목시계는 인물 참조에 대체로 부합합니다. 캠핑카의 흰색 외판과 회색 줄무늬는 유지되지만, 새로 노출된 대형 엔진이 차량의 중요한 구성 요소처럼 추가되었습니다. 수리액 병은 프레임 밖이므로 소지 여부를 판단할 수 없습니다. 읽을 수 있는 글자나 추가 실제 인물은 보이지 않습니다.",
        "hard_violations": [
         "참조의 닫힌 운전석 측면을 대규모로 절개하고 그 자리에 거대한 노출 엔진을 만들어 넣었습니다. 이는 단순한 외판 손상 차이가 아니라, 명시적으로 금지된 숨은 기계부의 절개 전시를 장면에 추가한 것입니다."
        ],
        "physics": "태진은 서 있는 몸통에서 팔을 자연스럽게 뻗고 있으며, 보이는 관절에 명백한 불가능성은 없습니다. 발과 반대 손의 끝은 화면 밖입니다. 차량은 바퀴로 지지되고 펜치는 바닥에 놓여 있습니다. 노출된 기계는 차량 내부에 연결된 것으로 보이므로 부유한다고 단정할 수는 없지만, 운전석 측면과 바퀴 주변을 차지하는 배치는 참조 차량의 닫힌 구조를 크게 바꿉니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트가 명시적으로 금지한 숨겨진 기계적 손상의 단면(cutaway) 묘사가 캠핑카 측면에 적용됨",
     "[gemini-pro] 수리 용액 병을 들고 있어야 한다는 지시 누락",
     "[gpt-high] 참조의 닫힌 운전석 측면을 대규모로 절개하고 그 자리에 거대한 노출 엔진을 만들어 넣었습니다. 이는 단순한 외판 손상 차이가 아니라, 명시적으로 금지된 숨은 기계부의 절개 전시를 장면에 추가한 것입니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "미간을 찌푸린 표정과 가리키는 동작, 수리 용액 병을 든 상태 및 배경의 연속성 요소를 모두 정확하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "수리 용액 병을 들고 있어야 한다는 유지(carried state) 지시를 누락했으며, 명시적으로 금지된 차량 내부 투시(cutaway)를 묘사해 치명적인 오류를 범했습니다.  ★위반: [gemini-pro] 프롬프트가 명시적으로 금지한 숨겨진 기계적 손상의 단면(cutaway) 묘사가 캠핑카 측면에 적용됨 / [gemini-pro] 수리 용액 병을 들고 있어야 한다는 지시 누락 / [gpt-high] 참조의 닫힌 운전석 측면을 대규모로 절개하고 그 자리에 거대한 노출 엔진을 만들어 넣었습니다. 이는 단순한 외판 손상 차이가 아니라, 명시적으로 금지된 숨은 기계부의 절개 전시를 장면에 추가한 것입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 태진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4_sel.png",
    "asset_id": "5860b257-d2b9-4f82-934d-680ad19eef89",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:922276>",
    "asset_id": "d985233d-df72-452a-9cf8-93dc78e15a29",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-edb2-7745-a8d5-399f32d04022",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S50sh4"
  }
 },
 "S50sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:19:08.583065+00:00",
  "fingerprint": "b31e975b03a3bdfffb208df7b63ced996bfd294c4c0769aa2a3a32ca198d81f3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S50sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S50sh9_sel.png",
  "source_sha256": "09b8d49ec4b724e0cd6a66980c2c3422e6478f9bde3b9324b52792f58431c21f",
  "file": "S50sh9_cine.png",
  "staged_sha256": "ad9b0ed7b2acb723e6ae2f194664422189697d8f83d299ee3213e496e96f6532",
  "latency_ms": 10979
 },
 "S50sh10::signage": {
  "fp": "91250177605fb4e8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S50sh10": {
  "input_fingerprint": "1499fc55a05f8d6e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잔뜩 미간을 찌푸린 채 앰버의 행방을 묻듯 다급하게 입을 연 현우의 얼굴.\n\nLOCATION (lock): At the resting area adjoining the repair bay inside the auto-center warehouse, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rock-band posters (Attached to the wall behind the seated area) — Their printed faces remain partially visible and softly resolved behind 현우; used as Maintains location continuity without competing with his face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral interior light and restrained contrast, allowing urgency to register through expression rather than a lighting shift.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the recovery area's interior surfaces, rock-band posters, and the tools already present there. Exclude the repair bay as a substitute for this resting area.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage and newly added repair fluid. The workshop retains its rock-band posters, large speaker and large pliers. 현우: His injured leg remains bandaged, with visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잔뜩 미간을 찌푸린 채 앰버의 행방을 묻듯 다급하게 입을 연 현우의 얼굴.\n\nLOCATION (lock): At the resting area adjoining the repair bay inside the auto-center warehouse, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rock-band posters (Attached to the wall behind the seated area) — Their printed faces remain partially visible and softly resolved behind 현우; used as Maintains location continuity without competing with his face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral interior light and restrained contrast, allowing urgency to register through expression rather than a lighting shift.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the recovery area's interior surfaces, rock-band posters, and the tools already present there. Exclude the repair bay as a substitute for this resting area.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage and newly added repair fluid. The workshop retains its rock-band posters, large speaker and large pliers. 현우: His injured leg remains bandaged, with visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잔뜩 미간을 찌푸린 채 앰버의 행방을 묻듯 다급하게 입을 연 현우의 얼굴.\n\nLOCATION (lock): At the resting area adjoining the repair bay inside the auto-center warehouse, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rock-band posters (Attached to the wall behind the seated area) — Their printed faces remain partially visible and softly resolved behind 현우; used as Maintains location continuity without competing with his face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral interior light and restrained contrast, allowing urgency to register through expression rather than a lighting shift.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the recovery area's interior surfaces, rock-band posters, and the tools already present there. Exclude the repair bay as a substitute for this resting area.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage and newly added repair fluid. The workshop retains its rock-band posters, large speaker and large pliers. 현우: His injured leg remains bandaged, with visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 카메라 프레임 바깥의 특정 대상을 향해 다급하게 고정되어 있음.",
    "built_space": "레퍼런스 이미지에서 볼 수 있는 회색 콘크리트 벽면, 전선 배관, 선반 모서리, 그리고 흐릿하게 처리된 록 밴드 포스터가 알맞은 비율로 배경에 위치함.",
    "entities": "현우(18세 남성, 앳된 얼굴, 헝클어진 검은 머리)가 레퍼런스와 일치하는 회색 셔츠를 입고 있으며, 다급하게 미간을 찌푸리고 입을 연 표정이 정확히 묘사됨.",
    "hard_violations": [],
    "physics": "인물이 땅에 서서 안정적인 자세를 취하고 있으며 물리적 오류 없음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 프레임 바깥을 향하며 다급한 느낌을 줌.",
    "built_space": "뒤편에 여러 장의 밴드 포스터가 부착되어 있고 좌측에 스피커 모서리가 보이나, 벽면의 질감이 레퍼런스의 콘크리트 벽보다 매끄럽게 묘사됨.",
    "entities": "지시된 현우의 외모와 복장이 잘 일치하며, 찌푸린 표정과 열린 입도 묘사되었으나 배경 포스터의 인물들이 다소 시선을 분산시킴.",
    "hard_violations": [],
    "physics": "인물의 자세와 옷 주름이 자연스럽게 지탱되며 물리적 오류 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 샷과 다급하게 찌푸린 표정을 훌륭하게 표현했으며, 레퍼런스의 콘크리트 벽면과 포스터 등 배경 요소를 자연스럽게 연속성 있게 유지함."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "표정과 인물 묘사는 양호하나, 배경의 포스터들이 너무 크고 뚜렷하게 묘사되어 레퍼런스의 벽면 질감 및 환경과의 일치도가 다소 떨어짐."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 카메라 프레임 바깥의 특정 대상을 향해 다급하게 고정되어 있음.",
        "built_space": "레퍼런스 이미지에서 볼 수 있는 회색 콘크리트 벽면, 전선 배관, 선반 모서리, 그리고 흐릿하게 처리된 록 밴드 포스터가 알맞은 비율로 배경에 위치함.",
        "entities": "현우(18세 남성, 앳된 얼굴, 헝클어진 검은 머리)가 레퍼런스와 일치하는 회색 셔츠를 입고 있으며, 다급하게 미간을 찌푸리고 입을 연 표정이 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "인물이 땅에 서서 안정적인 자세를 취하고 있으며 물리적 오류 없음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 프레임 바깥을 향하며 다급한 느낌을 줌.",
        "built_space": "뒤편에 여러 장의 밴드 포스터가 부착되어 있고 좌측에 스피커 모서리가 보이나, 벽면의 질감이 레퍼런스의 콘크리트 벽보다 매끄럽게 묘사됨.",
        "entities": "지시된 현우의 외모와 복장이 잘 일치하며, 찌푸린 표정과 열린 입도 묘사되었으나 배경 포스터의 인물들이 다소 시선을 분산시킴.",
        "hard_violations": [],
        "physics": "인물의 자세와 옷 주름이 자연스럽게 지탱되며 물리적 오류 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 샷과 다급하게 찌푸린 표정을 훌륭하게 표현했으며, 레퍼런스의 콘크리트 벽면과 포스터 등 배경 요소를 자연스럽게 연속성 있게 유지함."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "표정과 인물 묘사는 양호하나, 배경의 포스터들이 너무 크고 뚜렷하게 묘사되어 레퍼런스의 벽면 질감 및 환경과의 일치도가 다소 떨어짐."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 카메라 프레임 바깥의 특정 대상을 향해 다급하게 고정되어 있음.",
        "built_space": "레퍼런스 이미지에서 볼 수 있는 회색 콘크리트 벽면, 전선 배관, 선반 모서리, 그리고 흐릿하게 처리된 록 밴드 포스터가 알맞은 비율로 배경에 위치함.",
        "entities": "현우(18세 남성, 앳된 얼굴, 헝클어진 검은 머리)가 레퍼런스와 일치하는 회색 셔츠를 입고 있으며, 다급하게 미간을 찌푸리고 입을 연 표정이 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "인물이 땅에 서서 안정적인 자세를 취하고 있으며 물리적 오류 없음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 프레임 바깥을 향하며 다급한 느낌을 줌.",
        "built_space": "뒤편에 여러 장의 밴드 포스터가 부착되어 있고 좌측에 스피커 모서리가 보이나, 벽면의 질감이 레퍼런스의 콘크리트 벽보다 매끄럽게 묘사됨.",
        "entities": "지시된 현우의 외모와 복장이 잘 일치하며, 찌푸린 표정과 열린 입도 묘사되었으나 배경 포스터의 인물들이 다소 시선을 분산시킴.",
        "hard_violations": [],
        "physics": "인물의 자세와 옷 주름이 자연스럽게 지탱되며 물리적 오류 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "강하게 찌푸린 미간과 다급히 열린 입의 클로즈업은 충실하지만, 밝고 매끈해진 벽과 크게 드러나는 포스터 얼굴들이 기존 장소의 재질과 절제된 배경 조건에서 벗어난다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 다급한 질문 표정을 클로즈업으로 담으면서 낡은 벽면과 흐릿한 밴드 포스터, 중립적인 실내광을 유지해 장소 연속성까지 더 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 얼굴을 거의 정면으로 두고 시선을 화면 오른쪽의 카메라 밖으로 약간 돌린다. 질문을 듣는 상대는 보이지 않지만 화면 밖 상대에게 말하는 방향으로 자연스럽다. 미간을 깊이 찌푸리고 입을 열어 발화 중인 순간을 보여준다. 무기나 방향성 소품은 없다.",
        "built_space": "배경에는 왼쪽 포스터 한 장, 머리에 가려 경계가 불분명한 중앙 포스터 영역, 오른쪽 포스터 한 장이 보인다. 왼쪽 아래에는 검은 직사각형 설비 일부가 잘려 있다. 벽은 참고의 얼룩진 콘크리트보다 밝고 매끈하며, 포스터 속 얼굴들이 상당히 큰 화면 면적을 차지한다. 좌석과 공구는 클로즈업 밖이라 배치나 개수를 확인할 수 없다. 불가능한 반사나 명백히 중복된 설비는 없다.",
        "entities": "실제 인물은 현우에 해당하는 젊은 동아시아계 남성 한 명뿐이다. 헝클어진 검은 머리와 낡은 회색 셔츠는 인물 참고와 부합한다. 얼굴은 참고보다 다소 길고 성숙하게 보인다. 배경 사람들은 벽면 포스터에 인쇄된 밴드 사진이다. 뚜렷한 치료 흔적은 확인되지 않으며, 다리 붕대와 신발 속 카드 및 캠퍼는 프레임 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 정상적으로 이어지며 몸통은 화면 아래로 계속된다. 하체 지지점은 클로즈업에 포함되지 않지만 떠 있는 몸이나 무지지 물체는 보이지 않는다. 포스터는 벽면에 붙어 있고, 손에 든 물건이나 공중 동작은 없다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 왼쪽의 카메라 밖 상대를 향한다. 상대 자체는 프레임에 없으며, 시선을 둔 채 미간을 강하게 모으고 입을 벌려 다급히 묻는 순간으로 읽힌다. 총구나 가리키는 소품은 없다.",
        "built_space": "왼쪽의 공연 포스터 한 장과 오른쪽 뒤의 인물 포스터 한 장이 분명하게 보이고, 머리 뒤에는 경계가 가려진 좁은 인쇄면이 더 드러난다. 왼쪽 가장자리에 금속 선반 일부, 아래에 열린 수납함처럼 보이는 구조물 일부, 오른쪽에 세로 배선과 벽 부착 설비 한 개가 보인다. 얼룩지고 거친 회색 벽이 이전 장소와 잘 연결된다. 좌석은 화면 밖이며 수리 중인 차량을 배경 중심으로 대신 제시하지 않는다. 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "현우로 읽히는 젊은 동아시아계 남성 한 명이 등장한다. 앳된 얼굴, 검은 헝클어진 머리, 마른 체격과 해진 회색 셔츠가 인물 참고에 가깝다. 다른 실제 인물은 없고 배경 인물들은 록밴드 포스터의 인쇄 사진이다. 피부에 작은 붉은 흔적은 있으나 명확한 처치 자국인지는 확인하기 어렵다. 다리 붕대, 신발 속 카드와 캠퍼는 이 클로즈업의 범위 밖이다. 포스터 하단의 문자 형태는 흐려 읽히지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨로 자연스럽게 지지되고 몸통은 프레임 아래로 이어진다. 발이나 좌석 접촉은 보이지 않지만 공중에 떠 있다는 징후도 없다. 벌어진 입과 긴장한 얼굴 근육은 실제 발화로 가능한 형태다. 포스터와 설비는 벽에 부착되어 있으며 무지지 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "강하게 찌푸린 미간과 다급히 열린 입의 클로즈업은 충실하지만, 밝고 매끈해진 벽과 크게 드러나는 포스터 얼굴들이 기존 장소의 재질과 절제된 배경 조건에서 벗어난다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우의 다급한 질문 표정을 클로즈업으로 담으면서 낡은 벽면과 흐릿한 밴드 포스터, 중립적인 실내광을 유지해 장소 연속성까지 더 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 얼굴을 거의 정면으로 두고 시선을 화면 오른쪽의 카메라 밖으로 약간 돌린다. 질문을 듣는 상대는 보이지 않지만 화면 밖 상대에게 말하는 방향으로 자연스럽다. 미간을 깊이 찌푸리고 입을 열어 발화 중인 순간을 보여준다. 무기나 방향성 소품은 없다.",
        "built_space": "배경에는 왼쪽 포스터 한 장, 머리에 가려 경계가 불분명한 중앙 포스터 영역, 오른쪽 포스터 한 장이 보인다. 왼쪽 아래에는 검은 직사각형 설비 일부가 잘려 있다. 벽은 참고의 얼룩진 콘크리트보다 밝고 매끈하며, 포스터 속 얼굴들이 상당히 큰 화면 면적을 차지한다. 좌석과 공구는 클로즈업 밖이라 배치나 개수를 확인할 수 없다. 불가능한 반사나 명백히 중복된 설비는 없다.",
        "entities": "실제 인물은 현우에 해당하는 젊은 동아시아계 남성 한 명뿐이다. 헝클어진 검은 머리와 낡은 회색 셔츠는 인물 참고와 부합한다. 얼굴은 참고보다 다소 길고 성숙하게 보인다. 배경 사람들은 벽면 포스터에 인쇄된 밴드 사진이다. 뚜렷한 치료 흔적은 확인되지 않으며, 다리 붕대와 신발 속 카드 및 캠퍼는 프레임 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 정상적으로 이어지며 몸통은 화면 아래로 계속된다. 하체 지지점은 클로즈업에 포함되지 않지만 떠 있는 몸이나 무지지 물체는 보이지 않는다. 포스터는 벽면에 붙어 있고, 손에 든 물건이나 공중 동작은 없다."
       },
       {
        "label": "A",
        "direction": "현우의 시선은 화면 왼쪽의 카메라 밖 상대를 향한다. 상대 자체는 프레임에 없으며, 시선을 둔 채 미간을 강하게 모으고 입을 벌려 다급히 묻는 순간으로 읽힌다. 총구나 가리키는 소품은 없다.",
        "built_space": "왼쪽의 공연 포스터 한 장과 오른쪽 뒤의 인물 포스터 한 장이 분명하게 보이고, 머리 뒤에는 경계가 가려진 좁은 인쇄면이 더 드러난다. 왼쪽 가장자리에 금속 선반 일부, 아래에 열린 수납함처럼 보이는 구조물 일부, 오른쪽에 세로 배선과 벽 부착 설비 한 개가 보인다. 얼룩지고 거친 회색 벽이 이전 장소와 잘 연결된다. 좌석은 화면 밖이며 수리 중인 차량을 배경 중심으로 대신 제시하지 않는다. 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "현우로 읽히는 젊은 동아시아계 남성 한 명이 등장한다. 앳된 얼굴, 검은 헝클어진 머리, 마른 체격과 해진 회색 셔츠가 인물 참고에 가깝다. 다른 실제 인물은 없고 배경 인물들은 록밴드 포스터의 인쇄 사진이다. 피부에 작은 붉은 흔적은 있으나 명확한 처치 자국인지는 확인하기 어렵다. 다리 붕대, 신발 속 카드와 캠퍼는 이 클로즈업의 범위 밖이다. 포스터 하단의 문자 형태는 흐려 읽히지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨로 자연스럽게 지지되고 몸통은 프레임 아래로 이어진다. 발이나 좌석 접촉은 보이지 않지만 공중에 떠 있다는 징후도 없다. 벌어진 입과 긴장한 얼굴 근육은 실제 발화로 가능한 형태다. 포스터와 설비는 벽에 부착되어 있으며 무지지 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.603
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.603
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1603
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 샷과 다급하게 찌푸린 표정을 훌륭하게 표현했으며, 레퍼런스의 콘크리트 벽면과 포스터 등 배경 요소를 자연스럽게 연속성 있게 유지함."
   },
   {
    "label": "B",
    "score": 1603,
    "verdict_ko": "표정과 인물 묘사는 양호하나, 배경의 포스터들이 너무 크고 뚜렷하게 묘사되어 레퍼런스의 벽면 질감 및 환경과의 일치도가 다소 떨어짐."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh9_sel.png",
    "asset_id": "ec2c3e08-2bfb-43a2-8bac-621006966707",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-ef5a-7fd1-a429-5dbb56888136",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S50sh9"
  }
 },
 "S50sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:20:07.318043+00:00",
  "fingerprint": "6024c97c369e9a28d82a5420a21ef468733ffc0ed64883480f69891ddad3f0a4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S50sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S50sh10_sel.png",
  "source_sha256": "8706ac6f0630f74fe6868d18c6e2c1df8d3cc88374b13fd59f963c05051f87d7",
  "file": "S50sh10_cine.png",
  "staged_sha256": "bb45feadacaaa45ac60eedf642b3f30efef44d41d4769f3a7746198597c586ee",
  "latency_ms": 11416
 },
 "S51sh10::signage": {
  "fp": "40137920c01153b5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::511bcff39f0e633e": {
  "subjects": [],
  "subject_text": "카센터 컨테이너 간이식당\n정비 시설 옆 컨테이너 안에 마련된 간이식당. 비좁은 직사각형 공간에 식탁과 좌석이 놓여 있고 한쪽에 출입문이 있다.",
  "identity": "canonical",
  "scope_id": "L211",
  "scope_role": "location_interior",
  "scope_sha": "01d05264bb880c65"
 },
 "S51sh10::bgfirst_bg": {
  "input_fingerprint": "8865703c2344ba96",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh10__bgfirst_bg.png",
  "asset_id": "683d25c0-58aa-44f7-a776-222b7b276425",
  "input_asset_ids": [
   "30650cc7-b31b-433e-b1bc-a519c8b82177",
   "78695b5c-f8e2-4047-9612-b490195deea7"
  ]
 },
 "S51sh10": {
  "input_fingerprint": "03ade6b204e060df",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bowls of rice soup are in the container dining area beside the repair shop. Charlie has been repaired and can rotate his arm again; the previous shoulder malfunction should no longer be shown as active. 현우: His leg remains bandaged and his body retains treatment marks. The contact card remains hidden in his shoe. 앰버: She is in the container dining area, eating rice soup.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bowls of rice soup are in the container dining area beside the repair shop. Charlie has been repaired and can rotate his arm again; the previous shoulder malfunction should no longer be shown as active. 현우: His leg remains bandaged and his body retains treatment marks. The contact card remains hidden in his shoe. 앰버: She is in the container dining area, eating rice soup.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bowls of rice soup are in the container dining area beside the repair shop. Charlie has been repaired and can rotate his arm again; the previous shoulder malfunction should no longer be shown as active. 현우: His leg remains bandaged and his body retains treatment marks. The contact card remains hidden in his shoe. 앰버: She is in the container dining area, eating rice soup.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh10__bgfirst_bg.png",
     "asset_id": "683d25c0-58aa-44f7-a776-222b7b276425",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S51sh10.png",
     "asset_id": "30650cc7-b31b-433e-b1bc-a519c8b82177",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B01.png",
     "asset_id": "78695b5c-f8e2-4047-9612-b490195deea7",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우가 화면 왼쪽 앞의 앰버를 향해 시선을 고정하고 소리치고 있음.",
    "built_space": "식당 내부. 나무 테이블과 배경의 창문, 벽면 구조가 원본 위치와 매우 흡사함.",
    "entities": "현우와 앰버의 외형이 참조와 일치함. 테이블 위 그릇과 금속 컵들이 보임.",
    "hard_violations": [],
    "physics": "현우가 상체를 기울여 자세를 유지하고 있으며, 테이블 위 사물들의 지지 상태가 정상적임."
   },
   {
    "label": "B",
    "direction": "현우가 화면 앞쪽의 앰버를 향해 시선을 맞추고 입을 크게 벌려 소리치고 있음.",
    "built_space": "식당 내부. 얕은 심도로 처리된 창문과 나무 테이블 등 위치 참조와 일치함.",
    "entities": "현우(검은 머리, 회색 셔츠)와 앰버(금발 뒷모습)가 참조와 일치. 테이블 위에 국밥 그릇이 있음.",
    "hard_violations": [],
    "physics": "현우가 테이블 쪽으로 몸을 자연스럽게 기울였으며, 그릇과 수저가 테이블 위에 안정적으로 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 클로즈업 프레이밍과 얕은 심도를 정확히 구현했으며, 눈을 부릅뜨고 고함치는 현우의 분노한 표정이 지시문에 완벽히 부합합니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "상황과 공간적 배경은 잘 구현했으나 샷 크기가 지시된 클로즈업보다는 미디엄 샷에 가깝습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우가 화면 앞쪽의 앰버를 향해 시선을 맞추고 입을 크게 벌려 소리치고 있음.",
        "built_space": "식당 내부. 얕은 심도로 처리된 창문과 나무 테이블 등 위치 참조와 일치함.",
        "entities": "현우(검은 머리, 회색 셔츠)와 앰버(금발 뒷모습)가 참조와 일치. 테이블 위에 국밥 그릇이 있음.",
        "hard_violations": [],
        "physics": "현우가 테이블 쪽으로 몸을 자연스럽게 기울였으며, 그릇과 수저가 테이블 위에 안정적으로 놓여 있음."
       },
       {
        "label": "A",
        "direction": "현우가 화면 왼쪽 앞의 앰버를 향해 시선을 고정하고 소리치고 있음.",
        "built_space": "식당 내부. 나무 테이블과 배경의 창문, 벽면 구조가 원본 위치와 매우 흡사함.",
        "entities": "현우와 앰버의 외형이 참조와 일치함. 테이블 위 그릇과 금속 컵들이 보임.",
        "hard_violations": [],
        "physics": "현우가 상체를 기울여 자세를 유지하고 있으며, 테이블 위 사물들의 지지 상태가 정상적임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 클로즈업 프레이밍과 얕은 심도를 정확히 구현했으며, 눈을 부릅뜨고 고함치는 현우의 분노한 표정이 지시문에 완벽히 부합합니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "상황과 공간적 배경은 잘 구현했으나 샷 크기가 지시된 클로즈업보다는 미디엄 샷에 가깝습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우가 화면 앞쪽의 앰버를 향해 시선을 맞추고 입을 크게 벌려 소리치고 있음.",
        "built_space": "식당 내부. 얕은 심도로 처리된 창문과 나무 테이블 등 위치 참조와 일치함.",
        "entities": "현우(검은 머리, 회색 셔츠)와 앰버(금발 뒷모습)가 참조와 일치. 테이블 위에 국밥 그릇이 있음.",
        "hard_violations": [],
        "physics": "현우가 테이블 쪽으로 몸을 자연스럽게 기울였으며, 그릇과 수저가 테이블 위에 안정적으로 놓여 있음."
       },
       {
        "label": "A",
        "direction": "현우가 화면 왼쪽 앞의 앰버를 향해 시선을 고정하고 소리치고 있음.",
        "built_space": "식당 내부. 나무 테이블과 배경의 창문, 벽면 구조가 원본 위치와 매우 흡사함.",
        "entities": "현우와 앰버의 외형이 참조와 일치함. 테이블 위 그릇과 금속 컵들이 보임.",
        "hard_violations": [],
        "physics": "현우가 상체를 기울여 자세를 유지하고 있으며, 테이블 위 사물들의 지지 상태가 정상적임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "앰버에게 시선을 꽂고 입을 크게 벌린 현우의 얼굴을 밀착해 담아, 분노한 얼굴의 클로즈업이라는 최우선 지시를 더 충실히 구현한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "앰버를 향한 고함과 장소의 구조는 잘 맞지만, 앰버의 상반신과 넓은 식탁까지 보여 주어 요구한 얼굴 클로즈업보다 구도가 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 두 눈은 화면 왼쪽 전경의 금발 앰버를 향한다. 입을 크게 벌리고 눈썹을 강하게 모아 상대에게 고함치는 순간으로 읽힌다. 앰버도 현우 쪽으로 머리를 돌리고 있으나 눈은 보이지 않는다. 겨누는 무기나 이동하는 물체는 없다.",
        "built_space": "현우 뒤로 금속 창틀의 창 한 구간, 회색 벽, 목재 식탁과 벤치 일부, 오른쪽 밥솥 한 대가 흐리게 보인다. 전경 식탁을 사이에 둔 두 사람의 위치는 자연스럽다. 참조의 낡은 컨테이너 식당 재질은 유지되지만, 창 왼쪽에 조리 선반이 있던 참조와 달리 여기서는 밥솥이 창 오른쪽에 보여 고정 설비 배치의 일치는 약하다. 배경을 좁게 남긴 얼굴 중심 구도는 지시에 가깝다.",
        "entities": "현우 한 명과 앰버의 흐린 머리 일부가 보이며 다른 손님은 없다. 현우의 헝클어진 검은 머리, 젊은 한국계 남성 외형, 회색 셔츠는 참조에 대체로 맞지만 얼굴은 조금 더 성숙하고 각져 보인다. 앰버는 금발과 밝은 피부만 확인되므로 얼굴 정체성과 정확한 나이는 판별하기 어렵다. 식탁에는 쌀이 보이는 국그릇 한 개, 숟가락 한 개, 젓가락 한 쌍과 금속 컵 한 개가 보인다. 다리 붕대와 신발 속 카드는 프레임 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 머리는 목과 어깨에 자연스럽게 이어지고 상체를 식탁 너머로 숙인 자세다. 골반과 좌석은 화면 밖이므로 직접적인 지지점은 확인할 수 없지만 몸이 공중에 떠 있는 묘사는 아니다. 국그릇과 컵은 식탁에 놓여 있고 숟가락은 그릇 안쪽과 가장자리에 걸쳐 지지된다. 얼굴과 입의 구조도 실제 고함 연기로 가능한 범위다."
       },
       {
        "label": "B",
        "direction": "현우는 왼쪽 전경에 앉은 앰버의 얼굴을 똑바로 바라보며 입을 크게 벌린다. 앰버 역시 현우 쪽을 향한다. 고함의 대상은 명확하지만 눈을 부릅뜬 느낌보다 눈썹을 찌푸린 분노가 더 두드러진다. 별도로 겨누거나 날아가는 물체는 없다.",
        "built_space": "왼쪽 출입문 한 개, 천장 형광등 한 개, 뒤쪽 큰 창 한 개, 창 왼쪽 금속 선반 한 개, 벽과 천장의 배관이 보인다. 목재 식탁과 벤치도 참조의 재질과 배치에 잘 부합한다. 현우와 앰버는 식탁의 서로 반대편에 있다. 다만 천장과 넓은 배경, 식탁 위 반찬까지 포함되어 좁고 부드러운 배경만 남기는 클로즈업 요구에서 벗어난다.",
        "entities": "현우와 앰버만 보인다. 현우의 검은 머리와 회색 셔츠는 참조에 가깝고 젊은 한국계 남성으로 보인다. 앰버는 금발, 어린 체격, 남색 상의가 맞지만 뒷모습 위주라 얼굴 특징은 확인하기 어렵다. 검은 국그릇과 받침 각 한 개, 숟가락 한 개, 흰 반찬 접시 두 개, 금속 컵 한 개, 수저통과 휴지함이 보인다. 국그릇 내용물과 앰버의 식사 동작은 확인되지 않는다. 붕대와 숨긴 카드는 프레임 밖이며 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 식탁 쪽으로 허리를 깊게 굽히고 있으며 상체는 화면 오른쪽의 잘린 몸통으로 이어진다. 하체와 좌석은 보이지 않아 지지 자세 전체를 확인할 수 없지만, 불가능한 부유로 볼 근거는 없다. 앰버의 앉은 상체도 자연스럽다. 그릇과 컵, 반찬 접시, 휴지함, 수저통은 식탁에 놓여 있고 숟가락은 국그릇에 기대어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "앰버에게 시선을 꽂고 입을 크게 벌린 현우의 얼굴을 밀착해 담아, 분노한 얼굴의 클로즈업이라는 최우선 지시를 더 충실히 구현한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "앰버를 향한 고함과 장소의 구조는 잘 맞지만, 앰버의 상반신과 넓은 식탁까지 보여 주어 요구한 얼굴 클로즈업보다 구도가 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 두 눈은 화면 왼쪽 전경의 금발 앰버를 향한다. 입을 크게 벌리고 눈썹을 강하게 모아 상대에게 고함치는 순간으로 읽힌다. 앰버도 현우 쪽으로 머리를 돌리고 있으나 눈은 보이지 않는다. 겨누는 무기나 이동하는 물체는 없다.",
        "built_space": "현우 뒤로 금속 창틀의 창 한 구간, 회색 벽, 목재 식탁과 벤치 일부, 오른쪽 밥솥 한 대가 흐리게 보인다. 전경 식탁을 사이에 둔 두 사람의 위치는 자연스럽다. 참조의 낡은 컨테이너 식당 재질은 유지되지만, 창 왼쪽에 조리 선반이 있던 참조와 달리 여기서는 밥솥이 창 오른쪽에 보여 고정 설비 배치의 일치는 약하다. 배경을 좁게 남긴 얼굴 중심 구도는 지시에 가깝다.",
        "entities": "현우 한 명과 앰버의 흐린 머리 일부가 보이며 다른 손님은 없다. 현우의 헝클어진 검은 머리, 젊은 한국계 남성 외형, 회색 셔츠는 참조에 대체로 맞지만 얼굴은 조금 더 성숙하고 각져 보인다. 앰버는 금발과 밝은 피부만 확인되므로 얼굴 정체성과 정확한 나이는 판별하기 어렵다. 식탁에는 쌀이 보이는 국그릇 한 개, 숟가락 한 개, 젓가락 한 쌍과 금속 컵 한 개가 보인다. 다리 붕대와 신발 속 카드는 프레임 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 머리는 목과 어깨에 자연스럽게 이어지고 상체를 식탁 너머로 숙인 자세다. 골반과 좌석은 화면 밖이므로 직접적인 지지점은 확인할 수 없지만 몸이 공중에 떠 있는 묘사는 아니다. 국그릇과 컵은 식탁에 놓여 있고 숟가락은 그릇 안쪽과 가장자리에 걸쳐 지지된다. 얼굴과 입의 구조도 실제 고함 연기로 가능한 범위다."
       },
       {
        "label": "A",
        "direction": "현우는 왼쪽 전경에 앉은 앰버의 얼굴을 똑바로 바라보며 입을 크게 벌린다. 앰버 역시 현우 쪽을 향한다. 고함의 대상은 명확하지만 눈을 부릅뜬 느낌보다 눈썹을 찌푸린 분노가 더 두드러진다. 별도로 겨누거나 날아가는 물체는 없다.",
        "built_space": "왼쪽 출입문 한 개, 천장 형광등 한 개, 뒤쪽 큰 창 한 개, 창 왼쪽 금속 선반 한 개, 벽과 천장의 배관이 보인다. 목재 식탁과 벤치도 참조의 재질과 배치에 잘 부합한다. 현우와 앰버는 식탁의 서로 반대편에 있다. 다만 천장과 넓은 배경, 식탁 위 반찬까지 포함되어 좁고 부드러운 배경만 남기는 클로즈업 요구에서 벗어난다.",
        "entities": "현우와 앰버만 보인다. 현우의 검은 머리와 회색 셔츠는 참조에 가깝고 젊은 한국계 남성으로 보인다. 앰버는 금발, 어린 체격, 남색 상의가 맞지만 뒷모습 위주라 얼굴 특징은 확인하기 어렵다. 검은 국그릇과 받침 각 한 개, 숟가락 한 개, 흰 반찬 접시 두 개, 금속 컵 한 개, 수저통과 휴지함이 보인다. 국그릇 내용물과 앰버의 식사 동작은 확인되지 않는다. 붕대와 숨긴 카드는 프레임 밖이며 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 식탁 쪽으로 허리를 깊게 굽히고 있으며 상체는 화면 오른쪽의 잘린 몸통으로 이어진다. 하체와 좌석은 보이지 않아 지지 자세 전체를 확인할 수 없지만, 불가능한 부유로 볼 근거는 없다. 앰버의 앉은 상체도 자연스럽다. 그릇과 컵, 반찬 접시, 휴지함, 수저통은 식탁에 놓여 있고 숟가락은 국그릇에 기대어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.464,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.464,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1464
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 프레이밍과 얕은 심도를 정확히 구현했으며, 눈을 부릅뜨고 고함치는 현우의 분노한 표정이 지시문에 완벽히 부합합니다."
   },
   {
    "label": "A",
    "score": 1464,
    "verdict_ko": "상황과 공간적 배경은 잘 구현했으나 샷 크기가 지시된 클로즈업보다는 미디엄 샷에 가깝습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B01.png",
    "asset_id": "78695b5c-f8e2-4047-9612-b490195deea7",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-f0fa-7bda-8533-d16a24bffef1",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh10__bgfirst_bg.png",
   "bg_asset_id": "683d25c0-58aa-44f7-a776-222b7b276425",
   "bg_record_key": "S51sh10::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C03"
  ]
 },
 "S51sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:21:16.975759+00:00",
  "fingerprint": "637658d458cab300a7131c0eec4483a24feb4cedd32f7779d65df99cb4c0efbf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S51sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S51sh10_sel.png",
  "source_sha256": "c7ff208377bf96a87f85a560b6c08cc280c801a2e2f76d1714b5f9eb53f35f9b",
  "file": "S51sh10_cine.png",
  "staged_sha256": "a70e4726c1ad4bcd28e62f85d4793102ebda461d4c95143d76ebb5e15f746552",
  "latency_ms": 12762
 },
 "S51sh14::signage": {
  "fp": "64e7725cce9cf68e",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "삐뚤빼뚤한 글씨가 적힌 낡은 종이"
   }
  ],
  "dropped": []
 },
 "S51sh14": {
  "input_fingerprint": "a6730808414bbc49",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버가 쥔 삐뚤빼뚤한 글씨가 적힌 낡은 종이를 뚫어져라 내려다보는 현우의 시점 쇼트.\n\nLOCATION (lock): In the temporary sleeping area at the auto-repair premises, in morning light as the farewell note is examined. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리's farewell letter (Worn paper held by 앰버, bearing uneven handwriting) — The written face is visible to the camera and reads: 고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리; used as Central reading focus, surrounded by the holder's hands and partial torso rather than enlarged beyond natural scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the handwriting legible with restrained contrast and no perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's farewell note says that he is going to Haenam and will remember Amber as his friend. The camper is still at the repair shop; Charlie has departed alone. 앰버: She is presenting the farewell note.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버, 현우 right now, so 앰버, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버가 쥔 삐뚤빼뚤한 글씨가 적힌 낡은 종이를 뚫어져라 내려다보는 현우의 시점 쇼트.\n\nLOCATION (lock): In the temporary sleeping area at the auto-repair premises, in morning light as the farewell note is examined. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리's farewell letter (Worn paper held by 앰버, bearing uneven handwriting) — The written face is visible to the camera and reads: 고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리; used as Central reading focus, surrounded by the holder's hands and partial torso rather than enlarged beyond natural scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the handwriting legible with restrained contrast and no perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's farewell note says that he is going to Haenam and will remember Amber as his friend. The camper is still at the repair shop; Charlie has departed alone. 앰버: She is presenting the farewell note.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버, 현우 right now, so 앰버, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버가 쥔 삐뚤빼뚤한 글씨가 적힌 낡은 종이를 뚫어져라 내려다보는 현우의 시점 쇼트.\n\nLOCATION (lock): In the temporary sleeping area at the auto-repair premises, in morning light as the farewell note is examined. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리's farewell letter (Worn paper held by 앰버, bearing uneven handwriting) — The written face is visible to the camera and reads: 고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리; used as Central reading focus, surrounded by the holder's hands and partial torso rather than enlarged beyond natural scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the handwriting legible with restrained contrast and no perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's farewell note says that he is going to Haenam and will remember Amber as his friend. The camper is still at the repair shop; Charlie has departed alone. 앰버: She is presenting the farewell note.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버, 현우 right now, so 앰버, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 앰버의 등 뒤에서 그녀가 들고 있는 편지를 바라보고 있으며, 지시된 현우의 시점이 아님.",
    "built_space": "담요와 깡통들이 있는 공간으로, 자동차 정비소의 특징적인 구조물은 뚜렷하게 보이지 않음.",
    "entities": "앰버의 뒷모습과 손, 프롬프트가 요구한 정확한 문구가 적힌 낡은 종이가 보임. 현우는 화면에 등장하지 않음.",
    "hard_violations": [
     "[gemini-pro] 현우의 시점 쇼트를 지시한 프롬프트와 달리 카메라가 앰버의 등 뒤에 위치하여 카메라 시점을 위반함",
     "[gemini-pro] 앰버가 현우에게 편지를 제시(presenting)하는 액션이 아니라 스스로 읽는 모습으로 연출됨",
     "[gpt-high] 판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
    ],
    "physics": "앰버의 두 손이 편지지의 양옆을 쥐고 안정적으로 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 왼쪽 전경에 위치한 현우의 시점(어깨 너머)에서 편지를 정면으로 들고 있는 앰버를 내려다보고 있음.",
    "built_space": "야전 침대와 타이어가 배치된 자동차 정비소 내부의 임시 수면 공간.",
    "entities": "현우의 뒷모습/어깨 일부, 레퍼런스와 일치하는 외모 및 복장의 앰버, 지시된 문구가 적힌 종이(줄바꿈에 '게' 단어가 한 번 중복됨)가 모두 존재함.",
    "hard_violations": [
     "[gpt-high] 현우의 주관 시점이어야 하는 화면에 현우로 보이는 성인 남성의 머리와 어깨를 삽입하여, 이 숏에서 허용한 인물 노출과 시점 배치를 위반했다.",
     "[gpt-high] 판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
    ],
    "physics": "앰버가 종이의 양끝을 두 손으로 쥐고 현우를 향해 들어 올려 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우의 시점 쇼트와 편지를 보여주는(presenting) 액션을 무시하고 앰버의 등 뒤에서 촬영하여 핵심적인 카메라 위치 및 연출 지시를 완전히 위반했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 시점에서 앰버가 편지를 제시하는 구도를 정확히 구현했으며, 편지의 사소한 오탈자를 제외하면 캐릭터 레퍼런스와 공간적 배경까지 훌륭하게 반영했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 앰버의 등 뒤에서 그녀가 들고 있는 편지를 바라보고 있으며, 지시된 현우의 시점이 아님.",
        "built_space": "담요와 깡통들이 있는 공간으로, 자동차 정비소의 특징적인 구조물은 뚜렷하게 보이지 않음.",
        "entities": "앰버의 뒷모습과 손, 프롬프트가 요구한 정확한 문구가 적힌 낡은 종이가 보임. 현우는 화면에 등장하지 않음.",
        "hard_violations": [
         "현우의 시점 쇼트를 지시한 프롬프트와 달리 카메라가 앰버의 등 뒤에 위치하여 카메라 시점을 위반함",
         "앰버가 현우에게 편지를 제시(presenting)하는 액션이 아니라 스스로 읽는 모습으로 연출됨"
        ],
        "physics": "앰버의 두 손이 편지지의 양옆을 쥐고 안정적으로 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 왼쪽 전경에 위치한 현우의 시점(어깨 너머)에서 편지를 정면으로 들고 있는 앰버를 내려다보고 있음.",
        "built_space": "야전 침대와 타이어가 배치된 자동차 정비소 내부의 임시 수면 공간.",
        "entities": "현우의 뒷모습/어깨 일부, 레퍼런스와 일치하는 외모 및 복장의 앰버, 지시된 문구가 적힌 종이(줄바꿈에 '게' 단어가 한 번 중복됨)가 모두 존재함.",
        "hard_violations": [],
        "physics": "앰버가 종이의 양끝을 두 손으로 쥐고 현우를 향해 들어 올려 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우의 시점 쇼트와 편지를 보여주는(presenting) 액션을 무시하고 앰버의 등 뒤에서 촬영하여 핵심적인 카메라 위치 및 연출 지시를 완전히 위반했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 시점에서 앰버가 편지를 제시하는 구도를 정확히 구현했으며, 편지의 사소한 오탈자를 제외하면 캐릭터 레퍼런스와 공간적 배경까지 훌륭하게 반영했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 앰버의 등 뒤에서 그녀가 들고 있는 편지를 바라보고 있으며, 지시된 현우의 시점이 아님.",
        "built_space": "담요와 깡통들이 있는 공간으로, 자동차 정비소의 특징적인 구조물은 뚜렷하게 보이지 않음.",
        "entities": "앰버의 뒷모습과 손, 프롬프트가 요구한 정확한 문구가 적힌 낡은 종이가 보임. 현우는 화면에 등장하지 않음.",
        "hard_violations": [
         "현우의 시점 쇼트를 지시한 프롬프트와 달리 카메라가 앰버의 등 뒤에 위치하여 카메라 시점을 위반함",
         "앰버가 현우에게 편지를 제시(presenting)하는 액션이 아니라 스스로 읽는 모습으로 연출됨"
        ],
        "physics": "앰버의 두 손이 편지지의 양옆을 쥐고 안정적으로 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 왼쪽 전경에 위치한 현우의 시점(어깨 너머)에서 편지를 정면으로 들고 있는 앰버를 내려다보고 있음.",
        "built_space": "야전 침대와 타이어가 배치된 자동차 정비소 내부의 임시 수면 공간.",
        "entities": "현우의 뒷모습/어깨 일부, 레퍼런스와 일치하는 외모 및 복장의 앰버, 지시된 문구가 적힌 종이(줄바꿈에 '게' 단어가 한 번 중복됨)가 모두 존재함.",
        "hard_violations": [],
        "physics": "앰버가 종이의 양끝을 두 손으로 쥐고 현우를 향해 들어 올려 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우의 시점 클로즈업 대신 현우의 머리와 어깨가 들어온 대면 구도이며, 금지된 판독 가능한 글씨도 노출된다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "낡은 편지와 앰버의 손·부분 상체에 집중한 내려다보는 클로즈업은 더 충실하지만, 글씨를 읽을 수 없게 하라는 최종 지시를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 편지 뒤에 서서 맞은편 현우 쪽을 약간 내려다본다. 편지의 글씨 면은 현우와 카메라 쪽을 향하므로 제시 방향은 맞지만, 카메라가 현우의 눈 위치가 아니라 머리와 어깨 뒤에 있어 요구된 주관 시점은 아니다. 현우의 눈은 보이지 않아 실제 응시 대상은 확인할 수 없다.",
        "built_space": "뒤에는 간이침대 한 개, 그 위의 베개와 접힌 침구, 상단 창문 한 개, 오른쪽 바퀴 한 개와 기계 부품이 보인다. 앰버는 침대 앞에 있고 현우의 일부가 왼쪽 전경을 차지한다. 낡은 벽과 주간광은 장소 분위기에 부합하지만, 이전 사진의 포스터와 배관 등 동일 장소를 확인할 고정 특징은 보이지 않는다. 반사나 중복 설비 문제는 없다.",
        "entities": "낡고 접힌 편지 한 장을 금발 여자아이가 두 손으로 쥐고 있다. 아이의 연령대, 밝은 피부, 둥근 얼굴, 남색 상의와 갈색 작업복은 앰버 참고와 대체로 맞지만 머리 위 장비는 없다. 왼쪽에는 검은 머리와 회색 셔츠를 입은 성인 남성의 머리·어깨가 추가로 보인다. 편지에는 요청된 작별 문구 대부분이 읽히지만, 마지막의 모든 글씨를 판독 불가능하게 하라는 지시와 충돌한다.",
        "hard_violations": [
         "현우의 주관 시점이어야 하는 화면에 현우로 보이는 성인 남성의 머리와 어깨를 삽입하여, 이 숏에서 허용한 인물 노출과 시점 배치를 위반했다.",
         "판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
        ],
        "physics": "편지 양쪽 가장자리를 앰버의 손가락이 실제로 잡고 있으며 종이의 접힘과 처짐도 가능한 형태다. 침구는 침대에 놓여 있고 바퀴는 뒤쪽 구조물에 기대어 있다. 인물의 하체는 화면 밖이라 지면 접촉은 확인할 수 없지만, 공중에 뜬 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "앰버의 얼굴은 왼쪽 아래 편지 쪽으로 향하고 글씨 면도 앰버의 눈 쪽을 향한다. 카메라는 앰버의 옆뒤 위에서 같은 면을 내려다보므로 현우가 곁에서 편지를 들여다보는 시점으로 성립한다. 다만 맞은편 현우에게 편지를 내보이기보다는 앰버 자신이 읽는 동작에 가깝다.",
        "built_space": "화면 대부분은 회색 담요가 깔린 잠자리 한 곳이며, 가장자리에 녹색 프레임 일부가 보인다. 뒤쪽 작업면에는 노란 원통형 용기 두 개, 은색 열린 통 한 개, 어두운 통들과 큰 플라스틱 용기가 놓여 있다. 앰버는 잠자리 가장자리에 앉아 편지를 무릎 위로 들고 있는 것으로 보인다. 임시 숙박·정비 공간의 재질은 자연스럽지만 이전 사진의 고정 설비와 직접 대조할 부분은 부족하다. 불가능한 반사는 없다.",
        "entities": "낡은 종이 한 장, 이를 쥔 두 손, 앰버의 금발과 얼굴 일부, 남색 상의와 갈색 멜빵 작업복이 보인다. 보이는 체격과 피부는 어린 앰버와 대체로 맞으며, 얼굴이 대부분 제외되어 정확한 얼굴 정체성은 판단하기 어렵다. 다른 인물은 없다. 작별 편지의 내용은 식별되지만 글씨는 비교적 정돈되어 있고, 무엇보다 최종 비가독성 지시를 따르지 않는다.",
        "hard_violations": [
         "판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
        ],
        "physics": "양손이 종이의 좌우 가장자리를 엄지와 나머지 손가락으로 집어 지지한다. 팔과 손목의 연결, 종이의 기울기와 접힌 형태는 자연스럽다. 아래쪽 작업복 차림의 다리와 침구가 앉은 자세를 뒷받침하며, 배경 용기들은 작업면 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우의 시점 클로즈업 대신 현우의 머리와 어깨가 들어온 대면 구도이며, 금지된 판독 가능한 글씨도 노출된다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "낡은 편지와 앰버의 손·부분 상체에 집중한 내려다보는 클로즈업은 더 충실하지만, 글씨를 읽을 수 없게 하라는 최종 지시를 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 편지 뒤에 서서 맞은편 현우 쪽을 약간 내려다본다. 편지의 글씨 면은 현우와 카메라 쪽을 향하므로 제시 방향은 맞지만, 카메라가 현우의 눈 위치가 아니라 머리와 어깨 뒤에 있어 요구된 주관 시점은 아니다. 현우의 눈은 보이지 않아 실제 응시 대상은 확인할 수 없다.",
        "built_space": "뒤에는 간이침대 한 개, 그 위의 베개와 접힌 침구, 상단 창문 한 개, 오른쪽 바퀴 한 개와 기계 부품이 보인다. 앰버는 침대 앞에 있고 현우의 일부가 왼쪽 전경을 차지한다. 낡은 벽과 주간광은 장소 분위기에 부합하지만, 이전 사진의 포스터와 배관 등 동일 장소를 확인할 고정 특징은 보이지 않는다. 반사나 중복 설비 문제는 없다.",
        "entities": "낡고 접힌 편지 한 장을 금발 여자아이가 두 손으로 쥐고 있다. 아이의 연령대, 밝은 피부, 둥근 얼굴, 남색 상의와 갈색 작업복은 앰버 참고와 대체로 맞지만 머리 위 장비는 없다. 왼쪽에는 검은 머리와 회색 셔츠를 입은 성인 남성의 머리·어깨가 추가로 보인다. 편지에는 요청된 작별 문구 대부분이 읽히지만, 마지막의 모든 글씨를 판독 불가능하게 하라는 지시와 충돌한다.",
        "hard_violations": [
         "현우의 주관 시점이어야 하는 화면에 현우로 보이는 성인 남성의 머리와 어깨를 삽입하여, 이 숏에서 허용한 인물 노출과 시점 배치를 위반했다.",
         "판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
        ],
        "physics": "편지 양쪽 가장자리를 앰버의 손가락이 실제로 잡고 있으며 종이의 접힘과 처짐도 가능한 형태다. 침구는 침대에 놓여 있고 바퀴는 뒤쪽 구조물에 기대어 있다. 인물의 하체는 화면 밖이라 지면 접촉은 확인할 수 없지만, 공중에 뜬 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "앰버의 얼굴은 왼쪽 아래 편지 쪽으로 향하고 글씨 면도 앰버의 눈 쪽을 향한다. 카메라는 앰버의 옆뒤 위에서 같은 면을 내려다보므로 현우가 곁에서 편지를 들여다보는 시점으로 성립한다. 다만 맞은편 현우에게 편지를 내보이기보다는 앰버 자신이 읽는 동작에 가깝다.",
        "built_space": "화면 대부분은 회색 담요가 깔린 잠자리 한 곳이며, 가장자리에 녹색 프레임 일부가 보인다. 뒤쪽 작업면에는 노란 원통형 용기 두 개, 은색 열린 통 한 개, 어두운 통들과 큰 플라스틱 용기가 놓여 있다. 앰버는 잠자리 가장자리에 앉아 편지를 무릎 위로 들고 있는 것으로 보인다. 임시 숙박·정비 공간의 재질은 자연스럽지만 이전 사진의 고정 설비와 직접 대조할 부분은 부족하다. 불가능한 반사는 없다.",
        "entities": "낡은 종이 한 장, 이를 쥔 두 손, 앰버의 금발과 얼굴 일부, 남색 상의와 갈색 멜빵 작업복이 보인다. 보이는 체격과 피부는 어린 앰버와 대체로 맞으며, 얼굴이 대부분 제외되어 정확한 얼굴 정체성은 판단하기 어렵다. 다른 인물은 없다. 작별 편지의 내용은 식별되지만 글씨는 비교적 정돈되어 있고, 무엇보다 최종 비가독성 지시를 따르지 않는다.",
        "hard_violations": [
         "판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
        ],
        "physics": "양손이 종이의 좌우 가장자리를 엄지와 나머지 손가락으로 집어 지지한다. 팔과 손목의 연결, 종이의 기울기와 접힌 형태는 자연스럽다. 아래쪽 작업복 차림의 다리와 침구가 앉은 자세를 뒷받침하며, 배경 용기들은 작업면 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.333,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.083,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 현우의 시점 쇼트를 지시한 프롬프트와 달리 카메라가 앰버의 등 뒤에 위치하여 카메라 시점을 위반함",
     "[gemini-pro] 앰버가 현우에게 편지를 제시(presenting)하는 액션이 아니라 스스로 읽는 모습으로 연출됨",
     "[gpt-high] 판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
    ],
    "B": [
     "[gpt-high] 현우의 주관 시점이어야 하는 화면에 현우로 보이는 성인 남성의 머리와 어깨를 삽입하여, 이 숏에서 허용한 인물 노출과 시점 배치를 위반했다.",
     "[gpt-high] 판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1083,
   "B": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1083,
    "verdict_ko": "현우의 시점 쇼트와 편지를 보여주는(presenting) 액션을 무시하고 앰버의 등 뒤에서 촬영하여 핵심적인 카메라 위치 및 연출 지시를 완전히 위반했습니다.  ★위반: [gemini-pro] 현우의 시점 쇼트를 지시한 프롬프트와 달리 카메라가 앰버의 등 뒤에 위치하여 카메라 시점을 위반함 / [gemini-pro] 앰버가 현우에게 편지를 제시(presenting)하는 액션이 아니라 스스로 읽는 모습으로 연출됨 / [gpt-high] 판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "현우의 시점에서 앰버가 편지를 제시하는 구도를 정확히 구현했으며, 편지의 사소한 오탈자를 제외하면 캐릭터 레퍼런스와 공간적 배경까지 훌륭하게 반영했습니다.  ★위반: [gpt-high] 현우의 주관 시점이어야 하는 화면에 현우로 보이는 성인 남성의 머리와 어깨를 삽입하여, 이 숏에서 허용한 인물 노출과 시점 배치를 위반했다. / [gpt-high] 판독 가능한 글씨를 금지한 최종 지시에도 편지의 한글 문장이 선명하게 노출된다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh10_sel.png",
    "asset_id": "35691702-920c-4c9f-95d0-a4377b405d22",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-f43d-7e0d-9f7b-b449d7ee4169",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S50sh10"
  }
 },
 "S51sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:22:28.854024+00:00",
  "fingerprint": "25aa88defb22438a2755fda0fa62cb4832e6cf82ac4ad98f9104557771fe02d7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S51sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S51sh14_sel.png",
  "source_sha256": "fb893ac5c041ab95358c122085badc06f3d933dd8e87e0a315a8815e6afbc1c3",
  "file": "S51sh14_cine.png",
  "staged_sha256": "2a50b59cd52dae02a54c22f5200d8668647442ef126972c5778b19ea031b02e2",
  "latency_ms": 9788
 },
 "S51sh19::signage": {
  "fp": "077a34b083b69dc3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S51sh19::bgfirst_bg": {
  "input_fingerprint": "56bfb11813fc2114",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh19__bgfirst_bg.png",
  "asset_id": "3517c78d-2559-4fa2-9d40-d3795bb05053",
  "input_asset_ids": [
   "9849057f-ede4-459a-804c-8659aaf33808",
   "2d5b2b8c-9d71-4f7f-acce-b7f26107156a"
  ]
 },
 "S51sh19": {
  "input_fingerprint": "c724b45416a3e1fe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is operational again and has the luggage loaded for departure. Additional food and Amber's medicine have been placed aboard. 앰버: She is aboard the departing camper. 라울: He is aboard the departing camper, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is operational again and has the luggage loaded for departure. Additional food and Amber's medicine have been placed aboard. 앰버: She is aboard the departing camper. 라울: He is aboard the departing camper, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is operational again and has the luggage loaded for departure. Additional food and Amber's medicine have been placed aboard. 앰버: She is aboard the departing camper. 라울: He is aboard the departing camper, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh19__bgfirst_bg.png",
     "asset_id": "3517c78d-2559-4fa2-9d40-d3795bb05053",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S51sh19.png",
     "asset_id": "9849057f-ede4-459a-804c-8659aaf33808",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B02.png",
     "asset_id": "2d5b2b8c-9d71-4f7f-acce-b7f26107156a",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "앰버와 라울 모두 우측 창밖의 보이지 않는 타겟(은영과 태진)을 향해 시선을 두고 양손을 들고 있음.",
    "built_space": "캠핑카 내부에서 사선 구도(oblique view)로 촬영됨. 우측 창문 밖으로 레퍼런스와 동일한 자동차 정비소(파란색 외벽)가 보임.",
    "entities": "앰버는 금발 등 외형은 일치하나 멜빵바지 없이 파란 티셔츠만 입음. 라울은 꽁지머리와 회색 티셔츠 등 외형과 복장이 일치함.",
    "hard_violations": [],
    "physics": "두 인물 모두 좌석에 안정적으로 앉아 안전벨트를 매고 양팔을 허공으로 뻗고 있음."
   },
   {
    "label": "B",
    "direction": "앰버와 라울 모두 창밖이 아닌 카메라 렌즈를 정면으로 바라보며 손을 흔들고 있음.",
    "built_space": "캠핑카 내부 중앙 통로를 정면으로 바라보는 앵글로, 지시된 사선 구도를 위반함. 창밖 풍경은 레퍼런스와 다름.",
    "entities": "앰버는 파란 티셔츠와 멜빵바지를, 라울은 회색 티셔츠를 착용하여 두 인물 모두 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "좌석에 앉아 양팔을 뻗고 있으며 신체와 닿은 표면 지지가 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 사선 구도와 창밖을 향한 시선을 정확히 연출했으며 창밖의 정비소 배경도 완벽히 구현했으나, 앰버의 멜빵바지가 누락됨."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물들의 복장은 레퍼런스와 일치하나, 지시된 사선 구도를 무시하고 정면 렌즈를 응시하여 샷 텍스트의 핵심 연출을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 라울 모두 우측 창밖의 보이지 않는 타겟(은영과 태진)을 향해 시선을 두고 양손을 들고 있음.",
        "built_space": "캠핑카 내부에서 사선 구도(oblique view)로 촬영됨. 우측 창문 밖으로 레퍼런스와 동일한 자동차 정비소(파란색 외벽)가 보임.",
        "entities": "앰버는 금발 등 외형은 일치하나 멜빵바지 없이 파란 티셔츠만 입음. 라울은 꽁지머리와 회색 티셔츠 등 외형과 복장이 일치함.",
        "hard_violations": [],
        "physics": "두 인물 모두 좌석에 안정적으로 앉아 안전벨트를 매고 양팔을 허공으로 뻗고 있음."
       },
       {
        "label": "B",
        "direction": "앰버와 라울 모두 창밖이 아닌 카메라 렌즈를 정면으로 바라보며 손을 흔들고 있음.",
        "built_space": "캠핑카 내부 중앙 통로를 정면으로 바라보는 앵글로, 지시된 사선 구도를 위반함. 창밖 풍경은 레퍼런스와 다름.",
        "entities": "앰버는 파란 티셔츠와 멜빵바지를, 라울은 회색 티셔츠를 착용하여 두 인물 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "좌석에 앉아 양팔을 뻗고 있으며 신체와 닿은 표면 지지가 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 사선 구도와 창밖을 향한 시선을 정확히 연출했으며 창밖의 정비소 배경도 완벽히 구현했으나, 앰버의 멜빵바지가 누락됨."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물들의 복장은 레퍼런스와 일치하나, 지시된 사선 구도를 무시하고 정면 렌즈를 응시하여 샷 텍스트의 핵심 연출을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 라울 모두 우측 창밖의 보이지 않는 타겟(은영과 태진)을 향해 시선을 두고 양손을 들고 있음.",
        "built_space": "캠핑카 내부에서 사선 구도(oblique view)로 촬영됨. 우측 창문 밖으로 레퍼런스와 동일한 자동차 정비소(파란색 외벽)가 보임.",
        "entities": "앰버는 금발 등 외형은 일치하나 멜빵바지 없이 파란 티셔츠만 입음. 라울은 꽁지머리와 회색 티셔츠 등 외형과 복장이 일치함.",
        "hard_violations": [],
        "physics": "두 인물 모두 좌석에 안정적으로 앉아 안전벨트를 매고 양팔을 허공으로 뻗고 있음."
       },
       {
        "label": "B",
        "direction": "앰버와 라울 모두 창밖이 아닌 카메라 렌즈를 정면으로 바라보며 손을 흔들고 있음.",
        "built_space": "캠핑카 내부 중앙 통로를 정면으로 바라보는 앵글로, 지시된 사선 구도를 위반함. 창밖 풍경은 레퍼런스와 다름.",
        "entities": "앰버는 파란 티셔츠와 멜빵바지를, 라울은 회색 티셔츠를 착용하여 두 인물 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "좌석에 앉아 양팔을 뻗고 있으며 신체와 닿은 표면 지지가 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "두 아이의 양손 높이 들기와 의상은 충실하지만, 창밖이 아닌 차내 카메라를 정면으로 바라봐 은영과 태진을 향한 작별의 방향과 사선 실내 구도를 놓쳤다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "창밖을 향한 시선과 환한 웃음, 사선 상반신 구도가 작별 순간에 더 충실하나, 라울의 한 손이 낮고 앰버의 갈색 작업복이 빠졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 라울 모두 양손을 머리 위로 올리고 손바닥과 얼굴을 차내 카메라 쪽으로 향한다. 시선도 거의 렌즈를 향하며, 정비소가 보이는 왼쪽 창밖을 바라보지 않는다. 따라서 화면 밖 은영과 태진에게 보내는 인사보다는 차내 관객에게 하는 인사로 읽힌다.",
        "built_space": "중앙 통로 양쪽에 높은 등받이 좌석이 하나씩 있고 아이들은 각 좌석에 앉아 있다. 전경에는 탁자 상판 두 개, 양옆에는 큰 창이 하나씩, 뒤쪽에는 작은 창과 짐 공간이 보인다. 상부 목재 수납장, 좌우 독서등, 천창도 있다. 좌석과 신체의 배치는 가능하지만 카메라는 통로를 따라 거의 정면으로 놓여 요구한 사선 구도가 아니다. 왼쪽 창밖 골강판 건물은 참고 장소의 재질과 유사하나 동일한 외관인지는 제한적으로만 확인된다.",
        "entities": "보이는 사람은 어린 여자아이 앰버와 어린 남자아이 라울 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 티셔츠, 갈색 멜빵 작업복과 공구 벨트는 참고와 대체로 맞지만 머리 위 보호구는 없다. 라울의 피부색, 어린 얼굴, 뒤로 묶은 머리와 낡은 회색 티셔츠도 참고와 대체로 맞는다. 혼혈 배경 자체는 외관만으로 확정할 수 없다. 뒤쪽 여행 가방과 용기들은 보이지만 식량과 약의 내용물, 라울의 치료 상태는 확인되지 않는다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "두 아이의 하체는 일부 가려졌지만 엉덩이가 좌석에 놓이고 등 뒤에 등받이가 있어 앉은 자세의 지지가 성립한다. 네 팔은 어깨와 팔꿈치에서 자연스럽게 이어져 위로 올라가며, 손에 든 물건은 없다. 짐은 바닥과 후방 공간에 놓여 있고 용기들은 선반에 받쳐져 있다. 공중에 지지 없이 떠 있는 신체나 물건은 보이지 않는다. 차량의 출발 운동은 정지 화면에서 명확히 확인되지 않는다."
       },
       {
        "label": "B",
        "direction": "두 아이의 얼굴과 시선이 함께 오른쪽 창밖, 참고 장소와 닮은 건물이 있는 방향으로 향한다. 은영과 태진은 화면 밖이지만 인사의 목표가 차 밖에 있다는 관계는 명확하다. 앰버는 양손을 머리 위로 올렸고, 라울은 창가 쪽 손만 높이 올린 채 다른 손은 어깨 부근에서 펼쳐 보여 '양손을 높이'라는 지시는 부분적으로만 충족한다.",
        "built_space": "카메라는 아이들의 앞쪽 실내에서 창가를 비스듬히 본다. 두 아이는 하나의 넓은 벤치 좌석에 나란히 앉고 각자 뒤에 등받이와 머리받침이 있다. 오른쪽에는 큰 창과 그 뒤의 작은 창, 아래쪽 손잡이 하나가 보이며, 후방에는 빈 좌석들과 수납장, 천장에는 사각 천창 하나가 보인다. 창밖의 골강판 외벽, 창틀, 실외기와 가스통, 자갈 마당은 참고 장소와 잘 연결된다. 내부가 일부 넓게 드러나지만 인물의 상체와 손이 중심인 사선 미디엄 숏이다.",
        "entities": "보이는 사람은 앰버와 라울 두 명뿐이다. 앰버는 참고와 유사한 금발과 둥근 어린 얼굴, 남색 티셔츠를 갖췄지만 상체에 있어야 할 갈색 멜빵 작업복과 머리 위 보호구가 없다. 라울은 어린 얼굴, 짙은 피부, 뒤로 묶은 곱슬머리와 낡은 회색 티셔츠로 참고의 주요 특징을 유지한다. 두 아이 모두 약 10세로 읽히며 혼혈 배경의 정확한 구성은 외관만으로 확정할 수 없다. 안전벨트가 보이고 후방에 작은 가방이 있지만 식량, 약, 치료 흔적은 식별되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 아이의 엉덩이와 허벅지는 벤치 좌판에 지지되고 등은 등받이 쪽에 놓이며 안전벨트가 몸통을 지난다. 창밖으로 얼굴을 돌리고 팔을 들어 인사하는 자세는 앉은 상태에서 가능하다. 라울의 낮은 손도 팔꿈치를 굽혀 몸 옆에 든 자연스러운 자세이며, 해부학적 불가능이 아니라 동작 지시와의 차이다. 지지 없이 뜬 신체나 물건은 없다. 출발 중인 차량이라는 상태는 안전벨트와 탑승 상황으로 암시되지만 운동 자체는 뚜렷하지 않다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "두 아이의 양손 높이 들기와 의상은 충실하지만, 창밖이 아닌 차내 카메라를 정면으로 바라봐 은영과 태진을 향한 작별의 방향과 사선 실내 구도를 놓쳤다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "창밖을 향한 시선과 환한 웃음, 사선 상반신 구도가 작별 순간에 더 충실하나, 라울의 한 손이 낮고 앰버의 갈색 작업복이 빠졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버와 라울 모두 양손을 머리 위로 올리고 손바닥과 얼굴을 차내 카메라 쪽으로 향한다. 시선도 거의 렌즈를 향하며, 정비소가 보이는 왼쪽 창밖을 바라보지 않는다. 따라서 화면 밖 은영과 태진에게 보내는 인사보다는 차내 관객에게 하는 인사로 읽힌다.",
        "built_space": "중앙 통로 양쪽에 높은 등받이 좌석이 하나씩 있고 아이들은 각 좌석에 앉아 있다. 전경에는 탁자 상판 두 개, 양옆에는 큰 창이 하나씩, 뒤쪽에는 작은 창과 짐 공간이 보인다. 상부 목재 수납장, 좌우 독서등, 천창도 있다. 좌석과 신체의 배치는 가능하지만 카메라는 통로를 따라 거의 정면으로 놓여 요구한 사선 구도가 아니다. 왼쪽 창밖 골강판 건물은 참고 장소의 재질과 유사하나 동일한 외관인지는 제한적으로만 확인된다.",
        "entities": "보이는 사람은 어린 여자아이 앰버와 어린 남자아이 라울 두 명뿐이다. 앰버의 금발, 둥근 얼굴, 남색 티셔츠, 갈색 멜빵 작업복과 공구 벨트는 참고와 대체로 맞지만 머리 위 보호구는 없다. 라울의 피부색, 어린 얼굴, 뒤로 묶은 머리와 낡은 회색 티셔츠도 참고와 대체로 맞는다. 혼혈 배경 자체는 외관만으로 확정할 수 없다. 뒤쪽 여행 가방과 용기들은 보이지만 식량과 약의 내용물, 라울의 치료 상태는 확인되지 않는다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "두 아이의 하체는 일부 가려졌지만 엉덩이가 좌석에 놓이고 등 뒤에 등받이가 있어 앉은 자세의 지지가 성립한다. 네 팔은 어깨와 팔꿈치에서 자연스럽게 이어져 위로 올라가며, 손에 든 물건은 없다. 짐은 바닥과 후방 공간에 놓여 있고 용기들은 선반에 받쳐져 있다. 공중에 지지 없이 떠 있는 신체나 물건은 보이지 않는다. 차량의 출발 운동은 정지 화면에서 명확히 확인되지 않는다."
       },
       {
        "label": "A",
        "direction": "두 아이의 얼굴과 시선이 함께 오른쪽 창밖, 참고 장소와 닮은 건물이 있는 방향으로 향한다. 은영과 태진은 화면 밖이지만 인사의 목표가 차 밖에 있다는 관계는 명확하다. 앰버는 양손을 머리 위로 올렸고, 라울은 창가 쪽 손만 높이 올린 채 다른 손은 어깨 부근에서 펼쳐 보여 '양손을 높이'라는 지시는 부분적으로만 충족한다.",
        "built_space": "카메라는 아이들의 앞쪽 실내에서 창가를 비스듬히 본다. 두 아이는 하나의 넓은 벤치 좌석에 나란히 앉고 각자 뒤에 등받이와 머리받침이 있다. 오른쪽에는 큰 창과 그 뒤의 작은 창, 아래쪽 손잡이 하나가 보이며, 후방에는 빈 좌석들과 수납장, 천장에는 사각 천창 하나가 보인다. 창밖의 골강판 외벽, 창틀, 실외기와 가스통, 자갈 마당은 참고 장소와 잘 연결된다. 내부가 일부 넓게 드러나지만 인물의 상체와 손이 중심인 사선 미디엄 숏이다.",
        "entities": "보이는 사람은 앰버와 라울 두 명뿐이다. 앰버는 참고와 유사한 금발과 둥근 어린 얼굴, 남색 티셔츠를 갖췄지만 상체에 있어야 할 갈색 멜빵 작업복과 머리 위 보호구가 없다. 라울은 어린 얼굴, 짙은 피부, 뒤로 묶은 곱슬머리와 낡은 회색 티셔츠로 참고의 주요 특징을 유지한다. 두 아이 모두 약 10세로 읽히며 혼혈 배경의 정확한 구성은 외관만으로 확정할 수 없다. 안전벨트가 보이고 후방에 작은 가방이 있지만 식량, 약, 치료 흔적은 식별되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 아이의 엉덩이와 허벅지는 벤치 좌판에 지지되고 등은 등받이 쪽에 놓이며 안전벨트가 몸통을 지난다. 창밖으로 얼굴을 돌리고 팔을 들어 인사하는 자세는 앉은 상태에서 가능하다. 라울의 낮은 손도 팔꿈치를 굽혀 몸 옆에 든 자연스러운 자세이며, 해부학적 불가능이 아니라 동작 지시와의 차이다. 지지 없이 뜬 신체나 물건은 없다. 출발 중인 차량이라는 상태는 안전벨트와 탑승 상황으로 암시되지만 운동 자체는 뚜렷하지 않다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.196
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.196
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1196
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 사선 구도와 창밖을 향한 시선을 정확히 연출했으며 창밖의 정비소 배경도 완벽히 구현했으나, 앰버의 멜빵바지가 누락됨."
   },
   {
    "label": "B",
    "score": 1196,
    "verdict_ko": "인물들의 복장은 레퍼런스와 일치하나, 지시된 사선 구도를 무시하고 정면 렌즈를 응시하여 샷 텍스트의 핵심 연출을 위반함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B02.png",
    "asset_id": "2d5b2b8c-9d71-4f7f-acce-b7f26107156a",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-f5df-7c9f-adbe-4f7180dc3c00",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh19__bgfirst_bg.png",
   "bg_asset_id": "3517c78d-2559-4fa2-9d40-d3795bb05053",
   "bg_record_key": "S51sh19::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S51sh19::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:24:05.177271+00:00",
  "fingerprint": "99e2ac37c3d592ae037251fa09890dfa8b8bab3af1f746c963964d0f744eda9b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S51sh19_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S51sh19_sel.png",
  "source_sha256": "9f1dd6df4ec8d13e9152a8314f666787faa1304b0ae50bf3623dc205ef80c1e2",
  "file": "S51sh19_cine.png",
  "staged_sha256": "e807f55c04a0b38f9157e87f48d6f702e0d320c61bdbd8251735a540b3a8bba1",
  "latency_ms": 8954
 },
 "S52sh1::signage": {
  "fp": "9e37ac79ee8f4c1a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S52sh1": {
  "input_fingerprint": "bc2c52f9077fa81f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie now wears an oversized straw hat, oversized rubber boots and a colorful raincoat over his repaired, worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie now wears an oversized straw hat, oversized rubber boots and a colorful raincoat over his repaired, worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie now wears an oversized straw hat, oversized rubber boots and a colorful raincoat over his repaired, worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S52sh1__bgfirst_bg.png",
     "asset_id": "4bc8f5c0-751d-4e2e-8589-aaf2e188a7ae",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S52sh1.png",
     "asset_id": "364391b6-7633-4178-a784-c9f3499899a6",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_rural_walking_road_5c3d2f.png",
     "asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 정면을 향해 걸어오고 있으며, 도로는 프레임 우측 뒤로 뻗어 있음.",
    "built_space": "레퍼런스와 일치하는 시골 아스팔트 도로, 논밭, 산, 전신주가 올바르게 배치됨.",
    "entities": "찰리의 금속 전신, 밀짚모자, 알록달록한 우비, 장화 모두 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "오른발 앞꿈치와 왼발 뒤꿈치가 모두 바닥에 닿아 있어 앞발을 공중에 든 상태가 아님."
   },
   {
    "label": "B",
    "direction": "찰리가 프레임 우측 대각선 방향을 향해 걷고 있으며, 시선과 진행 방향이 도로와 일치함.",
    "built_space": "레퍼런스와 일치하는 시골 아스팔트 도로, 논밭, 산, 전신주가 올바르게 배치됨.",
    "entities": "찰리의 금속 전신, 밀짚모자, 알록달록한 우비, 장화 모두 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "왼발(뒷발)로 지면을 디디고 오른발(앞발)을 공중에 들어 올린 걷기 자세가 중력과 동작에 맞게 표현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지시된 비스듬한 구도(oblique view)와 앞발을 들고 뒷발로 밀어내는 정확한 mid-stride 자세를 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "캐릭터와 배경의 디테일은 좋으나, 두 발이 모두 땅에 닿아 있어 프롬프트가 요구한 앞발을 든 걷기 자세를 충족하지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 정면을 향해 걸어오고 있으며, 도로는 프레임 우측 뒤로 뻗어 있음.",
        "built_space": "레퍼런스와 일치하는 시골 아스팔트 도로, 논밭, 산, 전신주가 올바르게 배치됨.",
        "entities": "찰리의 금속 전신, 밀짚모자, 알록달록한 우비, 장화 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "오른발 앞꿈치와 왼발 뒤꿈치가 모두 바닥에 닿아 있어 앞발을 공중에 든 상태가 아님."
       },
       {
        "label": "B",
        "direction": "찰리가 프레임 우측 대각선 방향을 향해 걷고 있으며, 시선과 진행 방향이 도로와 일치함.",
        "built_space": "레퍼런스와 일치하는 시골 아스팔트 도로, 논밭, 산, 전신주가 올바르게 배치됨.",
        "entities": "찰리의 금속 전신, 밀짚모자, 알록달록한 우비, 장화 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "왼발(뒷발)로 지면을 디디고 오른발(앞발)을 공중에 들어 올린 걷기 자세가 중력과 동작에 맞게 표현됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지시된 비스듬한 구도(oblique view)와 앞발을 들고 뒷발로 밀어내는 정확한 mid-stride 자세를 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "캐릭터와 배경의 디테일은 좋으나, 두 발이 모두 땅에 닿아 있어 프롬프트가 요구한 앞발을 든 걷기 자세를 충족하지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 정면을 향해 걸어오고 있으며, 도로는 프레임 우측 뒤로 뻗어 있음.",
        "built_space": "레퍼런스와 일치하는 시골 아스팔트 도로, 논밭, 산, 전신주가 올바르게 배치됨.",
        "entities": "찰리의 금속 전신, 밀짚모자, 알록달록한 우비, 장화 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "오른발 앞꿈치와 왼발 뒤꿈치가 모두 바닥에 닿아 있어 앞발을 공중에 든 상태가 아님."
       },
       {
        "label": "B",
        "direction": "찰리가 프레임 우측 대각선 방향을 향해 걷고 있으며, 시선과 진행 방향이 도로와 일치함.",
        "built_space": "레퍼런스와 일치하는 시골 아스팔트 도로, 논밭, 산, 전신주가 올바르게 배치됨.",
        "entities": "찰리의 금속 전신, 밀짚모자, 알록달록한 우비, 장화 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "왼발(뒷발)로 지면을 디디고 오른발(앞발)을 공중에 들어 올린 걷기 자세가 중력과 동작에 맞게 표현됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "오른쪽 사선으로 향하는 전신 보행과 분리된 장화가 지시에 부합하지만, 뒷발로 바닥을 밀어내는 순간은 다소 약하다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "장소와 찰리의 외형은 잘 맞지만 카메라 쪽으로 걸어와, 오른쪽으로 이어지는 진행 방향과 비스듬한 의상 실루엣 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 몸통, 앞 장화가 화면 오른쪽을 향한다. 시선은 오른쪽 앞 노면으로 내려가며, 도로도 오른쪽 뒤로 뻗어 진행 방향과 대체로 연결된다. 다만 발끝은 도로의 소실점보다 화면 오른쪽 가장자리에 더 가깝게 향한다.",
        "built_space": "중앙 황색 점선 한 줄과 양쪽 흰 가장자리선이 있는 아스팔트 도로 한 개다. 오른쪽에는 전봇대가 한 줄로 반복되고, 왼쪽에는 물찬 논과 낮은 비닐하우스들, 뒤에는 산지가 있어 장소 참조와 잘 맞는다. 찰리 한 대가 중앙선 부근에 있으며 차량이나 다른 인물은 없다. 전신과 주변 도로가 보이지만 인물이 화면 높이의 상당 부분을 차지한다.",
        "entities": "찰리 한 대만 보인다. 긴 중량감 있는 팔, 상대적으로 짧은 다리, 마모된 샌드 베이지 장갑판, 흰 분절형 마스크 얼굴과 주황빛 눈이 참조와 일치한다. 큰 밀짚모자 한 개, 다색 무늬 우비 한 벌, 녹색 고무장화 두 짝을 착용했다. 모자와 두 장화가 비스듬한 각도에서 명확히 분리되어 보인다. 낮의 실제 도로와 금속·직물·고무 재질로 읽히며 명확히 읽히는 문구는 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽의 뒷장화 밑창이 노면에 닿아 몸을 지탱한다. 오른쪽 앞장화는 발끝이 들리고 뒤꿈치가 노면에 매우 가까워, 발을 내딛는 동작은 읽히지만 완전히 들린 단계는 선명하지 않다. 뒷발 뒤꿈치가 충분히 올라가지 않아 강한 밀어내기보다는 착지 직전의 보행에 가깝다. 모자는 머리에, 우비는 어깨와 몸통에, 장화는 다리에 지지되어 있으며 공중에 무지지로 뜬 물체는 없다."
       },
       {
        "label": "B",
        "direction": "얼굴과 가슴은 거의 카메라 정면을 향하고, 앞장화도 카메라 쪽으로 나와 있다. 시선은 정면에서 약간 화면 왼쪽으로 향한다. 오른쪽 뒤로 이어지는 도로는 찰리의 앞이 아니라 뒤에 놓여, 요구한 오른쪽 사선 진행과 맞지 않는다.",
        "built_space": "황색 중앙 점선 한 줄과 양쪽 흰 가장자리선이 있는 도로 한 개, 오른쪽 전봇대 한 줄, 왼쪽 논과 낮은 농업 시설, 뒤쪽 산지가 보인다. 참조 장소의 주요 재료와 배치는 유지된다. 찰리는 중앙선 가까이 혼자 있고 전신이 들어오지만, 정면으로 크게 서 있어 빈 도로를 따라 비스듬히 걷는 구도보다 정면 인물 구도에 가깝다.",
        "entities": "찰리 한 대의 베이지색 낡은 금속 몸체, 긴 팔, 흰 마스크형 얼굴이 참조와 대체로 일치한다. 큰 밀짚모자 한 개, 다색 우비 한 벌, 녹색 장화 두 짝이 보인다. 우비는 참조보다 길게 내려오며, 정면 각도 때문에 장화와 다리의 실루엣이 일부 겹친다. 우비에 문자처럼 보이는 작은 무늬가 있지만 확실한 문구로 판독되지는 않는다. 다른 사람이나 차량은 없다.",
        "hard_violations": [],
        "physics": "앞으로 나온 장화는 밑창이 보이도록 들려 있고, 그 뒤에 일부 가려진 장화가 지면 쪽으로 내려가 몸을 받치는 보행 자세다. 접지 부위는 겹침과 그림자로 불명확하지만 전신이 근거 없이 떠 있다고 볼 증거는 없다. 앞발 들기는 읽히나 뒷발의 밀어내기는 잘 드러나지 않는다. 모자와 우비, 장화는 각각 머리·몸통·다리에 정상적으로 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "오른쪽 사선으로 향하는 전신 보행과 분리된 장화가 지시에 부합하지만, 뒷발로 바닥을 밀어내는 순간은 다소 약하다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "장소와 찰리의 외형은 잘 맞지만 카메라 쪽으로 걸어와, 오른쪽으로 이어지는 진행 방향과 비스듬한 의상 실루엣 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 몸통, 앞 장화가 화면 오른쪽을 향한다. 시선은 오른쪽 앞 노면으로 내려가며, 도로도 오른쪽 뒤로 뻗어 진행 방향과 대체로 연결된다. 다만 발끝은 도로의 소실점보다 화면 오른쪽 가장자리에 더 가깝게 향한다.",
        "built_space": "중앙 황색 점선 한 줄과 양쪽 흰 가장자리선이 있는 아스팔트 도로 한 개다. 오른쪽에는 전봇대가 한 줄로 반복되고, 왼쪽에는 물찬 논과 낮은 비닐하우스들, 뒤에는 산지가 있어 장소 참조와 잘 맞는다. 찰리 한 대가 중앙선 부근에 있으며 차량이나 다른 인물은 없다. 전신과 주변 도로가 보이지만 인물이 화면 높이의 상당 부분을 차지한다.",
        "entities": "찰리 한 대만 보인다. 긴 중량감 있는 팔, 상대적으로 짧은 다리, 마모된 샌드 베이지 장갑판, 흰 분절형 마스크 얼굴과 주황빛 눈이 참조와 일치한다. 큰 밀짚모자 한 개, 다색 무늬 우비 한 벌, 녹색 고무장화 두 짝을 착용했다. 모자와 두 장화가 비스듬한 각도에서 명확히 분리되어 보인다. 낮의 실제 도로와 금속·직물·고무 재질로 읽히며 명확히 읽히는 문구는 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽의 뒷장화 밑창이 노면에 닿아 몸을 지탱한다. 오른쪽 앞장화는 발끝이 들리고 뒤꿈치가 노면에 매우 가까워, 발을 내딛는 동작은 읽히지만 완전히 들린 단계는 선명하지 않다. 뒷발 뒤꿈치가 충분히 올라가지 않아 강한 밀어내기보다는 착지 직전의 보행에 가깝다. 모자는 머리에, 우비는 어깨와 몸통에, 장화는 다리에 지지되어 있으며 공중에 무지지로 뜬 물체는 없다."
       },
       {
        "label": "A",
        "direction": "얼굴과 가슴은 거의 카메라 정면을 향하고, 앞장화도 카메라 쪽으로 나와 있다. 시선은 정면에서 약간 화면 왼쪽으로 향한다. 오른쪽 뒤로 이어지는 도로는 찰리의 앞이 아니라 뒤에 놓여, 요구한 오른쪽 사선 진행과 맞지 않는다.",
        "built_space": "황색 중앙 점선 한 줄과 양쪽 흰 가장자리선이 있는 도로 한 개, 오른쪽 전봇대 한 줄, 왼쪽 논과 낮은 농업 시설, 뒤쪽 산지가 보인다. 참조 장소의 주요 재료와 배치는 유지된다. 찰리는 중앙선 가까이 혼자 있고 전신이 들어오지만, 정면으로 크게 서 있어 빈 도로를 따라 비스듬히 걷는 구도보다 정면 인물 구도에 가깝다.",
        "entities": "찰리 한 대의 베이지색 낡은 금속 몸체, 긴 팔, 흰 마스크형 얼굴이 참조와 대체로 일치한다. 큰 밀짚모자 한 개, 다색 우비 한 벌, 녹색 장화 두 짝이 보인다. 우비는 참조보다 길게 내려오며, 정면 각도 때문에 장화와 다리의 실루엣이 일부 겹친다. 우비에 문자처럼 보이는 작은 무늬가 있지만 확실한 문구로 판독되지는 않는다. 다른 사람이나 차량은 없다.",
        "hard_violations": [],
        "physics": "앞으로 나온 장화는 밑창이 보이도록 들려 있고, 그 뒤에 일부 가려진 장화가 지면 쪽으로 내려가 몸을 받치는 보행 자세다. 접지 부위는 겹침과 그림자로 불명확하지만 전신이 근거 없이 떠 있다고 볼 증거는 없다. 앞발 들기는 읽히나 뒷발의 밀어내기는 잘 드러나지 않는다. 모자와 우비, 장화는 각각 머리·몸통·다리에 정상적으로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.375,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1375
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시된 비스듬한 구도(oblique view)와 앞발을 들고 뒷발로 밀어내는 정확한 mid-stride 자세를 완벽하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1375,
    "verdict_ko": "캐릭터와 배경의 디테일은 좋으나, 두 발이 모두 땅에 닿아 있어 프롬프트가 요구한 앞발을 든 걷기 자세를 충족하지 못했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_rural_walking_road_5c3d2f.png",
    "asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-f92f-7b09-8cfd-56831e5cfb9c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S52sh1__bgfirst_bg.png",
   "bg_asset_id": "4bc8f5c0-751d-4e2e-8589-aaf2e188a7ae",
   "bg_record_key": "S52sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "rural_walking_road",
   "groupbg_asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S52sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:25:11.603646+00:00",
  "fingerprint": "7a2915179a7f6bb00f56dba68967d60c33c0a658e33ecd1a8462de42f3e8b569",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S52sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S52sh1_sel.png",
  "source_sha256": "b1645184de113222e849f84c15a4b5b8d79200156f68c1e90c5b4813dfe68c69",
  "file": "S52sh1_cine.png",
  "staged_sha256": "38aecba5228c0df98bacea2b233e15b54a341ea7895ea7f6a1bcd568998da72f",
  "latency_ms": 10862
 },
 "S53sh2::signage": {
  "fp": "6f3c543335c9f5c1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::0a3f52341ebda2d0": {
  "subjects": [],
  "subject_text": "수도권 도심 민병대 사무실\n책상 위에 큰 반도 지도가 펼쳐진 사무실. 다트판과 다트가 있고 책상 주변에 업무용 좌석이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L213",
  "scope_role": "location_interior",
  "scope_sha": "c19a59f424a5380d"
 },
 "groupbg::urban_militia_office": {
  "input_fingerprint": "8b9d4ffe98043a1c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "urban_militia_office",
    "tags": [
     "S53sh2",
     "S53sh3"
    ]
   },
   "context_sig": "6935d1a2b18c8a4c"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수도권 도심 민병대 사무실: 거대한 지도가 펼쳐져 있고 다트 게임판이 놓여 있는 작전 회의실 겸 휴게 공간. (특징: 책상 위를 덮을 만큼 커다란 한반도 종이 지도; 벽에 걸린 다트 과녁판과 꽂힌 다트핀)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 53. 수도권 도심 민병대 사무실\n- 책상 위에 한반도 지도가 펼쳐져 있다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수도권 도심 민병대 사무실: 거대한 지도가 펼쳐져 있고 다트 게임판이 놓여 있는 작전 회의실 겸 휴게 공간. (특징: 책상 위를 덮을 만큼 커다란 한반도 종이 지도; 벽에 걸린 다트 과녁판과 꽂힌 다트핀)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 53. 수도권 도심 민병대 사무실\n- 책상 위에 한반도 지도가 펼쳐져 있다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_urban_militia_office_ccf685.png",
  "asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26",
  "input_asset_ids": [
   "70416dc7-78a1-45aa-80e3-0fdad91a577a"
  ],
  "origin_tag": "S53sh2",
  "place_text": "At the map-covered desk inside an urban militia office, under ordinary office lighting.",
  "origin_inputs": {
   "place_text": "At the map-covered desk inside an urban militia office, under ordinary office lighting.",
   "time_of_day_en": "day",
   "conti_asset_id": "70416dc7-78a1-45aa-80e3-0fdad91a577a"
  }
 },
 "S53sh2::bgfirst_bg": {
  "input_fingerprint": "71eef7a3fbfb7c55",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2__bgfirst_bg.png",
  "asset_id": "e67876e3-065b-4e8d-bf17-983782ebf5b0",
  "input_asset_ids": [
   "70416dc7-78a1-45aa-80e3-0fdad91a577a",
   "91778f94-deb9-4780-814b-5f1f344f3b26"
  ]
 },
 "S53sh2": {
  "input_fingerprint": "c87ff8acef502815",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A map of the Korean Peninsula lies spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A map of the Korean Peninsula lies spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A map of the Korean Peninsula lies spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2__bgfirst_bg.png",
     "asset_id": "e67876e3-065b-4e8d-bf17-983782ebf5b0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S53sh2.png",
     "asset_id": "70416dc7-78a1-45aa-80e3-0fdad91a577a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_urban_militia_office_ccf685.png",
     "asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 앞쪽 아래를 향하고 있으며 분노하며 소리치는 방향성을 가짐.",
    "built_space": "위치 레퍼런스의 사무실 구조(창문, 파일 캐비닛, 열린 문, 선반, 다트판)가 잘 보이나, 요구된 비스듬한 구도가 아닌 책상이 가로로 놓여 있음.",
    "entities": "박철진의 얼굴과 머리 모양은 일치하지만, 복장이 레퍼런스(검은 제복과 붉은 완장)가 아닌 회색 정장임. 화면 우측 하단에 프롬프트에 없는 정장 입은 타인의 어깨가 나타남.",
    "hard_violations": [
     "[gemini-pro] invented people (프롬프트에 지시되지 않은 추가 인물의 신체 일부가 우측 전경에 등장함)",
     "[gpt-high] 오른쪽 전경에 박철진 외 다른 인물의 어깨·상체 일부를 추가했다.",
     "[gpt-high] 다트판의 숫자가 읽혀 모든 문자를 판독 불가능하게 하라는 지시를 위반했다."
    ],
    "physics": "두 주먹이 책상 위에 올려져 있으나, 강하게 내리친다기보다는 단순히 얹혀 있는 상태를 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 책상 앞쪽 아래를 향하며 강하게 에너지를 분출하는 방향성을 가짐.",
    "built_space": "위치 레퍼런스의 사무실(창문, 서랍장, 벽면 지도, 철제 선반) 배경을 잘 반영함. 프롬프트 지시대로 책상의 근경 모서리가 화면 하단을 비스듬히 가로지르고 있음.",
    "entities": "박철진의 얼굴, 핏발 선 눈, 벌린 입, 검은 제복과 붉은 완장 등 모든 외형과 복장이 레퍼런스 및 프롬프트와 완벽히 일치함. 책상 위에 한반도 지도가 올바르게 펼쳐져 있음.",
    "hard_violations": [
     "[gpt-high] 장소 참조에 없는 켜진 책상등 두 개를 추가해 고정된 장소의 소품 및 조명 구성을 바꾸었다."
    ],
    "physics": "어깨와 팔의 근육 긴장도를 통해 체중을 실어 두 주먹으로 책상 위 지도를 강하게 내리친 순간의 물리적 타격감과 지탱이 사실적으로 표현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "캐릭터의 외모 및 의상을 레퍼런스와 정확히 일치시켰고, 요구된 카메라 구도(비스듬한 책상 모서리, 지도)와 핏발 선 눈의 디테일까지 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "우측 하단에 프롬프트에 없는 인물의 신체 일부가 등장하여 치명적인 오류를 범했으며, 캐릭터의 의상(정장)도 레퍼런스와 불일치합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 앞쪽 아래를 향하고 있으며 분노하며 소리치는 방향성을 가짐.",
        "built_space": "위치 레퍼런스의 사무실 구조(창문, 파일 캐비닛, 열린 문, 선반, 다트판)가 잘 보이나, 요구된 비스듬한 구도가 아닌 책상이 가로로 놓여 있음.",
        "entities": "박철진의 얼굴과 머리 모양은 일치하지만, 복장이 레퍼런스(검은 제복과 붉은 완장)가 아닌 회색 정장임. 화면 우측 하단에 프롬프트에 없는 정장 입은 타인의 어깨가 나타남.",
        "hard_violations": [
         "invented people (프롬프트에 지시되지 않은 추가 인물의 신체 일부가 우측 전경에 등장함)"
        ],
        "physics": "두 주먹이 책상 위에 올려져 있으나, 강하게 내리친다기보다는 단순히 얹혀 있는 상태를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 책상 앞쪽 아래를 향하며 강하게 에너지를 분출하는 방향성을 가짐.",
        "built_space": "위치 레퍼런스의 사무실(창문, 서랍장, 벽면 지도, 철제 선반) 배경을 잘 반영함. 프롬프트 지시대로 책상의 근경 모서리가 화면 하단을 비스듬히 가로지르고 있음.",
        "entities": "박철진의 얼굴, 핏발 선 눈, 벌린 입, 검은 제복과 붉은 완장 등 모든 외형과 복장이 레퍼런스 및 프롬프트와 완벽히 일치함. 책상 위에 한반도 지도가 올바르게 펼쳐져 있음.",
        "hard_violations": [],
        "physics": "어깨와 팔의 근육 긴장도를 통해 체중을 실어 두 주먹으로 책상 위 지도를 강하게 내리친 순간의 물리적 타격감과 지탱이 사실적으로 표현됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "캐릭터의 외모 및 의상을 레퍼런스와 정확히 일치시켰고, 요구된 카메라 구도(비스듬한 책상 모서리, 지도)와 핏발 선 눈의 디테일까지 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "우측 하단에 프롬프트에 없는 인물의 신체 일부가 등장하여 치명적인 오류를 범했으며, 캐릭터의 의상(정장)도 레퍼런스와 불일치합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 앞쪽 아래를 향하고 있으며 분노하며 소리치는 방향성을 가짐.",
        "built_space": "위치 레퍼런스의 사무실 구조(창문, 파일 캐비닛, 열린 문, 선반, 다트판)가 잘 보이나, 요구된 비스듬한 구도가 아닌 책상이 가로로 놓여 있음.",
        "entities": "박철진의 얼굴과 머리 모양은 일치하지만, 복장이 레퍼런스(검은 제복과 붉은 완장)가 아닌 회색 정장임. 화면 우측 하단에 프롬프트에 없는 정장 입은 타인의 어깨가 나타남.",
        "hard_violations": [
         "invented people (프롬프트에 지시되지 않은 추가 인물의 신체 일부가 우측 전경에 등장함)"
        ],
        "physics": "두 주먹이 책상 위에 올려져 있으나, 강하게 내리친다기보다는 단순히 얹혀 있는 상태를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 책상 앞쪽 아래를 향하며 강하게 에너지를 분출하는 방향성을 가짐.",
        "built_space": "위치 레퍼런스의 사무실(창문, 서랍장, 벽면 지도, 철제 선반) 배경을 잘 반영함. 프롬프트 지시대로 책상의 근경 모서리가 화면 하단을 비스듬히 가로지르고 있음.",
        "entities": "박철진의 얼굴, 핏발 선 눈, 벌린 입, 검은 제복과 붉은 완장 등 모든 외형과 복장이 레퍼런스 및 프롬프트와 완벽히 일치함. 책상 위에 한반도 지도가 올바르게 펼쳐져 있음.",
        "hard_violations": [],
        "physics": "어깨와 팔의 근육 긴장도를 통해 체중을 실어 두 주먹으로 책상 위 지도를 강하게 내리친 순간의 물리적 타격감과 지탱이 사실적으로 표현됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "군복·붉은 완장과 충혈된 얼굴, 두 주먹의 책상 접촉은 더 충실하지만, 참조에 없는 조명 두 개를 추가했고 얼굴 클로즈업보다 구도가 넓다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "사무실 배치는 가깝지만 금지된 추가 인물과 읽히는 다트판 숫자가 있으며, 정장 차림과 넓은 구도가 인물 참조 및 클로즈업 지시를 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 눈은 카메라보다 아래쪽, 책상 건너편의 낮은 지점을 향한다. 시선의 대상 인물은 보이지 않는다. 두 주먹은 아래로 향해 지도와 책상 위에 닿아 있어 타격 대상은 맞는다. 지도 인쇄면은 위를 향한다.",
        "built_space": "왼쪽에 블라인드 창과 낮은 서류장이 있고, 뒤에는 지도 게시판 하나와 금속 수납 선반이 보인다. 의자 등받이 일부가 몸 뒤에 있다. 목재 책상 모서리가 아래쪽에서 두 사선으로 만난다. 참조에 없는 켜진 책상등이 왼쪽 하나, 오른쪽 가장자리 하나 추가되어 있다. 책상과 상체가 상당히 넓게 노출되어 얼굴 중심 클로즈업보다 넓다.",
        "entities": "보이는 사람은 중년 한국인 남성 한 명이며, 짧은 검은 머리와 얼굴 윤곽은 참조에 가깝다. 검은 군복과 붉은 상완 완장이 일치한다. 입을 크게 벌리고 눈 주변이 붉으며 홍채와 동공은 정상적으로 보인다. 두 주먹 사이와 뒤에 위를 향한 인쇄 지도가 펼쳐져 있으나 한반도 윤곽은 부분적으로만 확인된다. 명확히 읽히는 문구나 화면 자막은 보이지 않는다.",
        "hard_violations": [
         "장소 참조에 없는 켜진 책상등 두 개를 추가해 고정된 장소의 소품 및 조명 구성을 바꾸었다."
        ],
        "physics": "양 주먹의 아래쪽이 지도와 책상에 접촉하고, 팔이 앞으로 기울어진 상체를 지지한다. 화면 밖 하체의 자세는 확인할 수 없지만 공중에 뜬 몸으로 보이지는 않는다. 타격 직후의 자세로 가능하나 왼쪽 전경 주먹이 원근상 크게 강조된다. 지도와 다른 물건은 책상이나 수납면에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "박철진은 화면 오른쪽 책상 건너편을 바라보며 소리친다. 오른쪽 전경에는 상대 인물의 어깨처럼 보이는 신체 일부가 있어 시선이 그쪽을 향하는 대면 구도로 읽힌다. 두 주먹은 아래쪽 지도와 책상을 향해 접촉한다. 지도 인쇄면은 정상적으로 위를 향한다.",
        "built_space": "왼쪽 블라인드 창과 낮은 서류장, 뒤쪽 지도 게시판 하나와 금속 선반, 오른쪽 열린 문 하나, 다트판 하나와 그 아래 수납대가 있어 참조 사무실의 배치는 상당히 가깝다. 오른쪽 의자 하나와 주인공 뒤 의자 일부가 보인다. 다만 넓은 책상과 방 전체의 시설을 보여 주어 요청한 얼굴 클로즈업이 아니며, 가까운 책상 모서리의 사선보다 지도 면적이 전경을 지배한다.",
        "entities": "주인공은 짧은 검은 머리의 중년 한국인 남성으로 얼굴은 참조와 유사하고, 입을 크게 벌린 분노 표정과 정상적인 눈을 보인다. 그러나 군복과 붉은 완장 대신 검은 정장, 흰 셔츠, 넥타이를 착용했다. 펼쳐진 한반도 지도는 분명하다. 오른쪽 전경에는 허용되지 않은 다른 인물의 어깨·상체 일부가 보이며 신원은 확인할 수 없다. 다트판에는 읽을 수 있는 숫자가 있다.",
        "hard_violations": [
         "오른쪽 전경에 박철진 외 다른 인물의 어깨·상체 일부를 추가했다.",
         "다트판의 숫자가 읽혀 모든 문자를 판독 불가능하게 하라는 지시를 위반했다."
        ],
        "physics": "양 주먹이 지도 위 책상에 닿아 있고 팔과 어깨의 연결도 자연스럽다. 앞으로 숙인 상체는 책상을 짚은 팔로 지지될 수 있어 두 주먹으로 내려친 순간으로 가능한 자세다. 지도와 서류는 책상에 놓여 있다. 오른쪽 추가 인물은 일부만 잘려 보여 지지 자세를 판단할 수 없으며, 떠 있다고 단정할 근거는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "군복·붉은 완장과 충혈된 얼굴, 두 주먹의 책상 접촉은 더 충실하지만, 참조에 없는 조명 두 개를 추가했고 얼굴 클로즈업보다 구도가 넓다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "사무실 배치는 가깝지만 금지된 추가 인물과 읽히는 다트판 숫자가 있으며, 정장 차림과 넓은 구도가 인물 참조 및 클로즈업 지시를 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 눈은 카메라보다 아래쪽, 책상 건너편의 낮은 지점을 향한다. 시선의 대상 인물은 보이지 않는다. 두 주먹은 아래로 향해 지도와 책상 위에 닿아 있어 타격 대상은 맞는다. 지도 인쇄면은 위를 향한다.",
        "built_space": "왼쪽에 블라인드 창과 낮은 서류장이 있고, 뒤에는 지도 게시판 하나와 금속 수납 선반이 보인다. 의자 등받이 일부가 몸 뒤에 있다. 목재 책상 모서리가 아래쪽에서 두 사선으로 만난다. 참조에 없는 켜진 책상등이 왼쪽 하나, 오른쪽 가장자리 하나 추가되어 있다. 책상과 상체가 상당히 넓게 노출되어 얼굴 중심 클로즈업보다 넓다.",
        "entities": "보이는 사람은 중년 한국인 남성 한 명이며, 짧은 검은 머리와 얼굴 윤곽은 참조에 가깝다. 검은 군복과 붉은 상완 완장이 일치한다. 입을 크게 벌리고 눈 주변이 붉으며 홍채와 동공은 정상적으로 보인다. 두 주먹 사이와 뒤에 위를 향한 인쇄 지도가 펼쳐져 있으나 한반도 윤곽은 부분적으로만 확인된다. 명확히 읽히는 문구나 화면 자막은 보이지 않는다.",
        "hard_violations": [
         "장소 참조에 없는 켜진 책상등 두 개를 추가해 고정된 장소의 소품 및 조명 구성을 바꾸었다."
        ],
        "physics": "양 주먹의 아래쪽이 지도와 책상에 접촉하고, 팔이 앞으로 기울어진 상체를 지지한다. 화면 밖 하체의 자세는 확인할 수 없지만 공중에 뜬 몸으로 보이지는 않는다. 타격 직후의 자세로 가능하나 왼쪽 전경 주먹이 원근상 크게 강조된다. 지도와 다른 물건은 책상이나 수납면에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "박철진은 화면 오른쪽 책상 건너편을 바라보며 소리친다. 오른쪽 전경에는 상대 인물의 어깨처럼 보이는 신체 일부가 있어 시선이 그쪽을 향하는 대면 구도로 읽힌다. 두 주먹은 아래쪽 지도와 책상을 향해 접촉한다. 지도 인쇄면은 정상적으로 위를 향한다.",
        "built_space": "왼쪽 블라인드 창과 낮은 서류장, 뒤쪽 지도 게시판 하나와 금속 선반, 오른쪽 열린 문 하나, 다트판 하나와 그 아래 수납대가 있어 참조 사무실의 배치는 상당히 가깝다. 오른쪽 의자 하나와 주인공 뒤 의자 일부가 보인다. 다만 넓은 책상과 방 전체의 시설을 보여 주어 요청한 얼굴 클로즈업이 아니며, 가까운 책상 모서리의 사선보다 지도 면적이 전경을 지배한다.",
        "entities": "주인공은 짧은 검은 머리의 중년 한국인 남성으로 얼굴은 참조와 유사하고, 입을 크게 벌린 분노 표정과 정상적인 눈을 보인다. 그러나 군복과 붉은 완장 대신 검은 정장, 흰 셔츠, 넥타이를 착용했다. 펼쳐진 한반도 지도는 분명하다. 오른쪽 전경에는 허용되지 않은 다른 인물의 어깨·상체 일부가 보이며 신원은 확인할 수 없다. 다트판에는 읽을 수 있는 숫자가 있다.",
        "hard_violations": [
         "오른쪽 전경에 박철진 외 다른 인물의 어깨·상체 일부를 추가했다.",
         "다트판의 숫자가 읽혀 모든 문자를 판독 불가능하게 하라는 지시를 위반했다."
        ],
        "physics": "양 주먹이 지도 위 책상에 닿아 있고 팔과 어깨의 연결도 자연스럽다. 앞으로 숙인 상체는 책상을 짚은 팔로 지지될 수 있어 두 주먹으로 내려친 순간으로 가능한 자세다. 지도와 서류는 책상에 놓여 있다. 오른쪽 추가 인물은 일부만 잘려 보여 지지 자세를 판단할 수 없으며, 떠 있다고 단정할 근거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.8,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.55,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people (프롬프트에 지시되지 않은 추가 인물의 신체 일부가 우측 전경에 등장함)",
     "[gpt-high] 오른쪽 전경에 박철진 외 다른 인물의 어깨·상체 일부를 추가했다.",
     "[gpt-high] 다트판의 숫자가 읽혀 모든 문자를 판독 불가능하게 하라는 지시를 위반했다."
    ],
    "B": [
     "[gpt-high] 장소 참조에 없는 켜진 책상등 두 개를 추가해 고정된 장소의 소품 및 조명 구성을 바꾸었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 550
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "캐릭터의 외모 및 의상을 레퍼런스와 정확히 일치시켰고, 요구된 카메라 구도(비스듬한 책상 모서리, 지도)와 핏발 선 눈의 디테일까지 완벽하게 구현했습니다.  ★위반: [gpt-high] 장소 참조에 없는 켜진 책상등 두 개를 추가해 고정된 장소의 소품 및 조명 구성을 바꾸었다."
   },
   {
    "label": "A",
    "score": 550,
    "verdict_ko": "우측 하단에 프롬프트에 없는 인물의 신체 일부가 등장하여 치명적인 오류를 범했으며, 캐릭터의 의상(정장)도 레퍼런스와 불일치합니다.  ★위반: [gemini-pro] invented people (프롬프트에 지시되지 않은 추가 인물의 신체 일부가 우측 전경에 등장함) / [gpt-high] 오른쪽 전경에 박철진 외 다른 인물의 어깨·상체 일부를 추가했다. / [gpt-high] 다트판의 숫자가 읽혀 모든 문자를 판독 불가능하게 하라는 지시를 위반했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_urban_militia_office_ccf685.png",
    "asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-fdf0-79b7-a077-16a483b3f779",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2__bgfirst_bg.png",
   "bg_asset_id": "e67876e3-065b-4e8d-bf17-983782ebf5b0",
   "bg_record_key": "S53sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "urban_militia_office",
   "groupbg_asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S53sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:26:18.515478+00:00",
  "fingerprint": "78086b08da5c331c72c5593c8110cd2645ed845c3019f148bff82135529653fe",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S53sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S53sh2_sel.png",
  "source_sha256": "40dbac73539cf7c19690692b448e6a2650e93404f2cdc4a0e36bd68746c3f288",
  "file": "S53sh2_cine.png",
  "staged_sha256": "fc9b1f945b6672901297013866015e33a308f9c889caaf5ca00140724252287d",
  "latency_ms": 10218
 },
 "S53sh3::signage": {
  "fp": "a6d7c9c683d47ee0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S53sh3": {
  "input_fingerprint": "012b5cbc82e7b79e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 박철진의 귓가에 입을 바짝 댄 채 은밀하게 속삭이는 수하 1의 조심스러운 상체.\n\nLOCATION (lock): Beside the commander's desk inside the urban militia office, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Remaining beneath the pair during the whispered report) — The same side edge recedes beneath their upper bodies; used as Maintains spatial continuity with the fist impact and anchors their close positioning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding office illumination and contrast so the lowered voices, not a lighting shift, create the secrecy.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The map of the Korean Peninsula remains spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 박철진의 귓가에 입을 바짝 댄 채 은밀하게 속삭이는 수하 1의 조심스러운 상체.\n\nLOCATION (lock): Beside the commander's desk inside the urban militia office, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Remaining beneath the pair during the whispered report) — The same side edge recedes beneath their upper bodies; used as Maintains spatial continuity with the fist impact and anchors their close positioning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding office illumination and contrast so the lowered voices, not a lighting shift, create the secrecy.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The map of the Korean Peninsula remains spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 박철진의 귓가에 입을 바짝 댄 채 은밀하게 속삭이는 수하 1의 조심스러운 상체.\n\nLOCATION (lock): Beside the commander's desk inside the urban militia office, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Remaining beneath the pair during the whispered report) — The same side edge recedes beneath their upper bodies; used as Maintains spatial continuity with the fist impact and anchors their close positioning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding office illumination and contrast so the lowered voices, not a lighting shift, create the secrecy.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The map of the Korean Peninsula remains spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수하 1이 박철진의 왼쪽 귀를 향해 고개를 숙이고 오른손을 올려 속삭임. 박철진은 시선을 앞쪽 아래의 지도로 향함.",
    "built_space": "사무실 내부. 왼쪽에 창문, 뒤쪽에 철제 선반과 벽에 붙은 지도가 있음. 이전 샷에 있던 데스크 램프와 철망 게시판이 누락되고 구조가 변경됨.",
    "entities": "박철진(충혈된 눈과 땀방울 묘사 누락), 수하 1(모자 착용), 책상 위 한반도 지도. 책상 위에 수하 1의 것과 동일해 보이는 여분의 모자가 놓여 있음.",
    "hard_violations": [
     "[gemini-pro] 발명/중복된 사물 (책상 위에 놓인 여분의 모자)",
     "[gemini-pro] 물리적으로 불가능한 해부학 (얼굴과 융합되고 형태가 무너진 수하 1의 오른손)",
     "[gpt-high] 수하가 참조의 군모를 착용하고 있는데도 책상 위에 별도의 군모를 하나 더 생성해, 지정되지 않은 중복 소품을 추가했다."
    ],
    "physics": "박철진은 양팔을 책상에 기댐. 수하 1은 왼팔을 책상에 짚어 몸을 지탱하며 오른손을 들어 올림."
   },
   {
    "label": "B",
    "direction": "수하 1이 박철진의 왼쪽 귀에 바짝 다가가 오른손을 입가에 대고 속삭임. 박철진은 앞을 주시함.",
    "built_space": "사무실 내부. 왼쪽의 램프와 캐비닛, 오른쪽의 철망 게시판과 철제 선반 등 이전 샷의 공간 구성과 배치가 완벽하게 일치함.",
    "entities": "박철진(충혈된 눈, 땀, 주먹 쥔 자세 등 이전 샷 완벽 유지), 수하 1(모자 미착용, 박철진과 외양 및 머리스타일이 매우 흡사하게 렌더링됨), 책상 위 한반도 지도.",
    "hard_violations": [
     "[gpt-high] 왼쪽 배경 책상의 명패에 판독 가능한 한글이 노출되어, 어떤 읽을 수 있는 글자도 허용하지 않는 조건을 위반한다."
    ],
    "physics": "박철진의 오른 주먹이 책상 위를 짚어 체중을 지탱함. 수하 1은 상체를 기울여 자세를 유지한 채 오른손을 입가에 자연스럽게 위치시킴."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "이전 샷의 배경(조명, 철망 등)과 박철진의 강렬한 인상(충혈된 눈)을 상실했으며, 책상 위에 중복된 모자가 있고 손가락 해부학이 무너져 실격 사유가 됩니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "수하 1의 모자가 누락되고 얼굴이 박철진과 다소 비슷하게 묘사되었으나, 이전 샷의 사무실 배경 디테일과 박철진의 상태(충혈된 눈, 주먹 쥔 포즈)를 완벽하게 유지하여 훨씬 우수합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수하 1이 박철진의 왼쪽 귀를 향해 고개를 숙이고 오른손을 올려 속삭임. 박철진은 시선을 앞쪽 아래의 지도로 향함.",
        "built_space": "사무실 내부. 왼쪽에 창문, 뒤쪽에 철제 선반과 벽에 붙은 지도가 있음. 이전 샷에 있던 데스크 램프와 철망 게시판이 누락되고 구조가 변경됨.",
        "entities": "박철진(충혈된 눈과 땀방울 묘사 누락), 수하 1(모자 착용), 책상 위 한반도 지도. 책상 위에 수하 1의 것과 동일해 보이는 여분의 모자가 놓여 있음.",
        "hard_violations": [
         "발명/중복된 사물 (책상 위에 놓인 여분의 모자)",
         "물리적으로 불가능한 해부학 (얼굴과 융합되고 형태가 무너진 수하 1의 오른손)"
        ],
        "physics": "박철진은 양팔을 책상에 기댐. 수하 1은 왼팔을 책상에 짚어 몸을 지탱하며 오른손을 들어 올림."
       },
       {
        "label": "B",
        "direction": "수하 1이 박철진의 왼쪽 귀에 바짝 다가가 오른손을 입가에 대고 속삭임. 박철진은 앞을 주시함.",
        "built_space": "사무실 내부. 왼쪽의 램프와 캐비닛, 오른쪽의 철망 게시판과 철제 선반 등 이전 샷의 공간 구성과 배치가 완벽하게 일치함.",
        "entities": "박철진(충혈된 눈, 땀, 주먹 쥔 자세 등 이전 샷 완벽 유지), 수하 1(모자 미착용, 박철진과 외양 및 머리스타일이 매우 흡사하게 렌더링됨), 책상 위 한반도 지도.",
        "hard_violations": [],
        "physics": "박철진의 오른 주먹이 책상 위를 짚어 체중을 지탱함. 수하 1은 상체를 기울여 자세를 유지한 채 오른손을 입가에 자연스럽게 위치시킴."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "이전 샷의 배경(조명, 철망 등)과 박철진의 강렬한 인상(충혈된 눈)을 상실했으며, 책상 위에 중복된 모자가 있고 손가락 해부학이 무너져 실격 사유가 됩니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "수하 1의 모자가 누락되고 얼굴이 박철진과 다소 비슷하게 묘사되었으나, 이전 샷의 사무실 배경 디테일과 박철진의 상태(충혈된 눈, 주먹 쥔 포즈)를 완벽하게 유지하여 훨씬 우수합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수하 1이 박철진의 왼쪽 귀를 향해 고개를 숙이고 오른손을 올려 속삭임. 박철진은 시선을 앞쪽 아래의 지도로 향함.",
        "built_space": "사무실 내부. 왼쪽에 창문, 뒤쪽에 철제 선반과 벽에 붙은 지도가 있음. 이전 샷에 있던 데스크 램프와 철망 게시판이 누락되고 구조가 변경됨.",
        "entities": "박철진(충혈된 눈과 땀방울 묘사 누락), 수하 1(모자 착용), 책상 위 한반도 지도. 책상 위에 수하 1의 것과 동일해 보이는 여분의 모자가 놓여 있음.",
        "hard_violations": [
         "발명/중복된 사물 (책상 위에 놓인 여분의 모자)",
         "물리적으로 불가능한 해부학 (얼굴과 융합되고 형태가 무너진 수하 1의 오른손)"
        ],
        "physics": "박철진은 양팔을 책상에 기댐. 수하 1은 왼팔을 책상에 짚어 몸을 지탱하며 오른손을 들어 올림."
       },
       {
        "label": "B",
        "direction": "수하 1이 박철진의 왼쪽 귀에 바짝 다가가 오른손을 입가에 대고 속삭임. 박철진은 앞을 주시함.",
        "built_space": "사무실 내부. 왼쪽의 램프와 캐비닛, 오른쪽의 철망 게시판과 철제 선반 등 이전 샷의 공간 구성과 배치가 완벽하게 일치함.",
        "entities": "박철진(충혈된 눈, 땀, 주먹 쥔 자세 등 이전 샷 완벽 유지), 수하 1(모자 미착용, 박철진과 외양 및 머리스타일이 매우 흡사하게 렌더링됨), 책상 위 한반도 지도.",
        "hard_violations": [],
        "physics": "박철진의 오른 주먹이 책상 위를 짚어 체중을 지탱함. 수하 1은 상체를 기울여 자세를 유지한 채 오른손을 입가에 자연스럽게 위치시킴."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "귓가에 밀착한 속삭임과 중간숏은 맞지만, 배경 명패의 판독 가능한 한글이 금지 조건을 위반하고 수하의 모자·인상 및 사무실 배치도 참조에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "속삭이는 상체의 중간숏, 인물 외형과 기존 사무실의 연속성이 더 충실하지만, 수하가 착용한 모자 외에 책상 위 모자를 추가해 완전한 적합작은 아닙니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 수하는 고개를 왼쪽으로 돌리고 손으로 입을 가린 채 박철진의 왼쪽 귀 바로 옆으로 입을 가져간다. 보고의 방향과 대상이 맞는다. 박철진은 카메라를 정면으로 응시하지 않고 앞쪽 아래를 바라본다. 조준하는 무기나 이동 중인 물체는 없다.",
        "built_space": "전경의 나무 책상 하나와 펼친 지도가 두 사람의 상체 아래에 이어진다. 뒤에는 블라인드 창 하나, 오른쪽 벽 지도 하나, 왼쪽 금속 선반 두 구획과 서랍장, 별도 책상 하나와 탁상등 하나, 오른쪽 의자 하나가 보인다. 천장에는 긴 조명 두 개가 보인다. 책상 옆에 두 사람이 밀착하는 공간은 성립하지만, 참조의 넓은 왼쪽 창과 오른쪽 수납 선반 배치에 비해 공간 구성이 상당히 달라졌다. 반사는 없다.",
        "entities": "성인 동아시아계 남성 두 명만 보인다. 박철진의 중년 얼굴, 뒤로 넘긴 짧은 검은 머리, 어두운 군복과 붉은 완장은 이전 장면에 가깝다. 수하는 짧은 검은 머리와 어두운 군복·붉은 완장을 갖췄지만 참조의 모자가 없고 얼굴도 참조보다 성숙하고 각져 보인다. 책상에는 지도가 펼쳐져 있다. 왼쪽 배경 명패에는 식별 가능한 흰 한글 글자가 노출된다.",
        "hard_violations": [
         "왼쪽 배경 책상의 명패에 판독 가능한 한글이 노출되어, 어떤 읽을 수 있는 글자도 허용하지 않는 조건을 위반한다."
        ],
        "physics": "박철진의 전완과 쥔 주먹은 지도와 책상 표면에 닿아 상체를 지지한다. 수하는 몸통에서 이어지는 팔을 굽혀 손을 입가에 올렸고, 앞으로 기울인 자세에 해부학적 단절은 없다. 하체와 발은 화면 밖이므로 접지 상태를 직접 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 지도와 사무용품은 책상이나 선반에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "오른쪽 수하의 입은 박철진의 왼쪽 귀를 향해 가까이 붙어 있고, 손은 입과 귀 사이 바깥쪽을 가려 은밀한 보고를 만든다. 수하의 시선도 박철진의 머리 쪽으로 향한다. 박철진은 앞쪽 아래를 보며 듣고 있어 직접적인 카메라 응시가 없다.",
        "built_space": "전경 책상 하나의 가장자리와 지도 면이 두 사람 아래로 이어져 중간숏 안에서 위치를 설명한다. 왼쪽의 큰 블라인드 창, 아래 서류장과 바인더, 뒤쪽의 큰 벽 지도와 금속 선반은 이전 장면의 주요 공간 특징에 가깝다. 오른쪽에는 작은 창 하나, 작은 벽 지도 하나, 보조 책상 하나와 빈 의자 하나가 더 보인다. 박철진 뒤에도 의자 등받이가 보인다. 일반적인 낮 사무실 조명이 유지되며, 불가능한 반사는 없다.",
        "entities": "허용된 성인 동아시아계 남성 두 명만 등장한다. 박철진의 중년 얼굴, 짧은 검은 머리, 어두운 군복과 붉은 완장이 참조에 가깝다. 수하도 젊은 얼굴, 짧은 검은 머리와 회색 계열 군복, 참조의 군모를 유지한다. 한반도 지도는 책상 위에 펼쳐져 있다. 그러나 착용 중인 군모와 별개로 전경 지도 위에 군모 하나가 추가되어 있다. 배경 표지와 종이의 글자는 흐리거나 비스듬해 뚜렷하게 판독되지 않는다.",
        "hard_violations": [
         "수하가 참조의 군모를 착용하고 있는데도 책상 위에 별도의 군모를 하나 더 생성해, 지정되지 않은 중복 소품을 추가했다."
        ],
        "physics": "박철진은 책상에 댄 팔과 손으로 앞으로 기울어진 몸을 지지한다. 수하의 입을 가리는 손은 굽힌 팔과 자연스럽게 연결되고, 아래쪽 손도 책상 표면에 접촉한다. 하체는 프레임 밖이지만 두 상체가 지지 없이 떠 있는 형태는 아니다. 추가된 모자 역시 지도 위에 놓여 있어 물리적으로 부유하지는 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "귓가에 밀착한 속삭임과 중간숏은 맞지만, 배경 명패의 판독 가능한 한글이 금지 조건을 위반하고 수하의 모자·인상 및 사무실 배치도 참조에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "속삭이는 상체의 중간숏, 인물 외형과 기존 사무실의 연속성이 더 충실하지만, 수하가 착용한 모자 외에 책상 위 모자를 추가해 완전한 적합작은 아닙니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 수하는 고개를 왼쪽으로 돌리고 손으로 입을 가린 채 박철진의 왼쪽 귀 바로 옆으로 입을 가져간다. 보고의 방향과 대상이 맞는다. 박철진은 카메라를 정면으로 응시하지 않고 앞쪽 아래를 바라본다. 조준하는 무기나 이동 중인 물체는 없다.",
        "built_space": "전경의 나무 책상 하나와 펼친 지도가 두 사람의 상체 아래에 이어진다. 뒤에는 블라인드 창 하나, 오른쪽 벽 지도 하나, 왼쪽 금속 선반 두 구획과 서랍장, 별도 책상 하나와 탁상등 하나, 오른쪽 의자 하나가 보인다. 천장에는 긴 조명 두 개가 보인다. 책상 옆에 두 사람이 밀착하는 공간은 성립하지만, 참조의 넓은 왼쪽 창과 오른쪽 수납 선반 배치에 비해 공간 구성이 상당히 달라졌다. 반사는 없다.",
        "entities": "성인 동아시아계 남성 두 명만 보인다. 박철진의 중년 얼굴, 뒤로 넘긴 짧은 검은 머리, 어두운 군복과 붉은 완장은 이전 장면에 가깝다. 수하는 짧은 검은 머리와 어두운 군복·붉은 완장을 갖췄지만 참조의 모자가 없고 얼굴도 참조보다 성숙하고 각져 보인다. 책상에는 지도가 펼쳐져 있다. 왼쪽 배경 명패에는 식별 가능한 흰 한글 글자가 노출된다.",
        "hard_violations": [
         "왼쪽 배경 책상의 명패에 판독 가능한 한글이 노출되어, 어떤 읽을 수 있는 글자도 허용하지 않는 조건을 위반한다."
        ],
        "physics": "박철진의 전완과 쥔 주먹은 지도와 책상 표면에 닿아 상체를 지지한다. 수하는 몸통에서 이어지는 팔을 굽혀 손을 입가에 올렸고, 앞으로 기울인 자세에 해부학적 단절은 없다. 하체와 발은 화면 밖이므로 접지 상태를 직접 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 지도와 사무용품은 책상이나 선반에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "오른쪽 수하의 입은 박철진의 왼쪽 귀를 향해 가까이 붙어 있고, 손은 입과 귀 사이 바깥쪽을 가려 은밀한 보고를 만든다. 수하의 시선도 박철진의 머리 쪽으로 향한다. 박철진은 앞쪽 아래를 보며 듣고 있어 직접적인 카메라 응시가 없다.",
        "built_space": "전경 책상 하나의 가장자리와 지도 면이 두 사람 아래로 이어져 중간숏 안에서 위치를 설명한다. 왼쪽의 큰 블라인드 창, 아래 서류장과 바인더, 뒤쪽의 큰 벽 지도와 금속 선반은 이전 장면의 주요 공간 특징에 가깝다. 오른쪽에는 작은 창 하나, 작은 벽 지도 하나, 보조 책상 하나와 빈 의자 하나가 더 보인다. 박철진 뒤에도 의자 등받이가 보인다. 일반적인 낮 사무실 조명이 유지되며, 불가능한 반사는 없다.",
        "entities": "허용된 성인 동아시아계 남성 두 명만 등장한다. 박철진의 중년 얼굴, 짧은 검은 머리, 어두운 군복과 붉은 완장이 참조에 가깝다. 수하도 젊은 얼굴, 짧은 검은 머리와 회색 계열 군복, 참조의 군모를 유지한다. 한반도 지도는 책상 위에 펼쳐져 있다. 그러나 착용 중인 군모와 별개로 전경 지도 위에 군모 하나가 추가되어 있다. 배경 표지와 종이의 글자는 흐리거나 비스듬해 뚜렷하게 판독되지 않는다.",
        "hard_violations": [
         "수하가 참조의 군모를 착용하고 있는데도 책상 위에 별도의 군모를 하나 더 생성해, 지정되지 않은 중복 소품을 추가했다."
        ],
        "physics": "박철진은 책상에 댄 팔과 손으로 앞으로 기울어진 몸을 지지한다. 수하의 입을 가리는 손은 굽힌 팔과 자연스럽게 연결되고, 아래쪽 손도 책상 표면에 접촉한다. 하체는 프레임 밖이지만 두 상체가 지지 없이 떠 있는 형태는 아니다. 추가된 모자 역시 지도 위에 놓여 있어 물리적으로 부유하지는 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 발명/중복된 사물 (책상 위에 놓인 여분의 모자)",
     "[gemini-pro] 물리적으로 불가능한 해부학 (얼굴과 융합되고 형태가 무너진 수하 1의 오른손)",
     "[gpt-high] 수하가 참조의 군모를 착용하고 있는데도 책상 위에 별도의 군모를 하나 더 생성해, 지정되지 않은 중복 소품을 추가했다."
    ],
    "B": [
     "[gpt-high] 왼쪽 배경 책상의 명패에 판독 가능한 한글이 노출되어, 어떤 읽을 수 있는 글자도 허용하지 않는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1321,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "이전 샷의 배경(조명, 철망 등)과 박철진의 강렬한 인상(충혈된 눈)을 상실했으며, 책상 위에 중복된 모자가 있고 손가락 해부학이 무너져 실격 사유가 됩니다.  ★위반: [gemini-pro] 발명/중복된 사물 (책상 위에 놓인 여분의 모자) / [gemini-pro] 물리적으로 불가능한 해부학 (얼굴과 융합되고 형태가 무너진 수하 1의 오른손) / [gpt-high] 수하가 참조의 군모를 착용하고 있는데도 책상 위에 별도의 군모를 하나 더 생성해, 지정되지 않은 중복 소품을 추가했다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "수하 1의 모자가 누락되고 얼굴이 박철진과 다소 비슷하게 묘사되었으나, 이전 샷의 사무실 배경 디테일과 박철진의 상태(충혈된 눈, 주먹 쥔 포즈)를 완벽하게 유지하여 훨씬 우수합니다.  ★위반: [gpt-high] 왼쪽 배경 책상의 명패에 판독 가능한 한글이 노출되어, 어떤 읽을 수 있는 글자도 허용하지 않는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2_sel.png",
    "asset_id": "8b55cc6b-9014-4cd3-b8a0-5542a111d9ca",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:798636>",
    "asset_id": "840c8473-e2ef-41b8-ac3f-988134fc14e7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-02a3-79f7-8e06-c02ad9a1ba0a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S53sh2"
  }
 },
 "S53sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:28:06.596440+00:00",
  "fingerprint": "5350a248cfb1809c33eddca058476f3a38e95a0994abde02f85f20fba6ff927f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S53sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S53sh3_sel.png",
  "source_sha256": "725782823641f30d6aa7a6497d5ed1a334b4162903814fbcd23f6efbb13c08b6",
  "file": "S53sh3_cine.png",
  "staged_sha256": "1eed819a0e9273322923db0cad8ef90cd9cd3ce2df17734f98f2fb6bd1947efc",
  "latency_ms": 11882
 },
 "S54sh2::signage": {
  "fp": "9e517e5410cb2101",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S54sh2": {
  "input_fingerprint": "541f708cc3ff6341",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 국방장관을 향해 몸을 쑥 내민 채 두 눈을 번뜩이며 속삭이듯 입을 반쯤 벌린 윤성찬의 굳은 상체.\n\nLOCATION (lock): In the visitor conversation area inside the defense minister's office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination gives the faces controlled tonal separation without adding a specific fixture or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 국방장관 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 국방장관을 향해 몸을 쑥 내민 채 두 눈을 번뜩이며 속삭이듯 입을 반쯤 벌린 윤성찬의 굳은 상체.\n\nLOCATION (lock): In the visitor conversation area inside the defense minister's office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination gives the faces controlled tonal separation without adding a specific fixture or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 국방장관 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 국방장관을 향해 몸을 쑥 내민 채 두 눈을 번뜩이며 속삭이듯 입을 반쯤 벌린 윤성찬의 굳은 상체.\n\nLOCATION (lock): In the visitor conversation area inside the defense minister's office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination gives the faces controlled tonal separation without adding a specific fixture or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 국방장관 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "윤성찬의 시선과 앞으로 쑥 내민 상체가 국방장관의 얼굴과 귀를 정확히 향하고 있음.",
    "built_space": "이전 샷에 확립된 왼쪽 소파와 오른쪽 책상의 배치를 정확히 유지하며 주간의 채광을 잘 반영함.",
    "entities": "국방장관은 이전 샷의 복장과 자세를 유지함. 윤성찬은 얼굴과 안경이 참조 이미지와 일치하나 오버코트를 입지 않음.",
    "hard_violations": [],
    "physics": "윤성찬은 무릎을 굽히고 다리로 안정적으로 상체를 지탱하며, 국방장관은 소파에 기대어 앉아 있음."
   },
   {
    "label": "B",
    "direction": "윤성찬의 시선이 테이블 건너편에 앉은 국방장관을 향하고 있음.",
    "built_space": "집무실 내부는 구현되었으나, 이전 샷에서 왼쪽 소파에 있던 국방장관이 오른쪽 의자로 위치를 크게 이동함.",
    "entities": "두 인물의 얼굴이 참조와 일치하며, 윤성찬이 참조 이미지에 맞게 오버코트를 착용함.",
    "hard_violations": [],
    "physics": "두 인물 모두 바닥과 가구를 통해 물리적으로 무리 없이 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "샷 텍스트에 명시된 '속삭이듯 입을 반쯤 벌린' 표정과 밀착된 상체의 행동을 훌륭하게 연출하여 높은 우선순위 요건을 충족함 (오버코트 누락은 후순위 감점)."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "윤성찬의 의상 참조는 충실히 따랐으나, 핵심 지시인 '속삭이는' 행동을 누락하고 인물 간의 거리를 멀게 배치하여 최우선 연출 지시를 어김."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬의 시선과 앞으로 쑥 내민 상체가 국방장관의 얼굴과 귀를 정확히 향하고 있음.",
        "built_space": "이전 샷에 확립된 왼쪽 소파와 오른쪽 책상의 배치를 정확히 유지하며 주간의 채광을 잘 반영함.",
        "entities": "국방장관은 이전 샷의 복장과 자세를 유지함. 윤성찬은 얼굴과 안경이 참조 이미지와 일치하나 오버코트를 입지 않음.",
        "hard_violations": [],
        "physics": "윤성찬은 무릎을 굽히고 다리로 안정적으로 상체를 지탱하며, 국방장관은 소파에 기대어 앉아 있음."
       },
       {
        "label": "B",
        "direction": "윤성찬의 시선이 테이블 건너편에 앉은 국방장관을 향하고 있음.",
        "built_space": "집무실 내부는 구현되었으나, 이전 샷에서 왼쪽 소파에 있던 국방장관이 오른쪽 의자로 위치를 크게 이동함.",
        "entities": "두 인물의 얼굴이 참조와 일치하며, 윤성찬이 참조 이미지에 맞게 오버코트를 착용함.",
        "hard_violations": [],
        "physics": "두 인물 모두 바닥과 가구를 통해 물리적으로 무리 없이 지탱되고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "샷 텍스트에 명시된 '속삭이듯 입을 반쯤 벌린' 표정과 밀착된 상체의 행동을 훌륭하게 연출하여 높은 우선순위 요건을 충족함 (오버코트 누락은 후순위 감점)."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "윤성찬의 의상 참조는 충실히 따랐으나, 핵심 지시인 '속삭이는' 행동을 누락하고 인물 간의 거리를 멀게 배치하여 최우선 연출 지시를 어김."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬의 시선과 앞으로 쑥 내민 상체가 국방장관의 얼굴과 귀를 정확히 향하고 있음.",
        "built_space": "이전 샷에 확립된 왼쪽 소파와 오른쪽 책상의 배치를 정확히 유지하며 주간의 채광을 잘 반영함.",
        "entities": "국방장관은 이전 샷의 복장과 자세를 유지함. 윤성찬은 얼굴과 안경이 참조 이미지와 일치하나 오버코트를 입지 않음.",
        "hard_violations": [],
        "physics": "윤성찬은 무릎을 굽히고 다리로 안정적으로 상체를 지탱하며, 국방장관은 소파에 기대어 앉아 있음."
       },
       {
        "label": "B",
        "direction": "윤성찬의 시선이 테이블 건너편에 앉은 국방장관을 향하고 있음.",
        "built_space": "집무실 내부는 구현되었으나, 이전 샷에서 왼쪽 소파에 있던 국방장관이 오른쪽 의자로 위치를 크게 이동함.",
        "entities": "두 인물의 얼굴이 참조와 일치하며, 윤성찬이 참조 이미지에 맞게 오버코트를 착용함.",
        "hard_violations": [],
        "physics": "두 인물 모두 바닥과 가구를 통해 물리적으로 무리 없이 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "윤성찬의 굳은 상체와 장관을 향한 근접한 속삭임을 미디엄 숏에 더 가깝게 담았으나, 참조의 짙은 외투와 어깨끈이 빠졌다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "장관을 향한 전진 자세와 외투는 충실하지만, 하체와 공간까지 넓게 담은 측면 투 숏이라 상체 중심 구도와 두 눈의 긴장감이 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 윤성찬이 오른쪽에 앉은 장관의 얼굴을 향해 몸과 고개를 내밀고 시선을 맞춘다. 장관도 윤성찬을 바라본다. 입은 벌어져 있지만 윤성찬의 얼굴이 거의 옆모습이어서 두 눈의 날카로운 시선을 함께 확인하기 어렵다.",
        "built_space": "왼쪽 가죽 소파 한 개, 오른쪽 목재 팔걸이 가죽 의자 한 개 사이에 대화용 탁자 한 개가 있다. 뒤에는 창 구역 두 개와 중앙 지도 액자 한 개, 목재 수납장, 오른쪽 업무 책상 한 개와 업무용 의자 한 개, 책장 한 개, 깃발 두 개가 보여 참조 사무실의 주요 구성을 유지한다. 장관은 참조의 왼쪽 소파가 아니라 맞은편 의자에 앉아 있으나, 움직일 수 없는 인물로 지정되지는 않았다. 낮의 창빛이 보인다.",
        "entities": "윤성찬은 회색 머리와 안경, 깊은 얼굴 주름을 가진 고령의 한국인 남성으로 보이며 짙은 외투와 줄무늬 정장, 흰 셔츠와 넥타이가 참조와 대체로 맞는다. 참조의 어깨끈은 보이지 않는다. 장관은 단정한 검은 머리의 한국인 성인 남성으로, 이전 장면의 회색 정장과 흰 셔츠를 유지한다. 탁자에는 책과 화분이 있으며 뚜렷하게 읽히는 글자는 없다. 인물은 두 명뿐이다.",
        "hard_violations": [],
        "physics": "윤성찬은 탁자 위에 양 주먹을 대어 상체의 전진 하중을 지지한다. 하체는 아래로 이어지고 발은 화면 밖이므로 공중에 뜬 자세가 아니다. 장관은 의자 좌판에 앉아 손을 무릎 부근에 두고 있다. 책과 화분은 탁자에 놓여 있으며 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "오른쪽 윤성찬의 얼굴과 상체가 왼쪽 장관에게 가까이 기울어져 있고, 시선은 장관의 눈 부근을 향한다. 장관도 윤성찬 쪽을 바라본다. 윤성찬은 입을 조금 벌린 채 얼굴을 가까이 대어 낮게 속삭이는 순간으로 읽힌다. 눈은 자연스러운 인간의 눈이며 안경 너머로 집중된 표정이 보인다.",
        "built_space": "장관은 전경의 검은 가죽 소파에 앉아 있고 윤성찬은 그 오른쪽 옆에 서서 몸을 숙인다. 뒤쪽에는 가죽 좌석 일부 두 곳과 대화용 탁자 한 개가 보인다. 창 구역 두 개, 중앙 지도 액자 한 개, 왼쪽 산 그림 한 개와 스탠드 한 개, 오른쪽 업무 책상 한 개와 의자 한 개, 책장 한 개, 깃발 두 개가 참조의 배치와 재료를 대체로 따른다. 뒤쪽 왼편 좌석은 참조보다 더 드러나지만 중복된 고정 설비라고 단정할 근거는 부족하다. 창밖과 실내는 낮으로 표현되어 있다.",
        "entities": "윤성찬은 회색 머리, 안경, 이마와 눈가의 주름을 가진 고령 한국인 남성으로 참조 인상에 가깝다. 흰 셔츠와 넥타이, 줄무늬 조끼 및 바지는 보이지만 참조의 짙은 외투와 어깨끈은 없다. 장관의 얼굴과 짧은 검은 머리, 회색 정장 및 흰 셔츠는 이전 장면과 잘 연결된다. 장관이 귀에 댄 전화기는 이전 장면에도 있던 물건이다. 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "윤성찬은 골반과 다리가 화면 아래로 이어진 선 자세에서 허리를 굽힌 것으로 보인다. 발의 접지는 화면 밖이지만 상체가 하체와 자연스럽게 연결되어 있어 지지 없이 떠 있는 형상은 아니다. 장관의 엉덩이와 허벅지는 소파 좌판에 지지되고 전화기는 손으로 잡아 귀에 댄다. 탁자 위 물건과 책상 위 기기에도 받치는 표면이 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "윤성찬의 굳은 상체와 장관을 향한 근접한 속삭임을 미디엄 숏에 더 가깝게 담았으나, 참조의 짙은 외투와 어깨끈이 빠졌다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "장관을 향한 전진 자세와 외투는 충실하지만, 하체와 공간까지 넓게 담은 측면 투 숏이라 상체 중심 구도와 두 눈의 긴장감이 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 윤성찬이 오른쪽에 앉은 장관의 얼굴을 향해 몸과 고개를 내밀고 시선을 맞춘다. 장관도 윤성찬을 바라본다. 입은 벌어져 있지만 윤성찬의 얼굴이 거의 옆모습이어서 두 눈의 날카로운 시선을 함께 확인하기 어렵다.",
        "built_space": "왼쪽 가죽 소파 한 개, 오른쪽 목재 팔걸이 가죽 의자 한 개 사이에 대화용 탁자 한 개가 있다. 뒤에는 창 구역 두 개와 중앙 지도 액자 한 개, 목재 수납장, 오른쪽 업무 책상 한 개와 업무용 의자 한 개, 책장 한 개, 깃발 두 개가 보여 참조 사무실의 주요 구성을 유지한다. 장관은 참조의 왼쪽 소파가 아니라 맞은편 의자에 앉아 있으나, 움직일 수 없는 인물로 지정되지는 않았다. 낮의 창빛이 보인다.",
        "entities": "윤성찬은 회색 머리와 안경, 깊은 얼굴 주름을 가진 고령의 한국인 남성으로 보이며 짙은 외투와 줄무늬 정장, 흰 셔츠와 넥타이가 참조와 대체로 맞는다. 참조의 어깨끈은 보이지 않는다. 장관은 단정한 검은 머리의 한국인 성인 남성으로, 이전 장면의 회색 정장과 흰 셔츠를 유지한다. 탁자에는 책과 화분이 있으며 뚜렷하게 읽히는 글자는 없다. 인물은 두 명뿐이다.",
        "hard_violations": [],
        "physics": "윤성찬은 탁자 위에 양 주먹을 대어 상체의 전진 하중을 지지한다. 하체는 아래로 이어지고 발은 화면 밖이므로 공중에 뜬 자세가 아니다. 장관은 의자 좌판에 앉아 손을 무릎 부근에 두고 있다. 책과 화분은 탁자에 놓여 있으며 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "오른쪽 윤성찬의 얼굴과 상체가 왼쪽 장관에게 가까이 기울어져 있고, 시선은 장관의 눈 부근을 향한다. 장관도 윤성찬 쪽을 바라본다. 윤성찬은 입을 조금 벌린 채 얼굴을 가까이 대어 낮게 속삭이는 순간으로 읽힌다. 눈은 자연스러운 인간의 눈이며 안경 너머로 집중된 표정이 보인다.",
        "built_space": "장관은 전경의 검은 가죽 소파에 앉아 있고 윤성찬은 그 오른쪽 옆에 서서 몸을 숙인다. 뒤쪽에는 가죽 좌석 일부 두 곳과 대화용 탁자 한 개가 보인다. 창 구역 두 개, 중앙 지도 액자 한 개, 왼쪽 산 그림 한 개와 스탠드 한 개, 오른쪽 업무 책상 한 개와 의자 한 개, 책장 한 개, 깃발 두 개가 참조의 배치와 재료를 대체로 따른다. 뒤쪽 왼편 좌석은 참조보다 더 드러나지만 중복된 고정 설비라고 단정할 근거는 부족하다. 창밖과 실내는 낮으로 표현되어 있다.",
        "entities": "윤성찬은 회색 머리, 안경, 이마와 눈가의 주름을 가진 고령 한국인 남성으로 참조 인상에 가깝다. 흰 셔츠와 넥타이, 줄무늬 조끼 및 바지는 보이지만 참조의 짙은 외투와 어깨끈은 없다. 장관의 얼굴과 짧은 검은 머리, 회색 정장 및 흰 셔츠는 이전 장면과 잘 연결된다. 장관이 귀에 댄 전화기는 이전 장면에도 있던 물건이다. 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "윤성찬은 골반과 다리가 화면 아래로 이어진 선 자세에서 허리를 굽힌 것으로 보인다. 발의 접지는 화면 밖이지만 상체가 하체와 자연스럽게 연결되어 있어 지지 없이 떠 있는 형상은 아니다. 장관의 엉덩이와 허벅지는 소파 좌판에 지지되고 전화기는 손으로 잡아 귀에 댄다. 탁자 위 물건과 책상 위 기기에도 받치는 표면이 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.446
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.446
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1446
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "샷 텍스트에 명시된 '속삭이듯 입을 반쯤 벌린' 표정과 밀착된 상체의 행동을 훌륭하게 연출하여 높은 우선순위 요건을 충족함 (오버코트 누락은 후순위 감점)."
   },
   {
    "label": "B",
    "score": 1446,
    "verdict_ko": "윤성찬의 의상 참조는 충실히 따랐으나, 핵심 지시인 '속삭이는' 행동을 누락하고 인물 간의 거리를 멀게 배치하여 최우선 연출 지시를 어김."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 국방장관 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2_sel.png",
    "asset_id": "88e93f27-8ef8-49bf-a845-65a1b73f1826",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1183281>",
    "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-044c-7ee5-b06a-ea4b8fe202ae",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S43sh2"
  },
  "staged_characters_added": [
   "C21"
  ]
 },
 "S54sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:29:37.288564+00:00",
  "fingerprint": "e0edd1b0c01c3487507c186708b955f932f9d910d5a59a6343de68a3365960bd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S54sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S54sh2_sel.png",
  "source_sha256": "c86ab7a49ef47430f4ca9129103229c71766a6381b8e268ae8ea95284ce7842d",
  "file": "S54sh2_cine.png",
  "staged_sha256": "54af89fbb19783a73a013024e2330650966675d7dd8a3f0efa20eed1ba43c4d1",
  "latency_ms": 8934
 },
 "S54sh5::signage": {
  "fp": "9b9ae43c0e77a6db",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S54sh5": {
  "input_fingerprint": "ab4a90f0cc042ea9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 미소를 지운 채 목적지를 안다며 날카로운 눈빛으로 허공을 꿰뚫어 보는 윤성찬의 차가운 얼굴 클로즈업.\n\nLOCATION (lock): In the meeting area inside the defense minister's executive office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's established ambient illumination and restrained contrast, conveying coldness through expression rather than a new color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the minister's office surfaces, nearby furnishings, and daytime interior illumination. Exclude the militia commander's office furnishings and telephone-conversation setting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 미소를 지운 채 목적지를 안다며 날카로운 눈빛으로 허공을 꿰뚫어 보는 윤성찬의 차가운 얼굴 클로즈업.\n\nLOCATION (lock): In the meeting area inside the defense minister's executive office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's established ambient illumination and restrained contrast, conveying coldness through expression rather than a new color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the minister's office surfaces, nearby furnishings, and daytime interior illumination. Exclude the militia commander's office furnishings and telephone-conversation setting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 미소를 지운 채 목적지를 안다며 날카로운 눈빛으로 허공을 꿰뚫어 보는 윤성찬의 차가운 얼굴 클로즈업.\n\nLOCATION (lock): In the meeting area inside the defense minister's executive office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's established ambient illumination and restrained contrast, conveying coldness through expression rather than a new color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the minister's office surfaces, nearby furnishings, and daytime interior illumination. Exclude the militia commander's office furnishings and telephone-conversation setting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물은 프레임 왼쪽 허공을 바라봄.",
    "built_space": "검은색 가죽 소파, 테이블, 화분 등이 있으나 이전 샷의 배경 구도와 완전히 일치하지는 않음.",
    "entities": "윤성찬의 얼굴과 셔츠, 조끼는 부합하나 레퍼런스에 있던 넥타이가 없음.",
    "hard_violations": [],
    "physics": "자연스러운 자세로 화면 내에 안정적으로 위치함."
   },
   {
    "label": "B",
    "direction": "인물은 프레임 왼쪽 아래 방향의 허공을 날카롭게 응시함.",
    "built_space": "소파, 램프, 창문, 배경의 책상 등 이전 샷의 장관 집무실 환경과 배치가 정확하게 일치함.",
    "entities": "윤성찬의 얼굴, 안경, 셔츠, 조끼, 그리고 지정된 도트 무늬 넥타이까지 모두 일치함.",
    "hard_violations": [],
    "physics": "화면 밖을 향해 상체를 약간 기울인 자세가 자연스럽게 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "이전 샷의 의상(넥타이 포함)과 집무실 배경 요소를 정확하게 유지하며 지시된 날카로운 표정의 클로즈업을 훌륭히 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물의 표정과 클로즈업 숏은 무난하나, 고정 조건인 넥타이가 누락되었고 배경의 가구 배치가 레퍼런스와 다름."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 프레임 왼쪽 허공을 바라봄.",
        "built_space": "검은색 가죽 소파, 테이블, 화분 등이 있으나 이전 샷의 배경 구도와 완전히 일치하지는 않음.",
        "entities": "윤성찬의 얼굴과 셔츠, 조끼는 부합하나 레퍼런스에 있던 넥타이가 없음.",
        "hard_violations": [],
        "physics": "자연스러운 자세로 화면 내에 안정적으로 위치함."
       },
       {
        "label": "B",
        "direction": "인물은 프레임 왼쪽 아래 방향의 허공을 날카롭게 응시함.",
        "built_space": "소파, 램프, 창문, 배경의 책상 등 이전 샷의 장관 집무실 환경과 배치가 정확하게 일치함.",
        "entities": "윤성찬의 얼굴, 안경, 셔츠, 조끼, 그리고 지정된 도트 무늬 넥타이까지 모두 일치함.",
        "hard_violations": [],
        "physics": "화면 밖을 향해 상체를 약간 기울인 자세가 자연스럽게 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "이전 샷의 의상(넥타이 포함)과 집무실 배경 요소를 정확하게 유지하며 지시된 날카로운 표정의 클로즈업을 훌륭히 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물의 표정과 클로즈업 숏은 무난하나, 고정 조건인 넥타이가 누락되었고 배경의 가구 배치가 레퍼런스와 다름."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 프레임 왼쪽 허공을 바라봄.",
        "built_space": "검은색 가죽 소파, 테이블, 화분 등이 있으나 이전 샷의 배경 구도와 완전히 일치하지는 않음.",
        "entities": "윤성찬의 얼굴과 셔츠, 조끼는 부합하나 레퍼런스에 있던 넥타이가 없음.",
        "hard_violations": [],
        "physics": "자연스러운 자세로 화면 내에 안정적으로 위치함."
       },
       {
        "label": "B",
        "direction": "인물은 프레임 왼쪽 아래 방향의 허공을 날카롭게 응시함.",
        "built_space": "소파, 램프, 창문, 배경의 책상 등 이전 샷의 장관 집무실 환경과 배치가 정확하게 일치함.",
        "entities": "윤성찬의 얼굴, 안경, 셔츠, 조끼, 그리고 지정된 도트 무늬 넥타이까지 모두 일치함.",
        "hard_violations": [],
        "physics": "화면 밖을 향해 상체를 약간 기울인 자세가 자연스럽게 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴을 크게 담은 클로즈업과 화면 밖을 꿰뚫는 차가운 시선이 지문에 부합하며, 이전 장면의 안경·조끼·넥타이와 낮 실내광도 유지한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "인물과 집무실의 연속성은 대체로 맞지만, 가슴까지 넓힌 구도가 얼굴 클로즈업보다 느슨하고 시선의 날카로움이 약하며 이전 장면의 넥타이가 사라졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 화면 왼쪽으로 비스듬히 향하고, 눈은 렌즈 왼쪽의 화면 밖 공간을 응시한다. 보이는 특정 사람이나 물건을 바라보는 것이 아니라 허공의 한 지점에 집중하는 모습이다. 입을 다물고 미소를 거둔 표정이 날카로운 시선과 결합한다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "왼쪽에 검은 가죽 소파 하나, 뒤쪽 중앙에 안락의자 하나, 왼쪽 아래에 다른 좌석 일부가 보인다. 중앙에는 낮은 목재 상판 두 면이 나란히 드러나며, 오른쪽 뒤에는 집무용 책상과 모니터 하나, 책장 일부가 있다. 왼쪽 벽의 그림 하나와 스탠드 하나, 큰 화분 하나, 뒤쪽 블라인드 창들이 이전 집무실의 재료와 배치를 대체로 잇는다. 상판들의 연결 관계는 인물과 크롭 때문에 확정하기 어렵다. 인물은 회의 가구 앞쪽에서 몸을 기울이고 있고, 얼굴이 화면 높이 대부분을 차지한다. 불가능한 반사나 명백히 중복된 고정 설비는 보이지 않는다.",
        "entities": "보이는 사람은 윤성찬에 해당하는 고령의 한국인 남성 한 명뿐이다. 깊은 이마 주름과 눈가 주름, 뒤로 넘긴 회색 머리, 가는 금속테 안경, 얼굴 윤곽이 참고 인물과 가깝다. 흰 셔츠, 남색 줄무늬 조끼, 어두운 무늬 넥타이가 이전 장면과 일치한다. 다른 인물이나 전화기는 없고, 읽을 수 있는 글자도 없다. 낮의 창빛과 따뜻한 스탠드 빛이 함께 남아 있으며 새로운 청색 색조로 냉기를 표현하지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 있고 상체는 앞으로 약간 기울어 있다. 하체와 발의 지지는 클로즈업 밖이므로 확인할 수 없지만, 몸이 공중에 떠 있다는 증거는 없다. 안경은 코와 귀에 걸려 있고 옷은 몸을 따라 접힌다. 배경 소품들은 탁자나 가구 상판에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이며 눈동자는 렌즈 가까운 화면 왼쪽을 향한다. 미소 없이 입을 다물었지만, 먼 허공을 꿰뚫는 집중보다는 카메라 근처를 조용히 살피는 인상이 강하다. 시선이 향하는 구체적인 대상은 화면에 없다. 무기나 방향성 있는 휴대 소품은 없다.",
        "built_space": "왼쪽 뒤에 검은 가죽 소파 하나, 오른쪽 뒤에 안락의자 하나, 오른쪽 아래에 좌석 일부가 있다. 앞뒤로 놓인 목재 회의 탁자 하나와 보조 탁자 하나, 벽 그림 일부 하나, 스탠드 하나, 큰 화분 하나, 창과 낮은 수납장 일부가 보인다. 회의 탁자 위에는 책 또는 문서철 두 개와 쟁반, 작은 화분이 있다. 가죽·목재·벽면과 낮 조명은 참고 장소에 대체로 맞는다. 인물은 회의 좌석 영역에 있으나 가슴 아래까지 포함되어 요청된 얼굴 중심 클로즈업보다 넓다. 명백한 설비 중복이나 불가능한 반사는 없다.",
        "entities": "고령의 한국인 남성 한 명만 보이며 회색 머리, 금속테 안경, 이마와 눈가의 주름, 얼굴 생김새가 윤성찬 참고와 가깝다. 흰 셔츠와 남색 줄무늬 조끼는 유지되지만 셔츠 목이 열려 있고 이전 장면의 넥타이가 없다. 다른 인물과 전화기는 등장하지 않는다. 뒤쪽 상패와 책 표면은 흐려 읽을 수 없으며 자막이나 로고도 없다.",
        "hard_violations": [],
        "physics": "상체는 거의 세워져 있고 머리와 목, 어깨의 연결은 자연스럽다. 하체와 좌석 접촉점은 프레임 밖이어서 착석 여부를 확정할 수 없지만 공중 부양이나 비현실적인 자세는 보이지 않는다. 안경은 코와 귀가 지지하고, 책과 쟁반은 탁자 상판에 놓여 있으며 화분은 가구 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴을 크게 담은 클로즈업과 화면 밖을 꿰뚫는 차가운 시선이 지문에 부합하며, 이전 장면의 안경·조끼·넥타이와 낮 실내광도 유지한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "인물과 집무실의 연속성은 대체로 맞지만, 가슴까지 넓힌 구도가 얼굴 클로즈업보다 느슨하고 시선의 날카로움이 약하며 이전 장면의 넥타이가 사라졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 화면 왼쪽으로 비스듬히 향하고, 눈은 렌즈 왼쪽의 화면 밖 공간을 응시한다. 보이는 특정 사람이나 물건을 바라보는 것이 아니라 허공의 한 지점에 집중하는 모습이다. 입을 다물고 미소를 거둔 표정이 날카로운 시선과 결합한다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "왼쪽에 검은 가죽 소파 하나, 뒤쪽 중앙에 안락의자 하나, 왼쪽 아래에 다른 좌석 일부가 보인다. 중앙에는 낮은 목재 상판 두 면이 나란히 드러나며, 오른쪽 뒤에는 집무용 책상과 모니터 하나, 책장 일부가 있다. 왼쪽 벽의 그림 하나와 스탠드 하나, 큰 화분 하나, 뒤쪽 블라인드 창들이 이전 집무실의 재료와 배치를 대체로 잇는다. 상판들의 연결 관계는 인물과 크롭 때문에 확정하기 어렵다. 인물은 회의 가구 앞쪽에서 몸을 기울이고 있고, 얼굴이 화면 높이 대부분을 차지한다. 불가능한 반사나 명백히 중복된 고정 설비는 보이지 않는다.",
        "entities": "보이는 사람은 윤성찬에 해당하는 고령의 한국인 남성 한 명뿐이다. 깊은 이마 주름과 눈가 주름, 뒤로 넘긴 회색 머리, 가는 금속테 안경, 얼굴 윤곽이 참고 인물과 가깝다. 흰 셔츠, 남색 줄무늬 조끼, 어두운 무늬 넥타이가 이전 장면과 일치한다. 다른 인물이나 전화기는 없고, 읽을 수 있는 글자도 없다. 낮의 창빛과 따뜻한 스탠드 빛이 함께 남아 있으며 새로운 청색 색조로 냉기를 표현하지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 있고 상체는 앞으로 약간 기울어 있다. 하체와 발의 지지는 클로즈업 밖이므로 확인할 수 없지만, 몸이 공중에 떠 있다는 증거는 없다. 안경은 코와 귀에 걸려 있고 옷은 몸을 따라 접힌다. 배경 소품들은 탁자나 가구 상판에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이며 눈동자는 렌즈 가까운 화면 왼쪽을 향한다. 미소 없이 입을 다물었지만, 먼 허공을 꿰뚫는 집중보다는 카메라 근처를 조용히 살피는 인상이 강하다. 시선이 향하는 구체적인 대상은 화면에 없다. 무기나 방향성 있는 휴대 소품은 없다.",
        "built_space": "왼쪽 뒤에 검은 가죽 소파 하나, 오른쪽 뒤에 안락의자 하나, 오른쪽 아래에 좌석 일부가 있다. 앞뒤로 놓인 목재 회의 탁자 하나와 보조 탁자 하나, 벽 그림 일부 하나, 스탠드 하나, 큰 화분 하나, 창과 낮은 수납장 일부가 보인다. 회의 탁자 위에는 책 또는 문서철 두 개와 쟁반, 작은 화분이 있다. 가죽·목재·벽면과 낮 조명은 참고 장소에 대체로 맞는다. 인물은 회의 좌석 영역에 있으나 가슴 아래까지 포함되어 요청된 얼굴 중심 클로즈업보다 넓다. 명백한 설비 중복이나 불가능한 반사는 없다.",
        "entities": "고령의 한국인 남성 한 명만 보이며 회색 머리, 금속테 안경, 이마와 눈가의 주름, 얼굴 생김새가 윤성찬 참고와 가깝다. 흰 셔츠와 남색 줄무늬 조끼는 유지되지만 셔츠 목이 열려 있고 이전 장면의 넥타이가 없다. 다른 인물과 전화기는 등장하지 않는다. 뒤쪽 상패와 책 표면은 흐려 읽을 수 없으며 자막이나 로고도 없다.",
        "hard_violations": [],
        "physics": "상체는 거의 세워져 있고 머리와 목, 어깨의 연결은 자연스럽다. 하체와 좌석 접촉점은 프레임 밖이어서 착석 여부를 확정할 수 없지만 공중 부양이나 비현실적인 자세는 보이지 않는다. 안경은 코와 귀가 지지하고, 책과 쟁반은 탁자 상판에 놓여 있으며 화분은 가구 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.333,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.333,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1333
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "이전 샷의 의상(넥타이 포함)과 집무실 배경 요소를 정확하게 유지하며 지시된 날카로운 표정의 클로즈업을 훌륭히 구현함."
   },
   {
    "label": "A",
    "score": 1333,
    "verdict_ko": "인물의 표정과 클로즈업 숏은 무난하나, 고정 조건인 넥타이가 누락되었고 배경의 가구 배치가 레퍼런스와 다름."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S54sh2_sel.png",
    "asset_id": "023d6fcd-3454-4c7c-81b4-22ea316851fa",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-05f7-7425-8eda-173fc692f589",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S54sh2"
  }
 },
 "S54sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:30:34.933181+00:00",
  "fingerprint": "e570c24ec16c7044876b05ffc0906fca0ba4e0686d2e193a6cda081260020fc7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S54sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S54sh5_sel.png",
  "source_sha256": "d0918cf2d7372e56a8e594692b34d05bbb022ed029ab2f8b78d1f981e7c0ae52",
  "file": "S54sh5_cine.png",
  "staged_sha256": "2953a7e6bc80f9114f376c8c1b609b3235c290dce1bc389f3759c23fbaa9e710",
  "latency_ms": 11006
 },
 "S54sh6::signage": {
  "fp": "769f5b928f054d73",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S54sh6": {
  "input_fingerprint": "424295be97ff7c1d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 윤성찬의 말에 두 눈을 동그랗게 뜬 채 놀란 표정으로 입을 벌린 국방장관의 어리둥절한 얼굴.\n\nLOCATION (lock): At the minister's position in the executive office's conversation area, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office ambient illumination keeps the minister's startled eyes and parted mouth clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same ministerial office surfaces, furnishings, and daylight conditions. Exclude the militia office's wall knife and speakerphone setup.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 윤성찬의 말에 두 눈을 동그랗게 뜬 채 놀란 표정으로 입을 벌린 국방장관의 어리둥절한 얼굴.\n\nLOCATION (lock): At the minister's position in the executive office's conversation area, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office ambient illumination keeps the minister's startled eyes and parted mouth clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same ministerial office surfaces, furnishings, and daylight conditions. Exclude the militia office's wall knife and speakerphone setup.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 윤성찬의 말에 두 눈을 동그랗게 뜬 채 놀란 표정으로 입을 벌린 국방장관의 어리둥절한 얼굴.\n\nLOCATION (lock): At the minister's position in the executive office's conversation area, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office ambient illumination keeps the minister's startled eyes and parted mouth clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same ministerial office surfaces, furnishings, and daylight conditions. Exclude the militia office's wall knife and speakerphone setup.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "국방장관의 시선이 카메라 렌즈를 정면으로 향하고 있음.",
    "built_space": "낮의 자연광이 드는 사무실. 배경의 소파, 테이블, 창문 배치가 이전 샷의 공간과 일치함.",
    "entities": "국방장관의 얼굴과 복장이 레퍼런스와 일치하며, 지시문대로 눈을 크게 뜨고 입을 벌린 표정임.",
    "hard_violations": [],
    "physics": "특별한 물리적 오류나 지지되지 않은 객체 없음."
   },
   {
    "label": "B",
    "direction": "국방장관의 시선이 화면 우측 전경에 있는 윤성찬을 정확히 향하고 있음.",
    "built_space": "낮의 자연광이 드는 사무실. 배경의 가구와 창문이 이전 샷과 일치하며, 대화 구도에 맞게 인물들이 배치됨.",
    "entities": "국방장관의 놀란 표정과 외모가 일치하며, 화면 전경에 걸친 윤성찬의 뒷모습(안경, 흰 셔츠, 줄무늬 조끼) 또한 이전 샷과 완벽히 일치함.",
    "hard_violations": [
     "[gpt-high] 장관의 반응 얼굴만 보여주는 샷에 윤성찬의 머리와 상체를 전경 인물로 추가했다."
    ],
    "physics": "대화 중인 인물들의 자세와 위치가 자연스러우며 물리적 오류 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "놀란 표정은 잘 표현되었으나, 카메라 렌즈를 정면으로 응시하고 있어 지시문이 암시하는 대화 상황과 어울리지 않습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "화면 전경의 윤성찬을 바라보며 놀라는 국방장관의 표정을 자연스러운 시선과 정확한 클로즈업으로 훌륭하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "국방장관의 시선이 카메라 렌즈를 정면으로 향하고 있음.",
        "built_space": "낮의 자연광이 드는 사무실. 배경의 소파, 테이블, 창문 배치가 이전 샷의 공간과 일치함.",
        "entities": "국방장관의 얼굴과 복장이 레퍼런스와 일치하며, 지시문대로 눈을 크게 뜨고 입을 벌린 표정임.",
        "hard_violations": [],
        "physics": "특별한 물리적 오류나 지지되지 않은 객체 없음."
       },
       {
        "label": "B",
        "direction": "국방장관의 시선이 화면 우측 전경에 있는 윤성찬을 정확히 향하고 있음.",
        "built_space": "낮의 자연광이 드는 사무실. 배경의 가구와 창문이 이전 샷과 일치하며, 대화 구도에 맞게 인물들이 배치됨.",
        "entities": "국방장관의 놀란 표정과 외모가 일치하며, 화면 전경에 걸친 윤성찬의 뒷모습(안경, 흰 셔츠, 줄무늬 조끼) 또한 이전 샷과 완벽히 일치함.",
        "hard_violations": [],
        "physics": "대화 중인 인물들의 자세와 위치가 자연스러우며 물리적 오류 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "놀란 표정은 잘 표현되었으나, 카메라 렌즈를 정면으로 응시하고 있어 지시문이 암시하는 대화 상황과 어울리지 않습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "화면 전경의 윤성찬을 바라보며 놀라는 국방장관의 표정을 자연스러운 시선과 정확한 클로즈업으로 훌륭하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "국방장관의 시선이 카메라 렌즈를 정면으로 향하고 있음.",
        "built_space": "낮의 자연광이 드는 사무실. 배경의 소파, 테이블, 창문 배치가 이전 샷의 공간과 일치함.",
        "entities": "국방장관의 얼굴과 복장이 레퍼런스와 일치하며, 지시문대로 눈을 크게 뜨고 입을 벌린 표정임.",
        "hard_violations": [],
        "physics": "특별한 물리적 오류나 지지되지 않은 객체 없음."
       },
       {
        "label": "B",
        "direction": "국방장관의 시선이 화면 우측 전경에 있는 윤성찬을 정확히 향하고 있음.",
        "built_space": "낮의 자연광이 드는 사무실. 배경의 가구와 창문이 이전 샷과 일치하며, 대화 구도에 맞게 인물들이 배치됨.",
        "entities": "국방장관의 놀란 표정과 외모가 일치하며, 화면 전경에 걸친 윤성찬의 뒷모습(안경, 흰 셔츠, 줄무늬 조끼) 또한 이전 샷과 완벽히 일치함.",
        "hard_violations": [],
        "physics": "대화 중인 인물들의 자세와 위치가 자연스러우며 물리적 오류 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "놀란 표정과 윤성찬을 향한 시선은 명확하지만, 장관 얼굴 단독 클로즈업에 윤성찬의 머리와 상체를 추가하여 어깨너머 구도로 바꿨다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "장관만의 얼굴 클로즈업, 동그랗게 뜬 눈과 벌어진 입, 기존 집무실의 주간 환경을 충실히 구현했으나 시선은 렌즈 정면에 가깝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장관의 얼굴과 두 눈은 화면 오른쪽 전경의 윤성찬 얼굴을 향한다. 윤성찬도 장관 쪽으로 고개를 돌리고 있어 대화 상대를 바라보는 관계가 명확하다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "왼쪽 가죽 소파 한 개, 뒤쪽 안락의자 한 개, 중앙 낮은 탁자 한 개, 왼쪽 보조 탁자와 스탠드 각 한 개, 벽 그림 한 개, 큰 화분 한 개와 후면 창들이 보인다. 목재와 가죽, 창으로 들어오는 낮빛은 이전 장면과 유사하다. 장관은 대화 공간 앞쪽에 있고 윤성찬은 오른쪽 전경을 크게 차지한다. 반사상이나 명백한 고정 설비 중복은 없다.",
        "entities": "장관은 짧게 정돈한 검은 머리의 한국인 성인 남성으로 보이며, 얼굴과 체형, 짙은 회색 정장·흰 셔츠·무늬 넥타이가 인물 참조와 대체로 맞는다. 눈을 크게 뜨고 입을 벌린 놀란 표정도 맞는다. 오른쪽에는 회색 머리와 안경, 흰 셔츠와 짙은 조끼를 착용한 윤성찬이 추가로 보이며 이전 장면의 외형을 따른다. 다만 권위 있는 샷 텍스트가 보여주도록 지정한 대상은 장관의 얼굴이다. 읽을 수 있는 글자, 벽 칼, 스피커폰은 보이지 않는다.",
        "hard_violations": [
         "장관의 반응 얼굴만 보여주는 샷에 윤성찬의 머리와 상체를 전경 인물로 추가했다."
        ],
        "physics": "두 사람의 머리는 목과 상체에 자연스럽게 연결되어 있고 서로 겹치는 부분도 전후 거리로 설명된다. 하체와 좌면 접촉은 클로즈업 밖이므로 앉거나 서 있는 상태를 확정할 수 없다. 공중에 뜬 신체나 지지 없는 물체는 보이지 않으며, 배경 가구와 소품은 바닥 또는 탁자 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "장관은 얼굴을 거의 정면으로 두고 렌즈 근처를 응시한다. 화면 밖 윤성찬에게 반응하는 정면 역숏으로 해석할 수 있지만, 실제 상대의 위치가 보이지 않아 시선이 윤성찬에게 도달하는지는 확인되지 않는다. 무기나 가리키는 소품은 없다.",
        "built_space": "왼쪽 가죽 소파 한 개와 전경 좌석 일부, 뒤쪽 안락의자 한 개, 중앙 낮은 탁자 한 개, 왼쪽 스탠드와 보조 탁자 각 한 개, 벽 그림 한 개, 큰 화분 한 개가 보인다. 후면 창과 목재 수납장, 오른쪽 업무용 책상 한 개·모니터 한 개·선반이 이전 장면의 배치와 재질을 유지한다. 장관은 대화 공간 앞쪽에서 얼굴과 어깨만 크게 잡혀 있다. 불가능한 반사나 명백히 중복된 설비는 없다.",
        "entities": "한 명의 한국인 성인 남성 장관만 보인다. 검은 짧은 머리, 얼굴 윤곽, 짙은 회색 정장과 흰 셔츠, 화면 아래 일부 보이는 넥타이가 인물 참조와 대체로 일치한다. 두 눈을 동그랗게 뜨고 입을 벌린 어리둥절한 표정이며, 홍채와 동공은 자연스러운 사람 눈으로 표현되어 있다. 윤성찬이나 다른 인물은 추가되지 않았다. 읽을 수 있는 글자와 금지된 칼·스피커폰도 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨의 연결 및 옷의 걸림이 자연스럽다. 하체와 좌석 접촉은 지정된 클로즈업 밖에 있어 지지 자세를 판정할 수 없지만, 부유를 시사하는 장면은 아니다. 배경 모니터는 책상 위에, 스탠드는 보조 탁자 위에 놓여 있고 가구는 바닥에 지지되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "놀란 표정과 윤성찬을 향한 시선은 명확하지만, 장관 얼굴 단독 클로즈업에 윤성찬의 머리와 상체를 추가하여 어깨너머 구도로 바꿨다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "장관만의 얼굴 클로즈업, 동그랗게 뜬 눈과 벌어진 입, 기존 집무실의 주간 환경을 충실히 구현했으나 시선은 렌즈 정면에 가깝다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "장관의 얼굴과 두 눈은 화면 오른쪽 전경의 윤성찬 얼굴을 향한다. 윤성찬도 장관 쪽으로 고개를 돌리고 있어 대화 상대를 바라보는 관계가 명확하다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "왼쪽 가죽 소파 한 개, 뒤쪽 안락의자 한 개, 중앙 낮은 탁자 한 개, 왼쪽 보조 탁자와 스탠드 각 한 개, 벽 그림 한 개, 큰 화분 한 개와 후면 창들이 보인다. 목재와 가죽, 창으로 들어오는 낮빛은 이전 장면과 유사하다. 장관은 대화 공간 앞쪽에 있고 윤성찬은 오른쪽 전경을 크게 차지한다. 반사상이나 명백한 고정 설비 중복은 없다.",
        "entities": "장관은 짧게 정돈한 검은 머리의 한국인 성인 남성으로 보이며, 얼굴과 체형, 짙은 회색 정장·흰 셔츠·무늬 넥타이가 인물 참조와 대체로 맞는다. 눈을 크게 뜨고 입을 벌린 놀란 표정도 맞는다. 오른쪽에는 회색 머리와 안경, 흰 셔츠와 짙은 조끼를 착용한 윤성찬이 추가로 보이며 이전 장면의 외형을 따른다. 다만 권위 있는 샷 텍스트가 보여주도록 지정한 대상은 장관의 얼굴이다. 읽을 수 있는 글자, 벽 칼, 스피커폰은 보이지 않는다.",
        "hard_violations": [
         "장관의 반응 얼굴만 보여주는 샷에 윤성찬의 머리와 상체를 전경 인물로 추가했다."
        ],
        "physics": "두 사람의 머리는 목과 상체에 자연스럽게 연결되어 있고 서로 겹치는 부분도 전후 거리로 설명된다. 하체와 좌면 접촉은 클로즈업 밖이므로 앉거나 서 있는 상태를 확정할 수 없다. 공중에 뜬 신체나 지지 없는 물체는 보이지 않으며, 배경 가구와 소품은 바닥 또는 탁자 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "장관은 얼굴을 거의 정면으로 두고 렌즈 근처를 응시한다. 화면 밖 윤성찬에게 반응하는 정면 역숏으로 해석할 수 있지만, 실제 상대의 위치가 보이지 않아 시선이 윤성찬에게 도달하는지는 확인되지 않는다. 무기나 가리키는 소품은 없다.",
        "built_space": "왼쪽 가죽 소파 한 개와 전경 좌석 일부, 뒤쪽 안락의자 한 개, 중앙 낮은 탁자 한 개, 왼쪽 스탠드와 보조 탁자 각 한 개, 벽 그림 한 개, 큰 화분 한 개가 보인다. 후면 창과 목재 수납장, 오른쪽 업무용 책상 한 개·모니터 한 개·선반이 이전 장면의 배치와 재질을 유지한다. 장관은 대화 공간 앞쪽에서 얼굴과 어깨만 크게 잡혀 있다. 불가능한 반사나 명백히 중복된 설비는 없다.",
        "entities": "한 명의 한국인 성인 남성 장관만 보인다. 검은 짧은 머리, 얼굴 윤곽, 짙은 회색 정장과 흰 셔츠, 화면 아래 일부 보이는 넥타이가 인물 참조와 대체로 일치한다. 두 눈을 동그랗게 뜨고 입을 벌린 어리둥절한 표정이며, 홍채와 동공은 자연스러운 사람 눈으로 표현되어 있다. 윤성찬이나 다른 인물은 추가되지 않았다. 읽을 수 있는 글자와 금지된 칼·스피커폰도 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨의 연결 및 옷의 걸림이 자연스럽다. 하체와 좌석 접촉은 지정된 클로즈업 밖에 있어 지지 자세를 판정할 수 없지만, 부유를 시사하는 장면은 아니다. 배경 모니터는 책상 위에, 스탠드는 보조 탁자 위에 놓여 있고 가구는 바닥에 지지되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.444
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.194
   },
   "violations": {
    "B": [
     "[gpt-high] 장관의 반응 얼굴만 보여주는 샷에 윤성찬의 머리와 상체를 전경 인물로 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1571,
   "B": 1194
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "놀란 표정은 잘 표현되었으나, 카메라 렌즈를 정면으로 응시하고 있어 지시문이 암시하는 대화 상황과 어울리지 않습니다."
   },
   {
    "label": "B",
    "score": 1194,
    "verdict_ko": "화면 전경의 윤성찬을 바라보며 놀라는 국방장관의 표정을 자연스러운 시선과 정확한 클로즈업으로 훌륭하게 구현했습니다.  ★위반: [gpt-high] 장관의 반응 얼굴만 보여주는 샷에 윤성찬의 머리와 상체를 전경 인물로 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S54sh5_sel.png",
    "asset_id": "2d99f90f-f18f-4140-9390-6107769a1a31",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:835328>",
    "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-0799-7231-a3ef-de88e3d8a030",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S54sh5"
  },
  "staged_characters_added": [
   "C18"
  ]
 },
 "S54sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:31:37.315925+00:00",
  "fingerprint": "64663494f3f5e13f0ad871a32f6a89b6f185bd5d1c6af6f9836ec20708cb7150",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S54sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S54sh6_sel.png",
  "source_sha256": "8337c092412cbcc35f0793be0e56bf1ac79a0e866de093eb01419602f2d0aa4b",
  "file": "S54sh6_cine.png",
  "staged_sha256": "57a9e80d446511676b756e04a3b45bf9b6be0081a9acaa7e89e2502bd0a3f6e9",
  "latency_ms": 10162
 },
 "S55sh3::signage": {
  "fp": "729584fef3dbf4c4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S55sh3::bgfirst_bg": {
  "input_fingerprint": "cf09efd0cb5df35a",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S55sh3__bgfirst_bg.png",
  "asset_id": "f40aa134-4d1d-43ab-80e4-59ce7d1bc555",
  "input_asset_ids": [
   "30ef55b8-86fa-49ef-81a1-ac257a19f52a",
   "4be906f8-5377-4122-9528-ee6aaccbd804"
  ]
 },
 "S55sh3": {
  "input_fingerprint": "d5057f330244e3e7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rain shower has recently passed, and the operational camper carries the luggage, added food and medicine supplied at the repair shop. 라울: He is riding in the camper, soaked and exhausted, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rain shower has recently passed, and the operational camper carries the luggage, added food and medicine supplied at the repair shop. 라울: He is riding in the camper, soaked and exhausted, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rain shower has recently passed, and the operational camper carries the luggage, added food and medicine supplied at the repair shop. 라울: He is riding in the camper, soaked and exhausted, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S55sh3__bgfirst_bg.png",
     "asset_id": "f40aa134-4d1d-43ab-80e4-59ce7d1bc555",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S55sh3.png",
     "asset_id": "30ef55b8-86fa-49ef-81a1-ac257a19f52a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
     "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "오른손 검지손가락을 오른쪽 창밖 숲을 향해 정확히 뻗고 있으며, 시선도 같은 방향을 향함.",
    "built_space": "통로, 측면 창문, 뒷좌석 등 전체적인 캠퍼 구조는 레퍼런스와 일치하나, 라울이 뒷좌석이 아닌 통로 바닥에 앉아 있음.",
    "entities": "10세 소년, 꽁지머리, 흙 묻고 젖은 옷차림 등 라울의 캐릭터 설정 및 비에 젖은 상태 지시와 잘 일치함.",
    "hard_violations": [
     "[gemini-pro] 지정된 공간(뒷좌석)이 아닌 바닥에 인물을 배치함 (스테이징 지시 위반)",
     "[gpt-high] 라울이 점유해야 하는 후방 소파를 비워 둔 채 인물을 앞쪽 통로의 다른 좌석 위치에 배치했다."
    ],
    "physics": "바닥에 엉덩이를 대고 안정적으로 앉아 있으며 뻗은 팔의 무게 중심이 자연스러움."
   },
   {
    "label": "B",
    "direction": "오른손 검지손가락을 측면의 큰 창밖으로 뻗고 있으며, 시선도 창밖을 향하고 있음.",
    "built_space": "뒷좌석 패턴의 소파가 거대한 창문과 평행하게 배치되어 위치 레퍼런스의 실제 구조(정면을 향한 뒷좌석)를 완전히 왜곡함.",
    "entities": "라울의 외형은 일치하나, 옷과 피부가 마른 상태로 묘사되어 비에 젖은 지친 모습이라는 지시를 어김.",
    "hard_violations": [
     "[gemini-pro] 위치 레퍼런스에 없는 거대한 창문을 생성하고 소파의 구조적 배치를 자의적으로 왜곡함 (물리적으로 불가능한 스테이징)"
    ],
    "physics": "소파 쿠션 위에 앉아 체중을 지탱하고 있으나, 뻗은 오른팔의 어깨 연결부가 어색하게 처리됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "공간 레퍼런스의 구조와 비에 젖은 무드는 잘 살렸으나, 인물을 지정된 뒷좌석이 아닌 바닥에 배치하는 치명적인 스테이징 오류를 범함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "비에 젖은 상태라는 지시를 무시하고 캠퍼 내부 구조를 완전히 왜곡하여 위치 레퍼런스와 스테이징 조건을 모두 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른손 검지손가락을 오른쪽 창밖 숲을 향해 정확히 뻗고 있으며, 시선도 같은 방향을 향함.",
        "built_space": "통로, 측면 창문, 뒷좌석 등 전체적인 캠퍼 구조는 레퍼런스와 일치하나, 라울이 뒷좌석이 아닌 통로 바닥에 앉아 있음.",
        "entities": "10세 소년, 꽁지머리, 흙 묻고 젖은 옷차림 등 라울의 캐릭터 설정 및 비에 젖은 상태 지시와 잘 일치함.",
        "hard_violations": [
         "지정된 공간(뒷좌석)이 아닌 바닥에 인물을 배치함 (스테이징 지시 위반)"
        ],
        "physics": "바닥에 엉덩이를 대고 안정적으로 앉아 있으며 뻗은 팔의 무게 중심이 자연스러움."
       },
       {
        "label": "B",
        "direction": "오른손 검지손가락을 측면의 큰 창밖으로 뻗고 있으며, 시선도 창밖을 향하고 있음.",
        "built_space": "뒷좌석 패턴의 소파가 거대한 창문과 평행하게 배치되어 위치 레퍼런스의 실제 구조(정면을 향한 뒷좌석)를 완전히 왜곡함.",
        "entities": "라울의 외형은 일치하나, 옷과 피부가 마른 상태로 묘사되어 비에 젖은 지친 모습이라는 지시를 어김.",
        "hard_violations": [
         "위치 레퍼런스에 없는 거대한 창문을 생성하고 소파의 구조적 배치를 자의적으로 왜곡함 (물리적으로 불가능한 스테이징)"
        ],
        "physics": "소파 쿠션 위에 앉아 체중을 지탱하고 있으나, 뻗은 오른팔의 어깨 연결부가 어색하게 처리됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "공간 레퍼런스의 구조와 비에 젖은 무드는 잘 살렸으나, 인물을 지정된 뒷좌석이 아닌 바닥에 배치하는 치명적인 스테이징 오류를 범함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "비에 젖은 상태라는 지시를 무시하고 캠퍼 내부 구조를 완전히 왜곡하여 위치 레퍼런스와 스테이징 조건을 모두 위반함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "오른손 검지손가락을 오른쪽 창밖 숲을 향해 정확히 뻗고 있으며, 시선도 같은 방향을 향함.",
        "built_space": "통로, 측면 창문, 뒷좌석 등 전체적인 캠퍼 구조는 레퍼런스와 일치하나, 라울이 뒷좌석이 아닌 통로 바닥에 앉아 있음.",
        "entities": "10세 소년, 꽁지머리, 흙 묻고 젖은 옷차림 등 라울의 캐릭터 설정 및 비에 젖은 상태 지시와 잘 일치함.",
        "hard_violations": [
         "지정된 공간(뒷좌석)이 아닌 바닥에 인물을 배치함 (스테이징 지시 위반)"
        ],
        "physics": "바닥에 엉덩이를 대고 안정적으로 앉아 있으며 뻗은 팔의 무게 중심이 자연스러움."
       },
       {
        "label": "B",
        "direction": "오른손 검지손가락을 측면의 큰 창밖으로 뻗고 있으며, 시선도 창밖을 향하고 있음.",
        "built_space": "뒷좌석 패턴의 소파가 거대한 창문과 평행하게 배치되어 위치 레퍼런스의 실제 구조(정면을 향한 뒷좌석)를 완전히 왜곡함.",
        "entities": "라울의 외형은 일치하나, 옷과 피부가 마른 상태로 묘사되어 비에 젖은 지친 모습이라는 지시를 어김.",
        "hard_violations": [
         "위치 레퍼런스에 없는 거대한 창문을 생성하고 소파의 구조적 배치를 자의적으로 왜곡함 (물리적으로 불가능한 스테이징)"
        ],
        "physics": "소파 쿠션 위에 앉아 체중을 지탱하고 있으나, 뻗은 오른팔의 어깨 연결부가 어색하게 처리됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "후방 좌석에 앉아 오른쪽 위 창밖 숲을 가리키는 관계가 정확하지만, 무릎까지 포함한 구도가 요구한 상체 중심보다 넓고 흠뻑 젖고 지친 상태는 약하다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "숲을 향한 검지와 참조 공간의 세부는 잘 맞지만, 라울이 지정된 후방 좌석이 아니라 통로 앞쪽에 앉아 있어 핵심 배치를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "라울은 오른팔과 검지를 오른쪽으로 길게 뻗으며, 손끝의 연장선은 오른쪽 위 측면 창에 보이는 숲으로 이어진다. 눈도 같은 창밖 방향을 향한다.",
        "built_space": "후방 소파 하나, 오른쪽 측면 창 하나와 하단 잠금장치 하나, 뒤벽의 직사각형 설비 하나, 창 옆 작은 설비들이 보인다. 라울은 소파 오른쪽 부분에 앉아 몸을 앞으로 기울였고 좌판이 몸 아래에 이어진다. 통로 쪽에서 창을 비스듬히 보는 관계가 성립한다. 참조의 낡은 벽과 무늬 있는 소파는 유사하지만, 오른쪽 전경의 목재 수납장 마감은 참조의 밝은 설비와 다르다. 판정할 만한 거울 반사는 없다.",
        "entities": "사람은 라울에 해당하는 어린 남자아이 한 명뿐이다. 갈색 피부, 어린 얼굴, 가는 체격, 뒤로 묶은 곱슬머리와 회색 티셔츠·카고 반바지는 참조 인물과 대체로 부합하며, 지정된 혼혈 배경과 외관상 모순되지 않는다. 창밖에는 낮의 녹색 숲이 보인다. 피부에는 약간의 습기와 상처 흔적이 있지만 옷은 흠뻑 젖은 느낌이 약하고 치료 흔적은 뚜렷하지 않다. 짐·식량·약품은 이 구도에서 확인되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 소파 좌판에 놓이고 왼손은 옆 좌판을 짚어 앞으로 기울인 상체를 지탱한다. 오른팔은 어깨에서 자연스럽게 뻗어 있으며 검지로 가리키는 손의 형태도 가능하다. 지지 없이 뜬 신체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "라울의 뻗은 검지는 오른쪽 위 창 너머 나무들을 향한다. 얼굴과 눈도 오른쪽 창밖을 향해 있어 숲을 목적지로 삼는 동작은 읽힌다.",
        "built_space": "좌우 측면 창 두 개, 뒤쪽 소파 하나와 하부 서랍, 오른쪽 벽등 하나, 왼쪽 커튼 하나, 천장 개구부에 걸린 천과 왼쪽 창틀의 천이 보여 참조 공간의 구성과 재질을 잘 보존한다. 그러나 후방 소파는 비어 있고 라울은 그보다 훨씬 앞쪽 통로 가장자리에 앉아 있다. 몸 아래의 일부 받침은 후방 소파와 연결되지 않아 지정된 후방 좌석 점유 관계가 성립하지 않는다. 불가능한 반사는 보이지 않는다.",
        "entities": "사람은 어린 남자아이 한 명이며 얼굴, 갈색 피부, 뒤로 묶은 머리, 회색 티셔츠와 카고 반바지가 라울의 참조와 대체로 맞는다. 피부와 옷의 습기, 지친 표정은 A보다 뚜렷하다. 팔의 긁힌 상처는 보이지만 치료된 상태인지는 확인하기 어렵다. 오른쪽 창 너머에는 낮의 숲과 젖은 도로가 보이고 창에는 물방울이 남아 있다. 식량과 약품은 식별되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "라울이 점유해야 하는 후방 소파를 비워 둔 채 인물을 앞쪽 통로의 다른 좌석 위치에 배치했다."
        ],
        "physics": "왼손은 화면 왼쪽 아래의 좌석 가장자리로 보이는 받침을 짚고, 골반과 허벅지도 화면 아래로 잘리는 좌면 위에 놓인 자세로 읽힌다. 따라서 공중에 떠 있다고 볼 근거는 없다. 뻗은 팔과 검지의 동작도 물리적으로 가능하지만, 그 지지 위치가 요구된 후방 소파가 아니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "후방 좌석에 앉아 오른쪽 위 창밖 숲을 가리키는 관계가 정확하지만, 무릎까지 포함한 구도가 요구한 상체 중심보다 넓고 흠뻑 젖고 지친 상태는 약하다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "숲을 향한 검지와 참조 공간의 세부는 잘 맞지만, 라울이 지정된 후방 좌석이 아니라 통로 앞쪽에 앉아 있어 핵심 배치를 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "라울은 오른팔과 검지를 오른쪽으로 길게 뻗으며, 손끝의 연장선은 오른쪽 위 측면 창에 보이는 숲으로 이어진다. 눈도 같은 창밖 방향을 향한다.",
        "built_space": "후방 소파 하나, 오른쪽 측면 창 하나와 하단 잠금장치 하나, 뒤벽의 직사각형 설비 하나, 창 옆 작은 설비들이 보인다. 라울은 소파 오른쪽 부분에 앉아 몸을 앞으로 기울였고 좌판이 몸 아래에 이어진다. 통로 쪽에서 창을 비스듬히 보는 관계가 성립한다. 참조의 낡은 벽과 무늬 있는 소파는 유사하지만, 오른쪽 전경의 목재 수납장 마감은 참조의 밝은 설비와 다르다. 판정할 만한 거울 반사는 없다.",
        "entities": "사람은 라울에 해당하는 어린 남자아이 한 명뿐이다. 갈색 피부, 어린 얼굴, 가는 체격, 뒤로 묶은 곱슬머리와 회색 티셔츠·카고 반바지는 참조 인물과 대체로 부합하며, 지정된 혼혈 배경과 외관상 모순되지 않는다. 창밖에는 낮의 녹색 숲이 보인다. 피부에는 약간의 습기와 상처 흔적이 있지만 옷은 흠뻑 젖은 느낌이 약하고 치료 흔적은 뚜렷하지 않다. 짐·식량·약품은 이 구도에서 확인되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 소파 좌판에 놓이고 왼손은 옆 좌판을 짚어 앞으로 기울인 상체를 지탱한다. 오른팔은 어깨에서 자연스럽게 뻗어 있으며 검지로 가리키는 손의 형태도 가능하다. 지지 없이 뜬 신체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "라울의 뻗은 검지는 오른쪽 위 창 너머 나무들을 향한다. 얼굴과 눈도 오른쪽 창밖을 향해 있어 숲을 목적지로 삼는 동작은 읽힌다.",
        "built_space": "좌우 측면 창 두 개, 뒤쪽 소파 하나와 하부 서랍, 오른쪽 벽등 하나, 왼쪽 커튼 하나, 천장 개구부에 걸린 천과 왼쪽 창틀의 천이 보여 참조 공간의 구성과 재질을 잘 보존한다. 그러나 후방 소파는 비어 있고 라울은 그보다 훨씬 앞쪽 통로 가장자리에 앉아 있다. 몸 아래의 일부 받침은 후방 소파와 연결되지 않아 지정된 후방 좌석 점유 관계가 성립하지 않는다. 불가능한 반사는 보이지 않는다.",
        "entities": "사람은 어린 남자아이 한 명이며 얼굴, 갈색 피부, 뒤로 묶은 머리, 회색 티셔츠와 카고 반바지가 라울의 참조와 대체로 맞는다. 피부와 옷의 습기, 지친 표정은 A보다 뚜렷하다. 팔의 긁힌 상처는 보이지만 치료된 상태인지는 확인하기 어렵다. 오른쪽 창 너머에는 낮의 숲과 젖은 도로가 보이고 창에는 물방울이 남아 있다. 식량과 약품은 식별되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "라울이 점유해야 하는 후방 소파를 비워 둔 채 인물을 앞쪽 통로의 다른 좌석 위치에 배치했다."
        ],
        "physics": "왼손은 화면 왼쪽 아래의 좌석 가장자리로 보이는 받침을 짚고, 골반과 허벅지도 화면 아래로 잘리는 좌면 위에 놓인 자세로 읽힌다. 따라서 공중에 떠 있다고 볼 근거는 없다. 뻗은 팔과 검지의 동작도 물리적으로 가능하지만, 그 지지 위치가 요구된 후방 소파가 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 공간(뒷좌석)이 아닌 바닥에 인물을 배치함 (스테이징 지시 위반)",
     "[gpt-high] 라울이 점유해야 하는 후방 소파를 비워 둔 채 인물을 앞쪽 통로의 다른 좌석 위치에 배치했다."
    ],
    "B": [
     "[gemini-pro] 위치 레퍼런스에 없는 거대한 창문을 생성하고 소파의 구조적 배치를 자의적으로 왜곡함 (물리적으로 불가능한 스테이징)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "공간 레퍼런스의 구조와 비에 젖은 무드는 잘 살렸으나, 인물을 지정된 뒷좌석이 아닌 바닥에 배치하는 치명적인 스테이징 오류를 범함.  ★위반: [gemini-pro] 지정된 공간(뒷좌석)이 아닌 바닥에 인물을 배치함 (스테이징 지시 위반) / [gpt-high] 라울이 점유해야 하는 후방 소파를 비워 둔 채 인물을 앞쪽 통로의 다른 좌석 위치에 배치했다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "비에 젖은 상태라는 지시를 무시하고 캠퍼 내부 구조를 완전히 왜곡하여 위치 레퍼런스와 스테이징 조건을 모두 위반함.  ★위반: [gemini-pro] 위치 레퍼런스에 없는 거대한 창문을 생성하고 소파의 구조적 배치를 자의적으로 왜곡함 (물리적으로 불가능한 스테이징)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
    "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-0948-71e4-b198-7bcbfd68a85f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S55sh3__bgfirst_bg.png",
   "bg_asset_id": "f40aa134-4d1d-43ab-80e4-59ce7d1bc555",
   "bg_record_key": "S55sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S55sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:33:26.530316+00:00",
  "fingerprint": "58b58a096b672bf82fc4a43ac0b3bf0c7bad3338d38bf81b9e2f773502a299d2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S55sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S55sh3_sel.png",
  "source_sha256": "7e58532787c85c8cd470d20619ad57751989ddbab5433b16eece70c2f1f3b95d",
  "file": "S55sh3_cine.png",
  "staged_sha256": "174e009934e216f1fb373e6199b35be1eed31e663fe4a8cc926fce4ff7e403a2",
  "latency_ms": 10459
 },
 "S55sh5::signage": {
  "fp": "f6a6e7660f37d1bf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S55sh5": {
  "input_fingerprint": "48766cae24ec19f2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 솟구친 흙먼지를 뒤로한 채 숲길을 향해 급격히 꺾인 캠핑카의 바퀴 찰나.\n\nLOCATION (lock): At the turnoff from the provincial road onto a forest track, beside the camper's sharply turning wheels. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Forest route opening ahead of the turning wheel in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper front wheel (Sharply turned toward the forest route) — Seen from slightly behind the axle, exposing the changed relationship between its side and forward-facing tread; used as Primary motion detail, kept at natural scale against the adjoining vehicle body; Adjacent camper bodywork (Moving with the turning wheel) — A partial side section recedes above and behind the wheel; used as Provides a scale reference and establishes the wheel's steering angle relative to the vehicle; Forest route entrance (Being entered by the turning camper) — Opens ahead of the wheel toward the upper right; used as Makes the destination of the turn legible; Kicked-up earth (Thrown backward during the abrupt turn) — Trails behind the wheel toward the lower left; used as Reinforces the turn's force without obscuring wheel contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the wheel angle, road contact, and kicked-up earth distinct without stylized lighting changes.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper turns toward the forest after a passing shower, still carrying the loaded luggage, food and medicine.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 솟구친 흙먼지를 뒤로한 채 숲길을 향해 급격히 꺾인 캠핑카의 바퀴 찰나.\n\nLOCATION (lock): At the turnoff from the provincial road onto a forest track, beside the camper's sharply turning wheels. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Forest route opening ahead of the turning wheel in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper front wheel (Sharply turned toward the forest route) — Seen from slightly behind the axle, exposing the changed relationship between its side and forward-facing tread; used as Primary motion detail, kept at natural scale against the adjoining vehicle body; Adjacent camper bodywork (Moving with the turning wheel) — A partial side section recedes above and behind the wheel; used as Provides a scale reference and establishes the wheel's steering angle relative to the vehicle; Forest route entrance (Being entered by the turning camper) — Opens ahead of the wheel toward the upper right; used as Makes the destination of the turn legible; Kicked-up earth (Thrown backward during the abrupt turn) — Trails behind the wheel toward the lower left; used as Reinforces the turn's force without obscuring wheel contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the wheel angle, road contact, and kicked-up earth distinct without stylized lighting changes.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper turns toward the forest after a passing shower, still carrying the loaded luggage, food and medicine.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 솟구친 흙먼지를 뒤로한 채 숲길을 향해 급격히 꺾인 캠핑카의 바퀴 찰나.\n\nLOCATION (lock): At the turnoff from the provincial road onto a forest track, beside the camper's sharply turning wheels. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Forest route opening ahead of the turning wheel in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper front wheel (Sharply turned toward the forest route) — Seen from slightly behind the axle, exposing the changed relationship between its side and forward-facing tread; used as Primary motion detail, kept at natural scale against the adjoining vehicle body; Adjacent camper bodywork (Moving with the turning wheel) — A partial side section recedes above and behind the wheel; used as Provides a scale reference and establishes the wheel's steering angle relative to the vehicle; Forest route entrance (Being entered by the turning camper) — Opens ahead of the wheel toward the upper right; used as Makes the destination of the turn legible; Kicked-up earth (Thrown backward during the abrupt turn) — Trails behind the wheel toward the lower left; used as Reinforces the turn's force without obscuring wheel contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the wheel angle, road contact, and kicked-up earth distinct without stylized lighting changes.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper turns toward the forest after a passing shower, still carrying the loaded luggage, food and medicine.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "차체 앞부분은 화면 오른쪽에 있고 측면은 왼쪽으로 이어진다. 앞바퀴는 측면과 넓은 트레드가 함께 보이도록 꺾여 있으며 진행 방향은 우상단 흙길 입구 쪽으로 읽힌다. 흙먼지는 접지점에서 좌하단 포장도로 쪽으로 뻗는다. 사람이나 시선 대상은 없다.",
    "built_space": "앞바퀴 한 개와 휠하우스 한 개가 크게 보이고, 왼쪽 차체 아래에는 다른 바퀴 일부가 보인다. 아래쪽 포장도로와 오른쪽 흙길이 맞닿아 있어 진입 관계는 성립한다. 다만 참조의 깊게 갈라진 포장면과 바위·낙엽이 쌓인 오른쪽 둔덕 대신 비교적 평탄한 길 가장자리와 흰 도로선이 보인다. 건물이나 반사면은 없다.",
    "entities": "은색 차량의 앞바퀴, 인접 차체, 숲길, 자갈과 흙먼지가 보이며 인물이나 읽을 수 있는 글자는 없다. 차체는 참조의 짙은 하부 패널과 수납함이 있는 캠핑카보다 매끈한 승합차 외장에 가깝다. 짐·식량·의약품은 이 외부 클로즈업에서 보이지 않아 적재 상태를 판단할 수 없다. 낮 장면이지만 소나기 직후의 젖은 표면은 뚜렷하지 않다.",
    "hard_violations": [],
    "physics": "앞타이어 하단이 흙과 포장 경계에 닿아 있고 바퀴는 휠하우스와 차체 아래 장치에 연결되어 있다. 차체와 바퀴의 크기 관계도 자연스럽다. 공중의 흙과 작은 돌은 접지점 부근에서 좌하단으로 튀어나온 것으로 보여 급회전 중 지면을 긁는 작용으로 설명된다. 근거 없이 떠 있는 물체는 없다."
   },
   {
    "label": "A",
    "direction": "차체 측면은 앞바퀴에서 왼쪽 뒤로 물러나고 앞범퍼 일부는 오른쪽에 보인다. 바퀴의 측면과 트레드가 동시에 드러나 차체에 대한 조향각이 읽히며, 바퀴 앞쪽에는 우상단으로 이어지는 숲길이 열린다. 흙먼지와 자갈은 바퀴 뒤쪽인 좌하단으로 튄다. 사람이나 시선 대상은 없다.",
    "built_space": "앞바퀴와 휠하우스가 각각 하나씩 보이고, 차체 아래 왼쪽에는 다른 바퀴 하나가 부분적으로 보인다. 측면에는 큰 하부 수납 패널과 그 위쪽 패널이 이어진다. 전경의 갈라진 포장도로, 우상단의 좁은 흙길, 오른쪽 나무줄기와 돌·낙엽 둔덕이 참조 장소의 구성과 가깝다. 낮은 차축 뒤쪽 시점에서 앞바퀴와 진입로가 함께 보이는 공간 관계도 성립한다.",
    "entities": "참조와 유사한 회색 캠핑카 외장과 어두운 하부 패널, 조향된 앞바퀴, 숲길 입구, 튀는 흙과 자갈이 보인다. 인물과 읽을 수 있는 글자는 없다. 내부 적재물은 프레임 밖이므로 판단 대상이 아니다. 숲의 낮 조명은 맞지만 소나기 직후라는 단서는 표면에서 강하게 드러나지 않는다.",
    "hard_violations": [],
    "physics": "앞바퀴는 자갈 섞인 지면에 확실히 접촉하고 휠하우스 안쪽으로 연결되며, 다른 바퀴도 차체 아래 지면에 놓여 있다. 차량을 지탱하는 관계와 바퀴 크기는 자연스럽다. 떠 있는 자갈과 흙은 접지점 뒤에 집중되어 타이어가 흙을 밀어내며 튀긴 궤적으로 설명되고, 먼지가 접지부를 완전히 가리지 않는다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "바퀴 클로즈업과 우상단 진입·좌하단 흙먼지 배치는 충실하지만, 매끈한 승합차형 차체와 단순한 길 가장자리가 참조의 캠핑카 및 장소와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "차축 뒤쪽에서 본 꺾인 앞바퀴 클로즈업을 유지하면서 참조의 캠핑카 외장, 갈라진 도로와 돌 많은 숲길, 뒤로 튀는 흙을 가장 충실하게 재현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "차체 앞부분은 화면 오른쪽에 있고 측면은 왼쪽으로 이어진다. 앞바퀴는 측면과 넓은 트레드가 함께 보이도록 꺾여 있으며 진행 방향은 우상단 흙길 입구 쪽으로 읽힌다. 흙먼지는 접지점에서 좌하단 포장도로 쪽으로 뻗는다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴 한 개와 휠하우스 한 개가 크게 보이고, 왼쪽 차체 아래에는 다른 바퀴 일부가 보인다. 아래쪽 포장도로와 오른쪽 흙길이 맞닿아 있어 진입 관계는 성립한다. 다만 참조의 깊게 갈라진 포장면과 바위·낙엽이 쌓인 오른쪽 둔덕 대신 비교적 평탄한 길 가장자리와 흰 도로선이 보인다. 건물이나 반사면은 없다.",
        "entities": "은색 차량의 앞바퀴, 인접 차체, 숲길, 자갈과 흙먼지가 보이며 인물이나 읽을 수 있는 글자는 없다. 차체는 참조의 짙은 하부 패널과 수납함이 있는 캠핑카보다 매끈한 승합차 외장에 가깝다. 짐·식량·의약품은 이 외부 클로즈업에서 보이지 않아 적재 상태를 판단할 수 없다. 낮 장면이지만 소나기 직후의 젖은 표면은 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "앞타이어 하단이 흙과 포장 경계에 닿아 있고 바퀴는 휠하우스와 차체 아래 장치에 연결되어 있다. 차체와 바퀴의 크기 관계도 자연스럽다. 공중의 흙과 작은 돌은 접지점 부근에서 좌하단으로 튀어나온 것으로 보여 급회전 중 지면을 긁는 작용으로 설명된다. 근거 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "차체 측면은 앞바퀴에서 왼쪽 뒤로 물러나고 앞범퍼 일부는 오른쪽에 보인다. 바퀴의 측면과 트레드가 동시에 드러나 차체에 대한 조향각이 읽히며, 바퀴 앞쪽에는 우상단으로 이어지는 숲길이 열린다. 흙먼지와 자갈은 바퀴 뒤쪽인 좌하단으로 튄다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴와 휠하우스가 각각 하나씩 보이고, 차체 아래 왼쪽에는 다른 바퀴 하나가 부분적으로 보인다. 측면에는 큰 하부 수납 패널과 그 위쪽 패널이 이어진다. 전경의 갈라진 포장도로, 우상단의 좁은 흙길, 오른쪽 나무줄기와 돌·낙엽 둔덕이 참조 장소의 구성과 가깝다. 낮은 차축 뒤쪽 시점에서 앞바퀴와 진입로가 함께 보이는 공간 관계도 성립한다.",
        "entities": "참조와 유사한 회색 캠핑카 외장과 어두운 하부 패널, 조향된 앞바퀴, 숲길 입구, 튀는 흙과 자갈이 보인다. 인물과 읽을 수 있는 글자는 없다. 내부 적재물은 프레임 밖이므로 판단 대상이 아니다. 숲의 낮 조명은 맞지만 소나기 직후라는 단서는 표면에서 강하게 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞바퀴는 자갈 섞인 지면에 확실히 접촉하고 휠하우스 안쪽으로 연결되며, 다른 바퀴도 차체 아래 지면에 놓여 있다. 차량을 지탱하는 관계와 바퀴 크기는 자연스럽다. 떠 있는 자갈과 흙은 접지점 뒤에 집중되어 타이어가 흙을 밀어내며 튀긴 궤적으로 설명되고, 먼지가 접지부를 완전히 가리지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "바퀴 클로즈업과 우상단 진입·좌하단 흙먼지 배치는 충실하지만, 매끈한 승합차형 차체와 단순한 길 가장자리가 참조의 캠핑카 및 장소와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "차축 뒤쪽에서 본 꺾인 앞바퀴 클로즈업을 유지하면서 참조의 캠핑카 외장, 갈라진 도로와 돌 많은 숲길, 뒤로 튀는 흙을 가장 충실하게 재현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "차체 앞부분은 화면 오른쪽에 있고 측면은 왼쪽으로 이어진다. 앞바퀴는 측면과 넓은 트레드가 함께 보이도록 꺾여 있으며 진행 방향은 우상단 흙길 입구 쪽으로 읽힌다. 흙먼지는 접지점에서 좌하단 포장도로 쪽으로 뻗는다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴 한 개와 휠하우스 한 개가 크게 보이고, 왼쪽 차체 아래에는 다른 바퀴 일부가 보인다. 아래쪽 포장도로와 오른쪽 흙길이 맞닿아 있어 진입 관계는 성립한다. 다만 참조의 깊게 갈라진 포장면과 바위·낙엽이 쌓인 오른쪽 둔덕 대신 비교적 평탄한 길 가장자리와 흰 도로선이 보인다. 건물이나 반사면은 없다.",
        "entities": "은색 차량의 앞바퀴, 인접 차체, 숲길, 자갈과 흙먼지가 보이며 인물이나 읽을 수 있는 글자는 없다. 차체는 참조의 짙은 하부 패널과 수납함이 있는 캠핑카보다 매끈한 승합차 외장에 가깝다. 짐·식량·의약품은 이 외부 클로즈업에서 보이지 않아 적재 상태를 판단할 수 없다. 낮 장면이지만 소나기 직후의 젖은 표면은 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "앞타이어 하단이 흙과 포장 경계에 닿아 있고 바퀴는 휠하우스와 차체 아래 장치에 연결되어 있다. 차체와 바퀴의 크기 관계도 자연스럽다. 공중의 흙과 작은 돌은 접지점 부근에서 좌하단으로 튀어나온 것으로 보여 급회전 중 지면을 긁는 작용으로 설명된다. 근거 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "차체 측면은 앞바퀴에서 왼쪽 뒤로 물러나고 앞범퍼 일부는 오른쪽에 보인다. 바퀴의 측면과 트레드가 동시에 드러나 차체에 대한 조향각이 읽히며, 바퀴 앞쪽에는 우상단으로 이어지는 숲길이 열린다. 흙먼지와 자갈은 바퀴 뒤쪽인 좌하단으로 튄다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴와 휠하우스가 각각 하나씩 보이고, 차체 아래 왼쪽에는 다른 바퀴 하나가 부분적으로 보인다. 측면에는 큰 하부 수납 패널과 그 위쪽 패널이 이어진다. 전경의 갈라진 포장도로, 우상단의 좁은 흙길, 오른쪽 나무줄기와 돌·낙엽 둔덕이 참조 장소의 구성과 가깝다. 낮은 차축 뒤쪽 시점에서 앞바퀴와 진입로가 함께 보이는 공간 관계도 성립한다.",
        "entities": "참조와 유사한 회색 캠핑카 외장과 어두운 하부 패널, 조향된 앞바퀴, 숲길 입구, 튀는 흙과 자갈이 보인다. 인물과 읽을 수 있는 글자는 없다. 내부 적재물은 프레임 밖이므로 판단 대상이 아니다. 숲의 낮 조명은 맞지만 소나기 직후라는 단서는 표면에서 강하게 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞바퀴는 자갈 섞인 지면에 확실히 접촉하고 휠하우스 안쪽으로 연결되며, 다른 바퀴도 차체 아래 지면에 놓여 있다. 차량을 지탱하는 관계와 바퀴 크기는 자연스럽다. 떠 있는 자갈과 흙은 접지점 뒤에 집중되어 타이어가 흙을 밀어내며 튀긴 궤적으로 설명되고, 먼지가 접지부를 완전히 가리지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 7,
   "A": 9
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 7,
    "verdict_ko": "바퀴 클로즈업과 우상단 진입·좌하단 흙먼지 배치는 충실하지만, 매끈한 승합차형 차체와 단순한 길 가장자리가 참조의 캠핑카 및 장소와 다르다."
   },
   {
    "label": "A",
    "score": 9,
    "verdict_ko": "차축 뒤쪽에서 본 꺾인 앞바퀴 클로즈업을 유지하면서 참조의 캠핑카 외장, 갈라진 도로와 돌 많은 숲길, 뒤로 튀는 흙을 가장 충실하게 재현했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B01.png",
    "asset_id": "ffcb0b14-6163-46a6-a2d3-38e545216e88",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-0c80-727d-8da6-68f93443cf2f",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S55sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:18:39.099938+00:00",
  "fingerprint": "5b35a120db932620a374842422a1c958524683aa823751bcd8240327f438cfa3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S55sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S55sh5_sel.png",
  "source_sha256": "09159aa142b7a96071f17257237e4b8dcc3c937f5c0b2eb4801f15ee3086ae7a",
  "file": "S55sh5_cine.png",
  "staged_sha256": "9fe50594028551319fc3b509145cee82a368c94089c6f32ef63c4fd79a69e4ed",
  "latency_ms": 12362
 },
 "S56sh5::signage": {
  "fp": "822597267d87434b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::8860149eacdd8c67": {
  "subjects": [],
  "subject_text": "익산 늪지대의 캠핑카 고립 지점\n숲길 끝에 펼쳐진 황폐한 늪지대. 물기 많은 진흙 바닥 옆에 출입 금지 팻말과 심하게 부식된 방사능 구역 표지가 서 있다.",
  "identity": "canonical",
  "scope_id": "L215",
  "scope_role": "location_exterior",
  "scope_sha": "a6042d9cf76b5b12"
 },
 "groupbg::swamp_vehicle_edge": {
  "input_fingerprint": "7404e75b60cc272c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "swamp_vehicle_edge",
    "tags": [
     "S56sh5",
     "S61sh1",
     "S61sh5",
     "S61sh7"
    ]
   },
   "context_sig": "b16126b42fff63ca"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 현우의 차가 늪지대 앞에서 멈추려다 풍덩 빠진다.\n- 부근에 출입 금지 팻말과 함께 부식돼 식별이 안 되는 방사능 구역 표시\n- 늪지대. 박철진과 민병대원들이 만지고 둘러보고 있는 건... 현우 일행이 탔던 자동차다. 녹슨 방사능 구역 팻말.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 현우의 차가 늪지대 앞에서 멈추려다 풍덩 빠진다.\n- 부근에 출입 금지 팻말과 함께 부식돼 식별이 안 되는 방사능 구역 표시\n- 늪지대. 박철진과 민병대원들이 만지고 둘러보고 있는 건... 현우 일행이 탔던 자동차다. 녹슨 방사능 구역 팻말.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_swamp_vehicle_edge_78578e.png",
  "asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e",
  "input_asset_ids": [
   "e0803289-5e5b-4591-a71c-27038d182ed9"
  ],
  "origin_tag": "S56sh5",
  "place_text": "On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.",
  "origin_inputs": {
   "place_text": "On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.",
   "time_of_day_en": "day",
   "conti_asset_id": "e0803289-5e5b-4591-a71c-27038d182ed9"
  }
 },
 "S56sh5::bgfirst_bg": {
  "input_fingerprint": "78af173715b24daf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5__bgfirst_bg.png",
  "asset_id": "ec65b360-e758-4e20-aaa5-dfe753a8386e",
  "input_asset_ids": [
   "e0803289-5e5b-4591-a71c-27038d182ed9",
   "bf15ced8-25a7-46c4-b3de-7db0c008c53e"
  ]
 },
 "S56sh5": {
  "input_fingerprint": "8f08740f32503a92",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is stuck in the swamp, and large shoeprints mark the nearby mud beside a no-entry sign and a badly corroded radiation warning sign. Charlie is already in the forest wearing the oversized straw hat, rubber boots and colorful raincoat. 앰버: She is searching for wood near the large shoeprints and remains damp and tired from the preceding journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is stuck in the swamp, and large shoeprints mark the nearby mud beside a no-entry sign and a badly corroded radiation warning sign. Charlie is already in the forest wearing the oversized straw hat, rubber boots and colorful raincoat. 앰버: She is searching for wood near the large shoeprints and remains damp and tired from the preceding journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is stuck in the swamp, and large shoeprints mark the nearby mud beside a no-entry sign and a badly corroded radiation warning sign. Charlie is already in the forest wearing the oversized straw hat, rubber boots and colorful raincoat. 앰버: She is searching for wood near the large shoeprints and remains damp and tired from the preceding journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5__bgfirst_bg.png",
     "asset_id": "ec65b360-e758-4e20-aaa5-dfe753a8386e",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S56sh5.png",
     "asset_id": "e0803289-5e5b-4591-a71c-27038d182ed9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_swamp_vehicle_edge_78578e.png",
     "asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 고개를 숙여 자기 앞 오른쪽의 큰 진흙 자국 가까이를 바라본다. 몸도 그 자국을 향해 쪼그려 있으며 카메라를 응시하지 않는다. 양손은 자국 주변의 나무 조각에 닿아 있다.",
    "built_space": "배경 오른쪽에 진흙에 빠진 캠핑카 한 대, 그 오른쪽에 원형 진입금지 표지 하나와 사각 방사능 표지 하나가 있다. 캠핑카의 측면과 전면, 갈대와 물웅덩이, 왼쪽 전경의 나무 조각들이 장소 참조와 잘 대응한다. 앰버와 오른쪽 아래 자국은 연속된 진흙 지면 위에 놓인다. 전신과 주변을 포함하는 와이드 숏이지만 자국의 화면상 크기가 쪼그린 인물과 비슷해, 더 작게 두라는 지시가 약하게 구현됐다.",
    "entities": "금발의 어린 여자아이 한 명만 보이며 대략 열 살의 체격과 둥근 얼굴은 참조에 대체로 가깝다. 한국계 백인 혼혈 여부는 외모만으로 확정할 수 없다. 갈색 작업복과 끈 달린 부츠는 참조에 가깝지만 반팔 상의는 남색이 아닌 회갈색이고, 머리 위 장비와 공구 벨트는 보이지 않는다. 진흙투성이 모습과 지친 표정은 이어지는 상태에 부합한다. 큰 자국은 가로 홈이 반복되는 직사각형에 가까워 신발 밑창 윤곽이 불명확하며, 뒤쪽에도 유사한 흔적이 이어진다. 낮의 늪지, 캠핑카와 두 표지는 확인되고 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "양쪽 부츠가 진흙에 닿고 무릎과 엉덩이가 굽혀져 쪼그린 체중을 지탱한다. 손이 닿은 나무 조각은 지면에 놓여 있어 떠 있지 않다. 캠핑카는 진흙에 잠긴 바퀴로 지지되고 표지 기둥도 땅에 박혀 있다. 자국 안에 고인 물과 젖은 진흙 가장자리는 물리적인 표면으로 읽힌다."
   },
   {
    "label": "A",
    "direction": "앰버의 얼굴과 숙인 시선이 오른쪽 아래의 큰 신발 자국을 향한다. 몸도 오른쪽으로 돌아 자국을 관찰하는 자세이며, 팔은 굽힌 무릎 위에 자연스럽게 놓여 있다.",
    "built_space": "배경 오른쪽의 캠핑카 한 대와 그 오른쪽의 원형 진입금지 표지 하나, 사각 방사능 표지 하나가 참조의 상대 배치를 유지한다. 갈대, 얕은 물웅덩이, 진흙과 왼쪽의 목재도 같은 장소를 이룬다. 왼쪽의 쪼그린 인물 전신과 늪지 배경을 담는 와이드 구도이며, 오른쪽 아래 신발 자국 전체가 비스듬한 위쪽 시점으로 보인다. 자국은 인물보다 작고 둘 사이에는 끊김 없는 진흙 지면이 이어진다.",
    "entities": "금발의 어린 여자아이 한 명만 보이고 나이와 체격은 앰버 설정에 대체로 맞는다. 옆얼굴이라 참조의 정면 얼굴과 큰 눈을 정밀하게 비교하기는 어렵고, 혼혈 배경도 외모만으로 확정할 수 없다. 회색 후드 외투와 녹색 고무장화는 참조의 남색 반팔·갈색 작업복·끈 달린 부츠와 다르며 머리 위 장비도 없다. 옷의 오염과 숙인 자세는 고단한 여정에 어울린다. 오른쪽 아래에는 발바닥 외곽과 밑창 무늬가 모두 식별되는 거대한 신발 자국 하나가 있다. 캠핑카와 두 표지, 낮의 늪지가 확인되며 추가 인물이나 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "앞쪽 고무장화의 밑창이 진흙에 넓게 닿고 뒤쪽 발도 몸 아래에서 접지해 쪼그린 몸을 지탱한다. 팔은 무릎에 기대고 손은 자연스럽게 내려와 있다. 나무는 지면에 놓여 있으며 캠핑카와 표지들도 지면의 지지를 받는다. 신발 자국은 진흙에 눌린 테두리와 밑창 홈으로 표현되어 붙여 넣은 평면 무늬나 떠 있는 물체로 보이지 않는다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "늪지와 쪼그린 동작은 충실하지만, 오른쪽 아래 자국이 인물보다 작게 읽히지 않고 신발보다는 길쭉한 궤도 흔적에 가까워 핵심 구성이 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "의상은 참조와 다르지만, 앰버가 바라보는 완전한 신발 자국을 오른쪽 아래에 인물보다 작게 배치해 우선순위가 높은 장면 구성을 더 정확히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 고개를 숙여 자기 앞 오른쪽의 큰 진흙 자국 가까이를 바라본다. 몸도 그 자국을 향해 쪼그려 있으며 카메라를 응시하지 않는다. 양손은 자국 주변의 나무 조각에 닿아 있다.",
        "built_space": "배경 오른쪽에 진흙에 빠진 캠핑카 한 대, 그 오른쪽에 원형 진입금지 표지 하나와 사각 방사능 표지 하나가 있다. 캠핑카의 측면과 전면, 갈대와 물웅덩이, 왼쪽 전경의 나무 조각들이 장소 참조와 잘 대응한다. 앰버와 오른쪽 아래 자국은 연속된 진흙 지면 위에 놓인다. 전신과 주변을 포함하는 와이드 숏이지만 자국의 화면상 크기가 쪼그린 인물과 비슷해, 더 작게 두라는 지시가 약하게 구현됐다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며 대략 열 살의 체격과 둥근 얼굴은 참조에 대체로 가깝다. 한국계 백인 혼혈 여부는 외모만으로 확정할 수 없다. 갈색 작업복과 끈 달린 부츠는 참조에 가깝지만 반팔 상의는 남색이 아닌 회갈색이고, 머리 위 장비와 공구 벨트는 보이지 않는다. 진흙투성이 모습과 지친 표정은 이어지는 상태에 부합한다. 큰 자국은 가로 홈이 반복되는 직사각형에 가까워 신발 밑창 윤곽이 불명확하며, 뒤쪽에도 유사한 흔적이 이어진다. 낮의 늪지, 캠핑카와 두 표지는 확인되고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "양쪽 부츠가 진흙에 닿고 무릎과 엉덩이가 굽혀져 쪼그린 체중을 지탱한다. 손이 닿은 나무 조각은 지면에 놓여 있어 떠 있지 않다. 캠핑카는 진흙에 잠긴 바퀴로 지지되고 표지 기둥도 땅에 박혀 있다. 자국 안에 고인 물과 젖은 진흙 가장자리는 물리적인 표면으로 읽힌다."
       },
       {
        "label": "B",
        "direction": "앰버의 얼굴과 숙인 시선이 오른쪽 아래의 큰 신발 자국을 향한다. 몸도 오른쪽으로 돌아 자국을 관찰하는 자세이며, 팔은 굽힌 무릎 위에 자연스럽게 놓여 있다.",
        "built_space": "배경 오른쪽의 캠핑카 한 대와 그 오른쪽의 원형 진입금지 표지 하나, 사각 방사능 표지 하나가 참조의 상대 배치를 유지한다. 갈대, 얕은 물웅덩이, 진흙과 왼쪽의 목재도 같은 장소를 이룬다. 왼쪽의 쪼그린 인물 전신과 늪지 배경을 담는 와이드 구도이며, 오른쪽 아래 신발 자국 전체가 비스듬한 위쪽 시점으로 보인다. 자국은 인물보다 작고 둘 사이에는 끊김 없는 진흙 지면이 이어진다.",
        "entities": "금발의 어린 여자아이 한 명만 보이고 나이와 체격은 앰버 설정에 대체로 맞는다. 옆얼굴이라 참조의 정면 얼굴과 큰 눈을 정밀하게 비교하기는 어렵고, 혼혈 배경도 외모만으로 확정할 수 없다. 회색 후드 외투와 녹색 고무장화는 참조의 남색 반팔·갈색 작업복·끈 달린 부츠와 다르며 머리 위 장비도 없다. 옷의 오염과 숙인 자세는 고단한 여정에 어울린다. 오른쪽 아래에는 발바닥 외곽과 밑창 무늬가 모두 식별되는 거대한 신발 자국 하나가 있다. 캠핑카와 두 표지, 낮의 늪지가 확인되며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 고무장화의 밑창이 진흙에 넓게 닿고 뒤쪽 발도 몸 아래에서 접지해 쪼그린 몸을 지탱한다. 팔은 무릎에 기대고 손은 자연스럽게 내려와 있다. 나무는 지면에 놓여 있으며 캠핑카와 표지들도 지면의 지지를 받는다. 신발 자국은 진흙에 눌린 테두리와 밑창 홈으로 표현되어 붙여 넣은 평면 무늬나 떠 있는 물체로 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "늪지와 쪼그린 동작은 충실하지만, 오른쪽 아래 자국이 인물보다 작게 읽히지 않고 신발보다는 길쭉한 궤도 흔적에 가까워 핵심 구성이 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "의상은 참조와 다르지만, 앰버가 바라보는 완전한 신발 자국을 오른쪽 아래에 인물보다 작게 배치해 우선순위가 높은 장면 구성을 더 정확히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 고개를 숙여 자기 앞 오른쪽의 큰 진흙 자국 가까이를 바라본다. 몸도 그 자국을 향해 쪼그려 있으며 카메라를 응시하지 않는다. 양손은 자국 주변의 나무 조각에 닿아 있다.",
        "built_space": "배경 오른쪽에 진흙에 빠진 캠핑카 한 대, 그 오른쪽에 원형 진입금지 표지 하나와 사각 방사능 표지 하나가 있다. 캠핑카의 측면과 전면, 갈대와 물웅덩이, 왼쪽 전경의 나무 조각들이 장소 참조와 잘 대응한다. 앰버와 오른쪽 아래 자국은 연속된 진흙 지면 위에 놓인다. 전신과 주변을 포함하는 와이드 숏이지만 자국의 화면상 크기가 쪼그린 인물과 비슷해, 더 작게 두라는 지시가 약하게 구현됐다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며 대략 열 살의 체격과 둥근 얼굴은 참조에 대체로 가깝다. 한국계 백인 혼혈 여부는 외모만으로 확정할 수 없다. 갈색 작업복과 끈 달린 부츠는 참조에 가깝지만 반팔 상의는 남색이 아닌 회갈색이고, 머리 위 장비와 공구 벨트는 보이지 않는다. 진흙투성이 모습과 지친 표정은 이어지는 상태에 부합한다. 큰 자국은 가로 홈이 반복되는 직사각형에 가까워 신발 밑창 윤곽이 불명확하며, 뒤쪽에도 유사한 흔적이 이어진다. 낮의 늪지, 캠핑카와 두 표지는 확인되고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "양쪽 부츠가 진흙에 닿고 무릎과 엉덩이가 굽혀져 쪼그린 체중을 지탱한다. 손이 닿은 나무 조각은 지면에 놓여 있어 떠 있지 않다. 캠핑카는 진흙에 잠긴 바퀴로 지지되고 표지 기둥도 땅에 박혀 있다. 자국 안에 고인 물과 젖은 진흙 가장자리는 물리적인 표면으로 읽힌다."
       },
       {
        "label": "A",
        "direction": "앰버의 얼굴과 숙인 시선이 오른쪽 아래의 큰 신발 자국을 향한다. 몸도 오른쪽으로 돌아 자국을 관찰하는 자세이며, 팔은 굽힌 무릎 위에 자연스럽게 놓여 있다.",
        "built_space": "배경 오른쪽의 캠핑카 한 대와 그 오른쪽의 원형 진입금지 표지 하나, 사각 방사능 표지 하나가 참조의 상대 배치를 유지한다. 갈대, 얕은 물웅덩이, 진흙과 왼쪽의 목재도 같은 장소를 이룬다. 왼쪽의 쪼그린 인물 전신과 늪지 배경을 담는 와이드 구도이며, 오른쪽 아래 신발 자국 전체가 비스듬한 위쪽 시점으로 보인다. 자국은 인물보다 작고 둘 사이에는 끊김 없는 진흙 지면이 이어진다.",
        "entities": "금발의 어린 여자아이 한 명만 보이고 나이와 체격은 앰버 설정에 대체로 맞는다. 옆얼굴이라 참조의 정면 얼굴과 큰 눈을 정밀하게 비교하기는 어렵고, 혼혈 배경도 외모만으로 확정할 수 없다. 회색 후드 외투와 녹색 고무장화는 참조의 남색 반팔·갈색 작업복·끈 달린 부츠와 다르며 머리 위 장비도 없다. 옷의 오염과 숙인 자세는 고단한 여정에 어울린다. 오른쪽 아래에는 발바닥 외곽과 밑창 무늬가 모두 식별되는 거대한 신발 자국 하나가 있다. 캠핑카와 두 표지, 낮의 늪지가 확인되며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 고무장화의 밑창이 진흙에 넓게 닿고 뒤쪽 발도 몸 아래에서 접지해 쪼그린 몸을 지탱한다. 팔은 무릎에 기대고 손은 자연스럽게 내려와 있다. 나무는 지면에 놓여 있으며 캠핑카와 표지들도 지면의 지지를 받는다. 신발 자국은 진흙에 눌린 테두리와 밑창 홈으로 표현되어 붙여 넣은 평면 무늬나 떠 있는 물체로 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 7,
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 7,
    "verdict_ko": "늪지와 쪼그린 동작은 충실하지만, 오른쪽 아래 자국이 인물보다 작게 읽히지 않고 신발보다는 길쭉한 궤도 흔적에 가까워 핵심 구성이 약하다."
   },
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "의상은 참조와 다르지만, 앰버가 바라보는 완전한 신발 자국을 오른쪽 아래에 인물보다 작게 배치해 우선순위가 높은 장면 구성을 더 정확히 구현한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_swamp_vehicle_edge_78578e.png",
    "asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-0e29-764e-b558-04aeb7e08c5b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5__bgfirst_bg.png",
   "bg_asset_id": "ec65b360-e758-4e20-aaa5-dfe753a8386e",
   "bg_record_key": "S56sh5::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "swamp_vehicle_edge",
   "groupbg_asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S56sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:34:34.627727+00:00",
  "fingerprint": "57fd35dc33150f079648cfe9b964a366de53b7177e4605bc82dcfb71ef1a63fb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S56sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S56sh5_sel.png",
  "source_sha256": "b64d413a711c7fc1a5192dda2cdfca8b758df15ed9fc90d09f523603a5020735",
  "file": "S56sh5_cine.png",
  "staged_sha256": "0053ea4cb7d3f2a6abbf11fd08f40497a40a1c89b411b6e1af0ee09fe3dff8ae",
  "latency_ms": 10585
 },
 "S56sh8::signage": {
  "fp": "8cb7151b94050d87",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S56sh8": {
  "input_fingerprint": "e21852974c9180dc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie retains the oversized straw hat, rubber boots and colorful raincoat as a capture net descends over him in the forest. The camper remains stuck in the swamp, with the large muddy shoeprints and corroded warning signs nearby.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie retains the oversized straw hat, rubber boots and colorful raincoat as a capture net descends over him in the forest. The camper remains stuck in the swamp, with the large muddy shoeprints and corroded warning signs nearby.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie retains the oversized straw hat, rubber boots and colorful raincoat as a capture net descends over him in the forest. The camper remains stuck in the swamp, with the large muddy shoeprints and corroded warning signs nearby.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8__bgfirst_bg.png",
     "asset_id": "7e75322c-4c67-4f9e-95d0-69b421b60a39",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S56sh8.png",
     "asset_id": "b891d386-6990-4b27-8c39-13f03b4127da",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_forest_trap_site_08ebbc.png",
     "asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 왼팔이 위를 향해 허공으로 뻗어 있으며, 시선은 떨어지는 그물과 정면 사이를 향함.",
    "built_space": "늪지 옆 진흙길 위에 피사체가 위치하며, 배경의 고장난 캠핑카와 우측의 방사능 경고판 등 레퍼런스의 공간 요소가 정확히 배치됨.",
    "entities": "찰리의 금속 몸체, 밀짚모자, 화려한 우비, 장화가 모두 일치하게 등장하며, 거대한 그물이 몸통 앞쪽을 덮고 있음.",
    "hard_violations": [],
    "physics": "두 발로 진흙 바닥을 안정적으로 딛고 무게 중심이 이동하는 찰나를 보여주며, 그물은 중력에 의해 몸 위로 쏟아지는 중임."
   },
   {
    "label": "B",
    "direction": "찰리의 왼팔이 비스듬히 허공을 향해 들려 있고 시선은 정면을 주시함.",
    "built_space": "진흙길, 배경의 캠핑카와 방사능 표지판 등 지정된 공간 요소들이 올바른 위치에 있음.",
    "entities": "찰리의 상단 장비(모자, 우비)와 금속 몸체가 프레임 안에 일치하게 표현되었으나, 그물망이 몸통이 아닌 등 뒤에 위치함.",
    "hard_violations": [
     "[gpt-high] 우비 주머니 부근에 읽을 수 있는 영문 로고가 노출되어 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "바닥을 딛고 서 있으나 움직임이 다소 경직된 정자세이며, 그물망은 몸과 접촉하지 않은 채 뒤쪽 허공에 팽팽하게 고정되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "그물이 찰리의 몸 앞쪽을 덮치는 찰나의 핵심 연출을 정확하게 구현했으며 포즈와 장비가 자연스럽게 묘사되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "몸 앞을 덮쳐야 할 그물망이 피사체 등 뒤에 배경처럼 배치되어 텍스트의 핵심 연출 조건을 완전히 충족하지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 왼팔이 위를 향해 허공으로 뻗어 있으며, 시선은 떨어지는 그물과 정면 사이를 향함.",
        "built_space": "늪지 옆 진흙길 위에 피사체가 위치하며, 배경의 고장난 캠핑카와 우측의 방사능 경고판 등 레퍼런스의 공간 요소가 정확히 배치됨.",
        "entities": "찰리의 금속 몸체, 밀짚모자, 화려한 우비, 장화가 모두 일치하게 등장하며, 거대한 그물이 몸통 앞쪽을 덮고 있음.",
        "hard_violations": [],
        "physics": "두 발로 진흙 바닥을 안정적으로 딛고 무게 중심이 이동하는 찰나를 보여주며, 그물은 중력에 의해 몸 위로 쏟아지는 중임."
       },
       {
        "label": "B",
        "direction": "찰리의 왼팔이 비스듬히 허공을 향해 들려 있고 시선은 정면을 주시함.",
        "built_space": "진흙길, 배경의 캠핑카와 방사능 표지판 등 지정된 공간 요소들이 올바른 위치에 있음.",
        "entities": "찰리의 상단 장비(모자, 우비)와 금속 몸체가 프레임 안에 일치하게 표현되었으나, 그물망이 몸통이 아닌 등 뒤에 위치함.",
        "hard_violations": [],
        "physics": "바닥을 딛고 서 있으나 움직임이 다소 경직된 정자세이며, 그물망은 몸과 접촉하지 않은 채 뒤쪽 허공에 팽팽하게 고정되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "그물이 찰리의 몸 앞쪽을 덮치는 찰나의 핵심 연출을 정확하게 구현했으며 포즈와 장비가 자연스럽게 묘사되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "몸 앞을 덮쳐야 할 그물망이 피사체 등 뒤에 배경처럼 배치되어 텍스트의 핵심 연출 조건을 완전히 충족하지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 왼팔이 위를 향해 허공으로 뻗어 있으며, 시선은 떨어지는 그물과 정면 사이를 향함.",
        "built_space": "늪지 옆 진흙길 위에 피사체가 위치하며, 배경의 고장난 캠핑카와 우측의 방사능 경고판 등 레퍼런스의 공간 요소가 정확히 배치됨.",
        "entities": "찰리의 금속 몸체, 밀짚모자, 화려한 우비, 장화가 모두 일치하게 등장하며, 거대한 그물이 몸통 앞쪽을 덮고 있음.",
        "hard_violations": [],
        "physics": "두 발로 진흙 바닥을 안정적으로 딛고 무게 중심이 이동하는 찰나를 보여주며, 그물은 중력에 의해 몸 위로 쏟아지는 중임."
       },
       {
        "label": "B",
        "direction": "찰리의 왼팔이 비스듬히 허공을 향해 들려 있고 시선은 정면을 주시함.",
        "built_space": "진흙길, 배경의 캠핑카와 방사능 표지판 등 지정된 공간 요소들이 올바른 위치에 있음.",
        "entities": "찰리의 상단 장비(모자, 우비)와 금속 몸체가 프레임 안에 일치하게 표현되었으나, 그물망이 몸통이 아닌 등 뒤에 위치함.",
        "hard_violations": [],
        "physics": "바닥을 딛고 서 있으나 움직임이 다소 경직된 정자세이며, 그물망은 몸과 접촉하지 않은 채 뒤쪽 허공에 팽팽하게 고정되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "중심 인물의 크기와 금속 몸체는 적절하지만, 그물이 몸통 앞을 덮치지 않고 뒤에 펼쳐져 있으며 우비에 영문 로고가 노출된다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "그물이 위에서 내려와 몸통 가까운 면을 접어 덮치는 핵심 동작은 정확하지만, 장화까지 담은 전신 구도는 지정된 미디엄 숏보다 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 정면에서 약간 화면 오른쪽을 향하고, 화면 오른쪽 팔은 위로 뻗어 손바닥을 펼친다. 그물은 화면 상단에서 어깨 뒤쪽으로 내려오지만 가슴과 배 앞에는 걸리지 않는다. 뻗은 손 위 공간에는 그물이 있으나, 몸통의 카메라 쪽 면을 덮치는 방향 관계는 구현되지 않았다.",
        "built_space": "왼쪽의 굵은 나무와 흰 꽃, 중앙의 진흙길과 물웅덩이, 오른쪽 뒤 늪에 기울어진 캠퍼 한 대가 보인다. 오른쪽에는 원형 표지 하나와 그 아래 사각 표지 하나가 있어 장소 참조의 주요 배치와 일치한다. 찰리는 길 전경에 서 있고 캠퍼는 작은 원경 요소로 유지된다. 하체 아래를 자른 구도는 두 후보 중 미디엄 숏에 더 가깝다.",
        "entities": "찰리 한 명만 보이며, 흰 분절형 마스크 얼굴과 붉은 기계 눈, 샌드 베이지 장갑판, 육중한 팔, 큰 밀짚모자와 다색 우비는 참조와 부합한다. 명시된 비인간 로봇이므로 사람의 연령·성별·민족성은 적용되지 않는다. 장화는 프레임 밖이다. 거대한 밧줄 그물과 숲, 꽃, 늪의 캠퍼 및 부식된 표지가 보인다. 우비의 화면 오른쪽 주머니 부근에는 흰 영문 로고가 드러나 있다.",
        "hard_violations": [
         "우비 주머니 부근에 읽을 수 있는 영문 로고가 노출되어 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "몸통과 팔은 연결된 기계 관절로 지지되고, 다리는 화면 아래로 이어진다. 발이 잘렸다는 이유로 몸이 떠 있다고 볼 근거는 없다. 모자는 머리에 놓여 있다. 그물은 상단 밖에서 내려오는 중으로 해석할 수 있어 무지지 정지 물체는 아니지만, 가슴 앞 접촉과 몸통을 따라 접히는 변형이 없어 요구된 충돌 순간은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 얼굴과 펼친 손을 화면 오른쪽 위로 향한다. 그물은 화면 왼쪽 상단에서 내려와 어깨를 지나 가슴과 배의 카메라 쪽 면을 가로지르고, 아래쪽 가장자리는 다리 옆으로 떨어진다. 그물이 실제로 겨냥해 덮는 대상이 찰리의 몸통으로 읽힌다. 다만 뻗은 손 바로 위의 하늘은 상당 부분 열려 있다.",
        "built_space": "왼쪽 굵은 나무와 꽃밭, 중앙 진흙길, 뒤쪽 늪과 캠퍼 한 대, 오른쪽의 원형·사각 표지 각 하나가 참조 장소에 맞게 배치된다. 찰리는 늪 앞 숲길에 있고 배경 캠퍼가 과도하게 커지지 않았다. 그러나 양쪽 장화와 주변 지면까지 모두 보이는 전신 구도여서 지정된 미디엄 숏보다 넓다.",
        "entities": "등장 인물은 찰리뿐이다. 흰 기계 마스크, 붉은 눈, 베이지 금속 몸통, 밀짚모자, 화려한 우비와 녹색 고무장화가 확인된다. 우비가 어깨와 팔을 많이 가려 참조의 거대한 장갑판과 긴 팔 비례는 덜 두드러진다. 그물, 꽃, 나무, 진흙과 물웅덩이, 늪에 박힌 캠퍼, 부식된 표지가 있다. 큰 신발 자국은 진흙 요철과 명확히 구별되지 않는다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 앞 장화와 오른쪽 뒤 장화가 지면에 닿고, 굽힌 무릎과 벌어진 발이 기울어진 몸통을 받친다. 들어 올린 팔은 어깨와 팔꿈치 관절에 연결되어 있다. 그물은 어깨와 가슴에 닿아 꺾이고 남은 부분은 중력 방향으로 늘어져 일부가 지면에 닿는다. 낙하 중 몸통에 걸리는 상황으로 가능한 접촉과 지지가 보이며, 근거 없이 공중에 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "중심 인물의 크기와 금속 몸체는 적절하지만, 그물이 몸통 앞을 덮치지 않고 뒤에 펼쳐져 있으며 우비에 영문 로고가 노출된다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "그물이 위에서 내려와 몸통 가까운 면을 접어 덮치는 핵심 동작은 정확하지만, 장화까지 담은 전신 구도는 지정된 미디엄 숏보다 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 정면에서 약간 화면 오른쪽을 향하고, 화면 오른쪽 팔은 위로 뻗어 손바닥을 펼친다. 그물은 화면 상단에서 어깨 뒤쪽으로 내려오지만 가슴과 배 앞에는 걸리지 않는다. 뻗은 손 위 공간에는 그물이 있으나, 몸통의 카메라 쪽 면을 덮치는 방향 관계는 구현되지 않았다.",
        "built_space": "왼쪽의 굵은 나무와 흰 꽃, 중앙의 진흙길과 물웅덩이, 오른쪽 뒤 늪에 기울어진 캠퍼 한 대가 보인다. 오른쪽에는 원형 표지 하나와 그 아래 사각 표지 하나가 있어 장소 참조의 주요 배치와 일치한다. 찰리는 길 전경에 서 있고 캠퍼는 작은 원경 요소로 유지된다. 하체 아래를 자른 구도는 두 후보 중 미디엄 숏에 더 가깝다.",
        "entities": "찰리 한 명만 보이며, 흰 분절형 마스크 얼굴과 붉은 기계 눈, 샌드 베이지 장갑판, 육중한 팔, 큰 밀짚모자와 다색 우비는 참조와 부합한다. 명시된 비인간 로봇이므로 사람의 연령·성별·민족성은 적용되지 않는다. 장화는 프레임 밖이다. 거대한 밧줄 그물과 숲, 꽃, 늪의 캠퍼 및 부식된 표지가 보인다. 우비의 화면 오른쪽 주머니 부근에는 흰 영문 로고가 드러나 있다.",
        "hard_violations": [
         "우비 주머니 부근에 읽을 수 있는 영문 로고가 노출되어 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "몸통과 팔은 연결된 기계 관절로 지지되고, 다리는 화면 아래로 이어진다. 발이 잘렸다는 이유로 몸이 떠 있다고 볼 근거는 없다. 모자는 머리에 놓여 있다. 그물은 상단 밖에서 내려오는 중으로 해석할 수 있어 무지지 정지 물체는 아니지만, 가슴 앞 접촉과 몸통을 따라 접히는 변형이 없어 요구된 충돌 순간은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 얼굴과 펼친 손을 화면 오른쪽 위로 향한다. 그물은 화면 왼쪽 상단에서 내려와 어깨를 지나 가슴과 배의 카메라 쪽 면을 가로지르고, 아래쪽 가장자리는 다리 옆으로 떨어진다. 그물이 실제로 겨냥해 덮는 대상이 찰리의 몸통으로 읽힌다. 다만 뻗은 손 바로 위의 하늘은 상당 부분 열려 있다.",
        "built_space": "왼쪽 굵은 나무와 꽃밭, 중앙 진흙길, 뒤쪽 늪과 캠퍼 한 대, 오른쪽의 원형·사각 표지 각 하나가 참조 장소에 맞게 배치된다. 찰리는 늪 앞 숲길에 있고 배경 캠퍼가 과도하게 커지지 않았다. 그러나 양쪽 장화와 주변 지면까지 모두 보이는 전신 구도여서 지정된 미디엄 숏보다 넓다.",
        "entities": "등장 인물은 찰리뿐이다. 흰 기계 마스크, 붉은 눈, 베이지 금속 몸통, 밀짚모자, 화려한 우비와 녹색 고무장화가 확인된다. 우비가 어깨와 팔을 많이 가려 참조의 거대한 장갑판과 긴 팔 비례는 덜 두드러진다. 그물, 꽃, 나무, 진흙과 물웅덩이, 늪에 박힌 캠퍼, 부식된 표지가 있다. 큰 신발 자국은 진흙 요철과 명확히 구별되지 않는다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 앞 장화와 오른쪽 뒤 장화가 지면에 닿고, 굽힌 무릎과 벌어진 발이 기울어진 몸통을 받친다. 들어 올린 팔은 어깨와 팔꿈치 관절에 연결되어 있다. 그물은 어깨와 가슴에 닿아 꺾이고 남은 부분은 중력 방향으로 늘어져 일부가 지면에 닿는다. 낙하 중 몸통에 걸리는 상황으로 가능한 접촉과 지지가 보이며, 근거 없이 공중에 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gpt-high] 우비 주머니 부근에 읽을 수 있는 영문 로고가 노출되어 글자와 로고를 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "그물이 찰리의 몸 앞쪽을 덮치는 찰나의 핵심 연출을 정확하게 구현했으며 포즈와 장비가 자연스럽게 묘사되었습니다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "몸 앞을 덮쳐야 할 그물망이 피사체 등 뒤에 배경처럼 배치되어 텍스트의 핵심 연출 조건을 완전히 충족하지 못했습니다.  ★위반: [gpt-high] 우비 주머니 부근에 읽을 수 있는 영문 로고가 노출되어 글자와 로고를 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_forest_trap_site_08ebbc.png",
    "asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-12ef-70bf-a7d4-98d8d06c3947",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8__bgfirst_bg.png",
   "bg_asset_id": "7e75322c-4c67-4f9e-95d0-69b421b60a39",
   "bg_record_key": "S56sh8::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "forest_trap_site",
   "groupbg_asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S56sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:35:42.508583+00:00",
  "fingerprint": "6327c8d02f473a86728245307f2e8d47bce377768463c44c5019c82bad6a93c4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S56sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S56sh8_sel.png",
  "source_sha256": "3f5b3eea851f4a8cbbf95409639ffab0ebf45e2cdb313fb8204026e3b3eaf58f",
  "file": "S56sh8_cine.png",
  "staged_sha256": "61def834a5b1f605dc146281524cf2f94758c64b6949efd82576f7fcd55f05e5",
  "latency_ms": 11408
 },
 "S56sh14::signage": {
  "fp": "b86d6b02e078402c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S56sh14": {
  "input_fingerprint": "402e9f6d1a3b963d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 구덩이 바닥에 엎어진 현우와 앰버를 둥글게 둘러싼 채 내려다보는 하회탈 병사들의 전경.\n\nLOCATION (lock): At an open trap pit in the forest near the marsh, with armed masked figures gathered around its rim. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Pit rim, walls, and floor (Open trap containing 현우 and 앰버) — The near rim, inner walls, and bottom are visible in one oblique overhead view; used as Establishes the vertical separation between captives and captors; Hahoe-shaped leather masks (Worn by the surrounding men in varied colors) — Near-side masks turn away and downward; far-side masks expose downward-angled fronts; used as Repeating but nonidentical details around the enclosing perimeter; Spears, axes, and bows (Held by the men around the pit) — Different oblique angles follow individual grips rather than a uniform radial pattern; used as Breaks the perimeter into distinct threatening silhouettes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled tonal separation keeps the pit floor and the encircling figures readable without inventing a separate light source below.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An arrow remains embedded in a tree beside the approach to the pit trap, and the camper remains stranded in the swamp. Charlie has been captured in the forest while wearing the oversized hat, boots and colorful raincoat. 현우: He is at the bottom of the pit trap with his earlier treated injuries retained. The contact card remains concealed in his shoe. 앰버: She is at the bottom of the pit trap.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 구덩이 바닥에 엎어진 현우와 앰버를 둥글게 둘러싼 채 내려다보는 하회탈 병사들의 전경.\n\nLOCATION (lock): At an open trap pit in the forest near the marsh, with armed masked figures gathered around its rim. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Pit rim, walls, and floor (Open trap containing 현우 and 앰버) — The near rim, inner walls, and bottom are visible in one oblique overhead view; used as Establishes the vertical separation between captives and captors; Hahoe-shaped leather masks (Worn by the surrounding men in varied colors) — Near-side masks turn away and downward; far-side masks expose downward-angled fronts; used as Repeating but nonidentical details around the enclosing perimeter; Spears, axes, and bows (Held by the men around the pit) — Different oblique angles follow individual grips rather than a uniform radial pattern; used as Breaks the perimeter into distinct threatening silhouettes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled tonal separation keeps the pit floor and the encircling figures readable without inventing a separate light source below.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An arrow remains embedded in a tree beside the approach to the pit trap, and the camper remains stranded in the swamp. Charlie has been captured in the forest while wearing the oversized hat, boots and colorful raincoat. 현우: He is at the bottom of the pit trap with his earlier treated injuries retained. The contact card remains concealed in his shoe. 앰버: She is at the bottom of the pit trap.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 구덩이 바닥에 엎어진 현우와 앰버를 둥글게 둘러싼 채 내려다보는 하회탈 병사들의 전경.\n\nLOCATION (lock): At an open trap pit in the forest near the marsh, with armed masked figures gathered around its rim. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Pit rim, walls, and floor (Open trap containing 현우 and 앰버) — The near rim, inner walls, and bottom are visible in one oblique overhead view; used as Establishes the vertical separation between captives and captors; Hahoe-shaped leather masks (Worn by the surrounding men in varied colors) — Near-side masks turn away and downward; far-side masks expose downward-angled fronts; used as Repeating but nonidentical details around the enclosing perimeter; Spears, axes, and bows (Held by the men around the pit) — Different oblique angles follow individual grips rather than a uniform radial pattern; used as Breaks the perimeter into distinct threatening silhouettes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled tonal separation keeps the pit floor and the encircling figures readable without inventing a separate light source below.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An arrow remains embedded in a tree beside the approach to the pit trap, and the camper remains stranded in the swamp. Charlie has been captured in the forest while wearing the oversized hat, boots and colorful raincoat. 현우: He is at the bottom of the pit trap with his earlier treated injuries retained. The contact card remains concealed in his shoe. 앰버: She is at the bottom of the pit trap.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "병사들의 시선과 무기(창, 도끼, 활)의 끝이 모두 구덩이 바닥에 쓰러져 있는 현우와 앰버를 향하고 있습니다.",
    "built_space": "늪지대 근처 숲에 파인 넓은 원형 흙 구덩이이며, 이전 샷 레퍼런스의 버려진 캠핑카가 배경에 올바르게 배치되어 있습니다.",
    "entities": "현우와 앰버는 의상과 외모 레퍼런스가 일치하며 바닥에 엎어져 있습니다. 병사들은 다양한 색상의 하회탈을 쓰고 지시된 무기들을 들고 있습니다.",
    "hard_violations": [],
    "physics": "모든 인물은 지면과 구덩이 바닥에 안정적으로 위치해 있으며, 병사들의 손은 각자의 무기를 자연스럽게 쥐고 있습니다."
   },
   {
    "label": "B",
    "direction": "병사들의 시선과 무기들이 구덩이 중앙을 향해 아래로 향하고 있습니다.",
    "built_space": "상단에 사각형의 나무판자가 덧대어진 구덩이 형태이며, 배경은 늪지대와 숲입니다.",
    "entities": "구덩이 안의 현우와 앰버는 레퍼런스와 일치하나, 가장자리에 선 병사들 중 좌측 상단 인물이 현우와 완전히 동일한 얼굴과 의상을 입고 있습니다. 일부 병사가 프롬프트에 없는 현대식 소총을 들고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 현우 캐릭터의 복제 (구덩이 안에 1명, 가장자리에 1명)",
     "[gemini-pro] 프롬프트에서 요구하지 않은 발명된 물건(현대식 총기 및 전술 조끼) 등장",
     "[gpt-high] 구덩이 바닥의 현우와 별개로, 가장자리에 현우와 같은 머리·회색 셔츠·카키색 바지의 비탈 청년이 추가되어 인물이 중복된다.",
     "[gpt-high] 지정된 창·도끼·활 외에 여러 병사에게 총기를 추가했다."
    ],
    "physics": "인물들은 지면과 바닥에 닿아 있으며 무기를 손에 쥐고 있으나, 일부 무기의 파지가 다소 어색합니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구덩이 안의 현우와 앰버, 그리고 이들을 둘러싼 하회탈 병사들의 배치와 무기가 프롬프트의 지시와 레퍼런스를 정확히 따르고 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "가장자리에 서 있는 병사 중 한 명이 현우와 똑같은 모습으로 복제되어 나타나는 치명적인 오류가 있으며, 지시되지 않은 현대식 총기가 등장합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "병사들의 시선과 무기(창, 도끼, 활)의 끝이 모두 구덩이 바닥에 쓰러져 있는 현우와 앰버를 향하고 있습니다.",
        "built_space": "늪지대 근처 숲에 파인 넓은 원형 흙 구덩이이며, 이전 샷 레퍼런스의 버려진 캠핑카가 배경에 올바르게 배치되어 있습니다.",
        "entities": "현우와 앰버는 의상과 외모 레퍼런스가 일치하며 바닥에 엎어져 있습니다. 병사들은 다양한 색상의 하회탈을 쓰고 지시된 무기들을 들고 있습니다.",
        "hard_violations": [],
        "physics": "모든 인물은 지면과 구덩이 바닥에 안정적으로 위치해 있으며, 병사들의 손은 각자의 무기를 자연스럽게 쥐고 있습니다."
       },
       {
        "label": "B",
        "direction": "병사들의 시선과 무기들이 구덩이 중앙을 향해 아래로 향하고 있습니다.",
        "built_space": "상단에 사각형의 나무판자가 덧대어진 구덩이 형태이며, 배경은 늪지대와 숲입니다.",
        "entities": "구덩이 안의 현우와 앰버는 레퍼런스와 일치하나, 가장자리에 선 병사들 중 좌측 상단 인물이 현우와 완전히 동일한 얼굴과 의상을 입고 있습니다. 일부 병사가 프롬프트에 없는 현대식 소총을 들고 있습니다.",
        "hard_violations": [
         "현우 캐릭터의 복제 (구덩이 안에 1명, 가장자리에 1명)",
         "프롬프트에서 요구하지 않은 발명된 물건(현대식 총기 및 전술 조끼) 등장"
        ],
        "physics": "인물들은 지면과 바닥에 닿아 있으며 무기를 손에 쥐고 있으나, 일부 무기의 파지가 다소 어색합니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구덩이 안의 현우와 앰버, 그리고 이들을 둘러싼 하회탈 병사들의 배치와 무기가 프롬프트의 지시와 레퍼런스를 정확히 따르고 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "가장자리에 서 있는 병사 중 한 명이 현우와 똑같은 모습으로 복제되어 나타나는 치명적인 오류가 있으며, 지시되지 않은 현대식 총기가 등장합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "병사들의 시선과 무기(창, 도끼, 활)의 끝이 모두 구덩이 바닥에 쓰러져 있는 현우와 앰버를 향하고 있습니다.",
        "built_space": "늪지대 근처 숲에 파인 넓은 원형 흙 구덩이이며, 이전 샷 레퍼런스의 버려진 캠핑카가 배경에 올바르게 배치되어 있습니다.",
        "entities": "현우와 앰버는 의상과 외모 레퍼런스가 일치하며 바닥에 엎어져 있습니다. 병사들은 다양한 색상의 하회탈을 쓰고 지시된 무기들을 들고 있습니다.",
        "hard_violations": [],
        "physics": "모든 인물은 지면과 구덩이 바닥에 안정적으로 위치해 있으며, 병사들의 손은 각자의 무기를 자연스럽게 쥐고 있습니다."
       },
       {
        "label": "B",
        "direction": "병사들의 시선과 무기들이 구덩이 중앙을 향해 아래로 향하고 있습니다.",
        "built_space": "상단에 사각형의 나무판자가 덧대어진 구덩이 형태이며, 배경은 늪지대와 숲입니다.",
        "entities": "구덩이 안의 현우와 앰버는 레퍼런스와 일치하나, 가장자리에 선 병사들 중 좌측 상단 인물이 현우와 완전히 동일한 얼굴과 의상을 입고 있습니다. 일부 병사가 프롬프트에 없는 현대식 소총을 들고 있습니다.",
        "hard_violations": [
         "현우 캐릭터의 복제 (구덩이 안에 1명, 가장자리에 1명)",
         "프롬프트에서 요구하지 않은 발명된 물건(현대식 총기 및 전술 조끼) 등장"
        ],
        "physics": "인물들은 지면과 바닥에 닿아 있으며 무기를 손에 쥐고 있으나, 일부 무기의 파지가 다소 어색합니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "비스듬한 부감 전경과 포위 구도는 맞지만, 바닥의 현우 외에 현우와 같은 모습의 비탈 인물을 가장자리에 추가했고 지정되지 않은 총기까지 등장시켰다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구덩이 안팎의 높이 차, 하회탈 병사들의 포위, 습지 배경과 무기 구성이 더 충실하지만, 앰버가 엎어진 대신 등을 대고 누웠고 일부 병사는 아래를 충분히 내려다보지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "가까운 병사들은 뒤통수나 탈의 옆면을 보이며 구덩이 쪽으로 고개를 돌린다. 맞은편 병사들도 대체로 안쪽을 향하지만 일부 탈의 정면은 바닥보다 카메라 높이를 향한다. 창과 도끼는 각기 다른 사선으로 구덩이 입구를 가로지르며, 모두 두 포로의 몸을 직접 겨냥하지는 않는다. 오른쪽 아래 총구는 앰버의 상체 가까운 방향을 향한다. 현우는 옆으로 얼굴을 돌렸고 앰버는 바닥을 향한다.",
        "built_space": "중앙에 사각 구덩이 하나가 있고, 입구 테두리와 내부의 판재·수직 지지목, 흙바닥이 보인다. 병사들은 가장자리 지면에 서고 두 포로는 그보다 낮은 바닥에 놓여 있어 수직 분리는 읽힌다. 가까운 테두리는 전경 병사들 때문에 일부 가려진다. 배경의 숲과 습지는 이어지지만, 참고에 확인되지 않는 규칙적인 목재 벽체가 장소의 주요 구조로 추가되어 있다. 반사면은 없다.",
        "entities": "바닥에는 헝클어진 검은 머리, 회색 셔츠, 카키색 바지의 젊은 남성과 금발, 남색 상의, 갈색 멜빵바지의 여자아이가 있어 현우와 앰버의 큰 외형·의상 특징에 부합한다. 작은 얼굴 크기로 정확한 얼굴 일치는 판별하기 어렵다. 현우의 셔츠에는 상처 흔적처럼 보이는 얼룩이 있으나 치료 상태는 명확하지 않다. 가장자리에는 약 15명이 있으며, 그중 회색 셔츠와 카키색 바지를 입은 검은 머리 청년 한 명은 탈을 쓰지 않아 현우의 중복 인물처럼 보인다. 나머지는 여러 색의 탈과 전술 조끼를 착용한다. 창·도끼·활 외에 여러 총기가 보인다. 앰버의 참고 머리 장비는 확인되지 않으며, 나무에 박힌 화살과 캠핑카도 이 화면에서는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "구덩이 바닥의 현우와 별개로, 가장자리에 현우와 같은 머리·회색 셔츠·카키색 바지의 비탈 청년이 추가되어 인물이 중복된다.",
         "지정된 창·도끼·활 외에 여러 병사에게 총기를 추가했다."
        ],
        "physics": "병사들의 보이는 발은 구덩이 밖 지면에 닿으며, 무기는 손으로 잡거나 몸에 걸쳐 지지한다. 현우는 옆구리와 굽힌 다리를 흙바닥에 대고 있고, 앰버는 상체와 팔다리를 바닥에 댄 채 엎어져 있다. 현우의 자세는 완전히 엎드린 상태보다는 옆으로 쓰러진 상태다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "가까운 병사들은 카메라에 등과 탈의 뒷끈을 보이며 구덩이 안을 향한다. 맞은편에서는 탈의 정면이 보이지만, 여러 병사의 고개가 거의 수평이라 두 포로를 내려다보는 시선이 일관되지는 않는다. 창끝과 도끼날은 구덩이 안쪽으로 서로 다른 각도로 뻗고, 일부는 포로보다 높은 벽이나 입구 쪽을 향한다. 활은 손에 든 상태이며 일제히 조준하거나 발사하는 모습은 아니다. 현우의 얼굴은 왼쪽 바닥을 향하고, 앰버의 얼굴은 위를 향한다.",
        "built_space": "흙으로 파인 구덩이 하나가 중앙에 있으며 뿌리가 드러난 벽, 바닥, 앞쪽 테두리 일부를 한 부감 화면에서 볼 수 있다. 병사들은 입구의 바깥 지면에 원형으로 배치되고, 현우와 앰버는 깊이 내려간 바닥에 있다. 전경 병사들이 가까운 테두리를 일부 가리지만 높이 차는 명료하다. 뒤쪽의 숲 가장자리, 습지, 고사목, 언덕과 기울어진 캠핑카 한 대가 이전 장소의 특징을 이어간다. 반사에 의한 공간 모순은 없다.",
        "entities": "현우는 검은 머리의 젊은 남성으로 회색 셔츠와 카키색 바지를 입고, 앰버는 금발 여자아이로 남색 상의와 갈색 멜빵바지를 입어 연령대·체격·의상 구분이 대체로 맞는다. 얼굴이 작아 정확한 동일 인물 여부와 혼혈 특징까지 확정하기는 어렵다. 앰버의 참고 머리 장비와 현우의 치료 흔적은 명확히 식별되지 않는다. 둘을 둘러싼 성인 남성들은 여러 색의 하회탈 형태 가면을 쓰고 창·도끼·활과 화살통을 지닌다. 가면 때문에 병사들의 구체적인 얼굴과 민족적 외형은 확인할 수 없다. 별도의 비탈 청년이나 찰리는 없으며, 배경 캠핑카는 습지에 남아 있다. 나무에 박힌 화살과 신발 안 카드는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "가장자리 병사들은 발을 지면에 딛고 있고, 창과 도끼는 손으로 자루를 잡으며 활도 손으로 지지한다. 화살통은 등에 걸려 있다. 현우는 몸통과 굽힌 다리가 바닥에 닿은 옆으로 엎어진 자세다. 앰버는 등과 다리가 바닥에 지지되어 물리적으로는 가능하지만, 요청한 엎어진 자세가 아니라 반듯하게 누운 자세다. 두 사람 모두 공중에 뜨지 않았으며 배경 캠핑카는 습지 지면에 기울어져 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "비스듬한 부감 전경과 포위 구도는 맞지만, 바닥의 현우 외에 현우와 같은 모습의 비탈 인물을 가장자리에 추가했고 지정되지 않은 총기까지 등장시켰다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구덩이 안팎의 높이 차, 하회탈 병사들의 포위, 습지 배경과 무기 구성이 더 충실하지만, 앰버가 엎어진 대신 등을 대고 누웠고 일부 병사는 아래를 충분히 내려다보지 않는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "가까운 병사들은 뒤통수나 탈의 옆면을 보이며 구덩이 쪽으로 고개를 돌린다. 맞은편 병사들도 대체로 안쪽을 향하지만 일부 탈의 정면은 바닥보다 카메라 높이를 향한다. 창과 도끼는 각기 다른 사선으로 구덩이 입구를 가로지르며, 모두 두 포로의 몸을 직접 겨냥하지는 않는다. 오른쪽 아래 총구는 앰버의 상체 가까운 방향을 향한다. 현우는 옆으로 얼굴을 돌렸고 앰버는 바닥을 향한다.",
        "built_space": "중앙에 사각 구덩이 하나가 있고, 입구 테두리와 내부의 판재·수직 지지목, 흙바닥이 보인다. 병사들은 가장자리 지면에 서고 두 포로는 그보다 낮은 바닥에 놓여 있어 수직 분리는 읽힌다. 가까운 테두리는 전경 병사들 때문에 일부 가려진다. 배경의 숲과 습지는 이어지지만, 참고에 확인되지 않는 규칙적인 목재 벽체가 장소의 주요 구조로 추가되어 있다. 반사면은 없다.",
        "entities": "바닥에는 헝클어진 검은 머리, 회색 셔츠, 카키색 바지의 젊은 남성과 금발, 남색 상의, 갈색 멜빵바지의 여자아이가 있어 현우와 앰버의 큰 외형·의상 특징에 부합한다. 작은 얼굴 크기로 정확한 얼굴 일치는 판별하기 어렵다. 현우의 셔츠에는 상처 흔적처럼 보이는 얼룩이 있으나 치료 상태는 명확하지 않다. 가장자리에는 약 15명이 있으며, 그중 회색 셔츠와 카키색 바지를 입은 검은 머리 청년 한 명은 탈을 쓰지 않아 현우의 중복 인물처럼 보인다. 나머지는 여러 색의 탈과 전술 조끼를 착용한다. 창·도끼·활 외에 여러 총기가 보인다. 앰버의 참고 머리 장비는 확인되지 않으며, 나무에 박힌 화살과 캠핑카도 이 화면에서는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "구덩이 바닥의 현우와 별개로, 가장자리에 현우와 같은 머리·회색 셔츠·카키색 바지의 비탈 청년이 추가되어 인물이 중복된다.",
         "지정된 창·도끼·활 외에 여러 병사에게 총기를 추가했다."
        ],
        "physics": "병사들의 보이는 발은 구덩이 밖 지면에 닿으며, 무기는 손으로 잡거나 몸에 걸쳐 지지한다. 현우는 옆구리와 굽힌 다리를 흙바닥에 대고 있고, 앰버는 상체와 팔다리를 바닥에 댄 채 엎어져 있다. 현우의 자세는 완전히 엎드린 상태보다는 옆으로 쓰러진 상태다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "가까운 병사들은 카메라에 등과 탈의 뒷끈을 보이며 구덩이 안을 향한다. 맞은편에서는 탈의 정면이 보이지만, 여러 병사의 고개가 거의 수평이라 두 포로를 내려다보는 시선이 일관되지는 않는다. 창끝과 도끼날은 구덩이 안쪽으로 서로 다른 각도로 뻗고, 일부는 포로보다 높은 벽이나 입구 쪽을 향한다. 활은 손에 든 상태이며 일제히 조준하거나 발사하는 모습은 아니다. 현우의 얼굴은 왼쪽 바닥을 향하고, 앰버의 얼굴은 위를 향한다.",
        "built_space": "흙으로 파인 구덩이 하나가 중앙에 있으며 뿌리가 드러난 벽, 바닥, 앞쪽 테두리 일부를 한 부감 화면에서 볼 수 있다. 병사들은 입구의 바깥 지면에 원형으로 배치되고, 현우와 앰버는 깊이 내려간 바닥에 있다. 전경 병사들이 가까운 테두리를 일부 가리지만 높이 차는 명료하다. 뒤쪽의 숲 가장자리, 습지, 고사목, 언덕과 기울어진 캠핑카 한 대가 이전 장소의 특징을 이어간다. 반사에 의한 공간 모순은 없다.",
        "entities": "현우는 검은 머리의 젊은 남성으로 회색 셔츠와 카키색 바지를 입고, 앰버는 금발 여자아이로 남색 상의와 갈색 멜빵바지를 입어 연령대·체격·의상 구분이 대체로 맞는다. 얼굴이 작아 정확한 동일 인물 여부와 혼혈 특징까지 확정하기는 어렵다. 앰버의 참고 머리 장비와 현우의 치료 흔적은 명확히 식별되지 않는다. 둘을 둘러싼 성인 남성들은 여러 색의 하회탈 형태 가면을 쓰고 창·도끼·활과 화살통을 지닌다. 가면 때문에 병사들의 구체적인 얼굴과 민족적 외형은 확인할 수 없다. 별도의 비탈 청년이나 찰리는 없으며, 배경 캠핑카는 습지에 남아 있다. 나무에 박힌 화살과 신발 안 카드는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "가장자리 병사들은 발을 지면에 딛고 있고, 창과 도끼는 손으로 자루를 잡으며 활도 손으로 지지한다. 화살통은 등에 걸려 있다. 현우는 몸통과 굽힌 다리가 바닥에 닿은 옆으로 엎어진 자세다. 앰버는 등과 다리가 바닥에 지지되어 물리적으로는 가능하지만, 요청한 엎어진 자세가 아니라 반듯하게 누운 자세다. 두 사람 모두 공중에 뜨지 않았으며 배경 캠핑카는 습지 지면에 기울어져 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.714
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.464
   },
   "violations": {
    "B": [
     "[gemini-pro] 현우 캐릭터의 복제 (구덩이 안에 1명, 가장자리에 1명)",
     "[gemini-pro] 프롬프트에서 요구하지 않은 발명된 물건(현대식 총기 및 전술 조끼) 등장",
     "[gpt-high] 구덩이 바닥의 현우와 별개로, 가장자리에 현우와 같은 머리·회색 셔츠·카키색 바지의 비탈 청년이 추가되어 인물이 중복된다.",
     "[gpt-high] 지정된 창·도끼·활 외에 여러 병사에게 총기를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 464
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "구덩이 안의 현우와 앰버, 그리고 이들을 둘러싼 하회탈 병사들의 배치와 무기가 프롬프트의 지시와 레퍼런스를 정확히 따르고 있습니다."
   },
   {
    "label": "B",
    "score": 464,
    "verdict_ko": "가장자리에 서 있는 병사 중 한 명이 현우와 똑같은 모습으로 복제되어 나타나는 치명적인 오류가 있으며, 지시되지 않은 현대식 총기가 등장합니다.  ★위반: [gemini-pro] 현우 캐릭터의 복제 (구덩이 안에 1명, 가장자리에 1명) / [gemini-pro] 프롬프트에서 요구하지 않은 발명된 물건(현대식 총기 및 전술 조끼) 등장 / [gpt-high] 구덩이 바닥의 현우와 별개로, 가장자리에 현우와 같은 머리·회색 셔츠·카키색 바지의 비탈 청년이 추가되어 인물이 중복된다. / [gpt-high] 지정된 창·도끼·활 외에 여러 병사에게 총기를 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8_sel.png",
    "asset_id": "36c176a1-fc86-4102-9890-fdfb373a8b27",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-17be-7dd8-9d30-399949e4e871",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S56sh8"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S56sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:37:06.417945+00:00",
  "fingerprint": "e47d18a6049192de3509e7718f012f14f86efd596d3cc9cdaeca8751647a9ddd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S56sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S56sh14_sel.png",
  "source_sha256": "12fb4c3fd83beeb435e3fd1a5caefd193266d67b793a294b34c1a64eece13072",
  "file": "S56sh14_cine.png",
  "staged_sha256": "80fc3e068865d17f6254d04be7b6ea4a57c45664f2a916f865324ad56b4f2c83",
  "latency_ms": 11448
 },
 "S57sh3::signage": {
  "fp": "8b9583166e8fa63b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::416cbd07cc13fbe5": {
  "subjects": [],
  "subject_text": "익산 한옥마을 거리와 공터\n낡은 한옥들이 이어진 거리와 넓게 트인 마을 공터. 기와지붕이 겹쳐 보이며, 건물 사이로 높은 물탱크 탑이 솟아 있다.",
  "identity": "canonical",
  "scope_id": "L218",
  "scope_role": "location_exterior",
  "scope_sha": "212f1a87b313475d"
 },
 "S57sh3::bgfirst_bg": {
  "input_fingerprint": "3c72f26b33c2e5bf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh3__bgfirst_bg.png",
  "asset_id": "24a0e20b-6878-4179-adef-b500d62cac48",
  "input_asset_ids": [
   "eea69cbc-ad3c-4b30-99fa-fe8f252efa26",
   "e716b949-d46e-4479-90bb-be052601519b"
  ]
 },
 "S57sh3": {
  "input_fingerprint": "636c8cfbb33c267f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A barred cage is mounted on the transporting truck, and a tall water-tank-like tower stands in the village. Water bottles are being distributed nearby.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 하회탈 병사들과 사람들 right now, so 하회탈 병사들과 사람들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 하회탈 병사들과 사람들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A barred cage is mounted on the transporting truck, and a tall water-tank-like tower stands in the village. Water bottles are being distributed nearby.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 하회탈 병사들과 사람들 right now, so 하회탈 병사들과 사람들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 하회탈 병사들과 사람들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A barred cage is mounted on the transporting truck, and a tall water-tank-like tower stands in the village. Water bottles are being distributed nearby.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 하회탈 병사들과 사람들 right now, so 하회탈 병사들과 사람들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 하회탈 병사들과 사람들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh3__bgfirst_bg.png",
     "asset_id": "24a0e20b-6878-4179-adef-b500d62cac48",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S57sh3.png",
     "asset_id": "eea69cbc-ad3c-4b30-99fa-fe8f252efa26",
     "role": "conti_light"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_ruined_village_square_sel.png",
     "asset_id": "e716b949-d46e-4479-90bb-be052601519b",
     "role": "location_seed_bg"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "군중의 시선과 무기가 화면 왼쪽 전경을 향하고 있으나, 대상이 되는 트럭의 실체가 명확하지 않음.",
    "built_space": "마을 배경은 전반적으로 일치하나, 교회 앞 흰색 조각상이 있어야 할 자리에 회색 옷을 입은 사람이 서 있음.",
    "entities": "하회탈을 쓴 군중과 무기는 존재하나, 트럭의 창살은 화면 좌측 가장자리에 흐릿하게만 나타남.",
    "hard_violations": [
     "[gemini-pro] 지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
    ],
    "physics": "인물들은 지면에 안정적으로 서 있으며, 무기를 쥐고 있는 손의 지탱 상태는 정상적임."
   },
   {
    "label": "B",
    "direction": "군중 전체가 화면 우측의 트럭과 카메라를 향해 뚜렷하게 시선과 무기를 겨누고 있음.",
    "built_space": "교회, 예수상, 급수탑, 건물 등 참조 이미지의 고정 구조물들이 정확한 위치와 형태로 재현됨.",
    "entities": "트럭의 창살과 갇힌 사람들, 무기를 든 군중이 잘 나타나나, 다수의 인물이 하회탈을 쓰지 않고 손이나 머리 위에 들고 있음.",
    "hard_violations": [
     "[gpt-high] 지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
    ],
    "physics": "땅에 서 있는 인물들의 무게 중심과 사물을 쥐고 있는 손의 형태 등 물리적 지탱이 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 위치의 조각상을 실제 사람으로 변형한 치명적 오류가 있으며, 트럭 내부에서 바라보는 구도 지시를 거의 구현하지 못했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "트럭의 쇠창살 구도와 배경 장소를 정확히 재현했으나, 일부 군중이 지시와 다르게 탈을 얼굴에 쓰지 않고 손에 들고 있는 점이 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "군중의 시선과 무기가 화면 왼쪽 전경을 향하고 있으나, 대상이 되는 트럭의 실체가 명확하지 않음.",
        "built_space": "마을 배경은 전반적으로 일치하나, 교회 앞 흰색 조각상이 있어야 할 자리에 회색 옷을 입은 사람이 서 있음.",
        "entities": "하회탈을 쓴 군중과 무기는 존재하나, 트럭의 창살은 화면 좌측 가장자리에 흐릿하게만 나타남.",
        "hard_violations": [
         "지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
        ],
        "physics": "인물들은 지면에 안정적으로 서 있으며, 무기를 쥐고 있는 손의 지탱 상태는 정상적임."
       },
       {
        "label": "B",
        "direction": "군중 전체가 화면 우측의 트럭과 카메라를 향해 뚜렷하게 시선과 무기를 겨누고 있음.",
        "built_space": "교회, 예수상, 급수탑, 건물 등 참조 이미지의 고정 구조물들이 정확한 위치와 형태로 재현됨.",
        "entities": "트럭의 창살과 갇힌 사람들, 무기를 든 군중이 잘 나타나나, 다수의 인물이 하회탈을 쓰지 않고 손이나 머리 위에 들고 있음.",
        "hard_violations": [],
        "physics": "땅에 서 있는 인물들의 무게 중심과 사물을 쥐고 있는 손의 형태 등 물리적 지탱이 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 위치의 조각상을 실제 사람으로 변형한 치명적 오류가 있으며, 트럭 내부에서 바라보는 구도 지시를 거의 구현하지 못했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "트럭의 쇠창살 구도와 배경 장소를 정확히 재현했으나, 일부 군중이 지시와 다르게 탈을 얼굴에 쓰지 않고 손에 들고 있는 점이 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "군중의 시선과 무기가 화면 왼쪽 전경을 향하고 있으나, 대상이 되는 트럭의 실체가 명확하지 않음.",
        "built_space": "마을 배경은 전반적으로 일치하나, 교회 앞 흰색 조각상이 있어야 할 자리에 회색 옷을 입은 사람이 서 있음.",
        "entities": "하회탈을 쓴 군중과 무기는 존재하나, 트럭의 창살은 화면 좌측 가장자리에 흐릿하게만 나타남.",
        "hard_violations": [
         "지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
        ],
        "physics": "인물들은 지면에 안정적으로 서 있으며, 무기를 쥐고 있는 손의 지탱 상태는 정상적임."
       },
       {
        "label": "B",
        "direction": "군중 전체가 화면 우측의 트럭과 카메라를 향해 뚜렷하게 시선과 무기를 겨누고 있음.",
        "built_space": "교회, 예수상, 급수탑, 건물 등 참조 이미지의 고정 구조물들이 정확한 위치와 형태로 재현됨.",
        "entities": "트럭의 창살과 갇힌 사람들, 무기를 든 군중이 잘 나타나나, 다수의 인물이 하회탈을 쓰지 않고 손이나 머리 위에 들고 있음.",
        "hard_violations": [],
        "physics": "땅에 서 있는 인물들의 무게 중심과 사물을 쥐고 있는 손의 형태 등 물리적 지탱이 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "트럭을 향한 함성과 무기 동작, 마을 구조는 잘 드러나지만 카메라가 철창 밖에서 차량 측면을 보는 구도여서 지정된 철창 내부 관찰 시점을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "철창 안에서 내려다보는 와이드 구도와 트럭 쪽으로 무기를 치켜든 군중을 충실히 구현하지만, 일부 정면형 가면과 병사·주민의 불분명한 구별은 아쉽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽과 중앙의 군중은 오른쪽 트럭 철창과 그 안의 사람들을 향해 얼굴을 들고 함성을 지른다. 앞사람의 몽둥이는 왼쪽 위로, 여러 창은 서로 다른 각도로 위를 향한다. 무기를 높이 드는 행동과 함성의 대상인 트럭의 관계는 명확하다.",
        "built_space": "왼쪽에 파란 지붕 교회 한 채와 십자가 하나, 입구 앞 조각상 하나가 있고, 뒤에는 녹슨 원통형 물탱크 탑 하나가 보인다. 중앙 비탈길과 낡은 저층 건물들은 참조 장소에 부합한다. 그러나 차량 철창이 오른쪽 화면의 거의 절반을 차지하고 측면 거울과 외벽이 함께 보이며, 카메라는 철창 밖에서 차량 옆면과 내부를 들여다보는 위치로 읽힌다. 철창 안에서 가느다란 가장자리 장애물 너머 군중을 보는 구도가 아니다.",
        "entities": "군중은 동아시아계 외양의 성인 남성이 주를 이루며, 하회탈 형태의 갈색 가면과 낡은 작업복을 착용한다. 일부는 탈을 얼굴에 쓰지 않고 이마 위로 올리거나 손에 들어 보인다. 창과 나무 몽둥이, 철창 차량, 물탱크 탑이 식별된다. 철창 안에는 어두운 사람 형상이 보인다. 병사와 주민의 복장 차이는 뚜렷하지 않으며 물병 배급은 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
        ],
        "physics": "앞쪽 몽둥이는 올라간 손이 손잡이를 감싸 지지하고, 창들도 손으로 자루를 붙잡고 있다. 들어 올린 가면 역시 손이 받친다. 뒤쪽 인물의 발은 길에 닿아 있으며, 앞쪽 인물들은 하체가 잘렸지만 상체의 기울기와 팔 동작에 명백한 부유나 불가능한 관절은 없다. 철창은 차량 차체에 고정되어 있다."
       },
       {
        "label": "B",
        "direction": "전경 군중은 대체로 얼굴을 위로 들어 카메라가 있는 높은 트럭 쪽에 함성을 보낸다. 일부는 옆 사람이나 트럭의 다른 지점을 보는 듯해 시선이 한 점에 복제되어 있지 않다. 몽둥이는 좌우 위쪽으로 벌어지고 창은 수직과 사선이 섞여 있어, 무기를 높이 치켜드는 동작을 보여준다.",
        "built_space": "왼쪽 가장자리의 굵은 세로 철창과 왼쪽 아래를 비스듬히 가로지르는 가로대 너머로 군중이 펼쳐져, 차량 철창 안의 높은 관찰 위치가 성립한다. 왼쪽 교회 한 채, 십자가 하나, 입구 앞 팔을 벌린 조각상 하나, 그 옆 원통형 물탱크 탑 하나가 보인다. 중앙 오르막길, 오른쪽 큰 창고와 가까운 기와지붕도 참조의 주요 배치를 따른다. 탑의 폭과 일부 건물 형태·간격은 참조와 차이가 있지만 동일 장소의 핵심 구조는 유지된다.",
        "entities": "동아시아계 외양의 성인 남녀 군중이 낡은 회색·갈색·남색 옷을 입고 있으며, 다수는 갈색 하회탈 형태의 가면을 얼굴에 착용한다. 맨얼굴인 주민들도 섞여 있다. 창날 달린 장대와 나무 몽둥이가 구별되며, 앞쪽 철창과 뒤쪽 물탱크 탑이 보인다. 병사를 특정할 복장 표지는 약하다. 트럭 차체와 포로, 물병 배급은 이 구도에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "전경의 창과 몽둥이에는 각각 자루를 움켜쥔 손과 연결된 팔이 보인다. 군중의 들린 팔, 젖혀진 목, 기울어진 몸통은 지상에서 환호하는 동작으로 가능하다. 중·후경 인물들은 발로 길을 딛고 있고, 전경 인물들의 잘린 하체를 부유로 볼 근거는 없다. 앞쪽 철창의 세로대와 가로대는 서로 연결되어 있으며 지지 없는 물체는 확인되지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "트럭을 향한 함성과 무기 동작, 마을 구조는 잘 드러나지만 카메라가 철창 밖에서 차량 측면을 보는 구도여서 지정된 철창 내부 관찰 시점을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "철창 안에서 내려다보는 와이드 구도와 트럭 쪽으로 무기를 치켜든 군중을 충실히 구현하지만, 일부 정면형 가면과 병사·주민의 불분명한 구별은 아쉽다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽과 중앙의 군중은 오른쪽 트럭 철창과 그 안의 사람들을 향해 얼굴을 들고 함성을 지른다. 앞사람의 몽둥이는 왼쪽 위로, 여러 창은 서로 다른 각도로 위를 향한다. 무기를 높이 드는 행동과 함성의 대상인 트럭의 관계는 명확하다.",
        "built_space": "왼쪽에 파란 지붕 교회 한 채와 십자가 하나, 입구 앞 조각상 하나가 있고, 뒤에는 녹슨 원통형 물탱크 탑 하나가 보인다. 중앙 비탈길과 낡은 저층 건물들은 참조 장소에 부합한다. 그러나 차량 철창이 오른쪽 화면의 거의 절반을 차지하고 측면 거울과 외벽이 함께 보이며, 카메라는 철창 밖에서 차량 옆면과 내부를 들여다보는 위치로 읽힌다. 철창 안에서 가느다란 가장자리 장애물 너머 군중을 보는 구도가 아니다.",
        "entities": "군중은 동아시아계 외양의 성인 남성이 주를 이루며, 하회탈 형태의 갈색 가면과 낡은 작업복을 착용한다. 일부는 탈을 얼굴에 쓰지 않고 이마 위로 올리거나 손에 들어 보인다. 창과 나무 몽둥이, 철창 차량, 물탱크 탑이 식별된다. 철창 안에는 어두운 사람 형상이 보인다. 병사와 주민의 복장 차이는 뚜렷하지 않으며 물병 배급은 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
        ],
        "physics": "앞쪽 몽둥이는 올라간 손이 손잡이를 감싸 지지하고, 창들도 손으로 자루를 붙잡고 있다. 들어 올린 가면 역시 손이 받친다. 뒤쪽 인물의 발은 길에 닿아 있으며, 앞쪽 인물들은 하체가 잘렸지만 상체의 기울기와 팔 동작에 명백한 부유나 불가능한 관절은 없다. 철창은 차량 차체에 고정되어 있다."
       },
       {
        "label": "A",
        "direction": "전경 군중은 대체로 얼굴을 위로 들어 카메라가 있는 높은 트럭 쪽에 함성을 보낸다. 일부는 옆 사람이나 트럭의 다른 지점을 보는 듯해 시선이 한 점에 복제되어 있지 않다. 몽둥이는 좌우 위쪽으로 벌어지고 창은 수직과 사선이 섞여 있어, 무기를 높이 치켜드는 동작을 보여준다.",
        "built_space": "왼쪽 가장자리의 굵은 세로 철창과 왼쪽 아래를 비스듬히 가로지르는 가로대 너머로 군중이 펼쳐져, 차량 철창 안의 높은 관찰 위치가 성립한다. 왼쪽 교회 한 채, 십자가 하나, 입구 앞 팔을 벌린 조각상 하나, 그 옆 원통형 물탱크 탑 하나가 보인다. 중앙 오르막길, 오른쪽 큰 창고와 가까운 기와지붕도 참조의 주요 배치를 따른다. 탑의 폭과 일부 건물 형태·간격은 참조와 차이가 있지만 동일 장소의 핵심 구조는 유지된다.",
        "entities": "동아시아계 외양의 성인 남녀 군중이 낡은 회색·갈색·남색 옷을 입고 있으며, 다수는 갈색 하회탈 형태의 가면을 얼굴에 착용한다. 맨얼굴인 주민들도 섞여 있다. 창날 달린 장대와 나무 몽둥이가 구별되며, 앞쪽 철창과 뒤쪽 물탱크 탑이 보인다. 병사를 특정할 복장 표지는 약하다. 트럭 차체와 포로, 물병 배급은 이 구도에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "전경의 창과 몽둥이에는 각각 자루를 움켜쥔 손과 연결된 팔이 보인다. 군중의 들린 팔, 젖혀진 목, 기울어진 몸통은 지상에서 환호하는 동작으로 가능하다. 중·후경 인물들은 발로 길을 딛고 있고, 전경 인물들의 잘린 하체를 부유로 볼 근거는 없다. 앞쪽 철창의 세로대와 가로대는 서로 연결되어 있으며 지지 없는 물체는 확인되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
    ],
    "B": [
     "[gpt-high] 지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "지정된 위치의 조각상을 실제 사람으로 변형한 치명적 오류가 있으며, 트럭 내부에서 바라보는 구도 지시를 거의 구현하지 못했습니다.  ★위반: [gemini-pro] 지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "트럭의 쇠창살 구도와 배경 장소를 정확히 재현했으나, 일부 군중이 지시와 다르게 탈을 얼굴에 쓰지 않고 손에 들고 있는 점이 아쉽습니다.  ★위반: [gpt-high] 지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_ruined_village_square_sel.png",
    "asset_id": "e716b949-d46e-4479-90bb-be052601519b",
    "role": "location_seed_bg"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-1979-7f87-95f5-4c2072351bcf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh3__bgfirst_bg.png",
   "bg_asset_id": "24a0e20b-6878-4179-adef-b500d62cac48",
   "bg_record_key": "S57sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "seed_bg"
  },
  "ref_mode": "seed-bg+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S57sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:25:38.622807+00:00",
  "fingerprint": "63dfc1dcaa1ab310f5272800f76bb0887fde779f627167ec870a093876519e97",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S57sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S57sh3_sel.png",
  "source_sha256": "e5bd12d9985e7bc18dc655de2f32c2a16c454b386f85056cdc0700e742a968d2",
  "file": "S57sh3_cine.png",
  "staged_sha256": "42b2402c2732d9106a0e95f3b77f1d8db77a83979dfb6b44a322f8038becd1dc",
  "latency_ms": 59883
 },
 "S57sh5::signage": {
  "fp": "4be797c62f935eee",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::church_arena_forecourt": {
  "input_fingerprint": "4d1a3d8873f88228",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "church_arena_forecourt",
    "tags": [
     "S57sh5",
     "S57sh7",
     "S60sh4"
    ]
   },
   "context_sig": "075c46da6f3309d3"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 마을 성당 앞 격투장과 관중석: 원형의 넓은 흙바닥 투기장과 이를 내려다보는 폭발 현장. (특징: 모래와 흙이 깔린 평평한 원형 공터; 테두리를 밝히는 횃불 조명들; 육중하고 기계 장갑이 덧대어진 개조형 전투 병기(B-200); 팔에 부착된 회전형 고사포 총구; 관중석에 피어오르는 박격포 폭발 화염과 뽀얀 흙먼지; 방진복을 입은 최신 용병들) / 익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 성당인 듯 보이는 건물 앞에 트럭이 서자\n- 두 팔을 벌린 예수 동상에도 하회탈 가면이 씌워진.\n- 성당 앞 마을 한가운데 만들어진 격투장.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 마을 성당 앞 격투장과 관중석: 원형의 넓은 흙바닥 투기장과 이를 내려다보는 폭발 현장. (특징: 모래와 흙이 깔린 평평한 원형 공터; 테두리를 밝히는 횃불 조명들; 육중하고 기계 장갑이 덧대어진 개조형 전투 병기(B-200); 팔에 부착된 회전형 고사포 총구; 관중석에 피어오르는 박격포 폭발 화염과 뽀얀 흙먼지; 방진복을 입은 최신 용병들) / 익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 성당인 듯 보이는 건물 앞에 트럭이 서자\n- 두 팔을 벌린 예수 동상에도 하회탈 가면이 씌워진.\n- 성당 앞 마을 한가운데 만들어진 격투장.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_church_arena_forecourt_18a654.png",
  "asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350",
  "input_asset_ids": [
   "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a"
  ],
  "origin_tag": "S57sh5",
  "place_text": "At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.",
  "origin_inputs": {
   "place_text": "At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.",
   "time_of_day_en": "day",
   "conti_asset_id": "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a"
  }
 },
 "S57sh5::bgfirst_bg": {
  "input_fingerprint": "28fd29c710e31d3e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5__bgfirst_bg.png",
  "asset_id": "5771d1f0-905b-4ad6-80b0-1c684d75f479",
  "input_asset_ids": [
   "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a",
   "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350"
  ]
 },
 "S57sh5": {
  "input_fingerprint": "f37b75b2aca6ca5b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck and its barred cage are stopped outside the church-like building. A Hahoe mask covers the face of the outstretched-armed Jesus statue. 현우: He is bound while being unloaded from the truck, retaining his treated injuries. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck and its barred cage are stopped outside the church-like building. A Hahoe mask covers the face of the outstretched-armed Jesus statue. 현우: He is bound while being unloaded from the truck, retaining his treated injuries. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck and its barred cage are stopped outside the church-like building. A Hahoe mask covers the face of the outstretched-armed Jesus statue. 현우: He is bound while being unloaded from the truck, retaining his treated injuries. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5__bgfirst_bg.png",
     "asset_id": "5771d1f0-905b-4ad6-80b0-1c684d75f479",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S57sh5.png",
     "asset_id": "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_church_arena_forecourt_18a654.png",
     "asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "지상의 병사가 현우를 트럭 밖(왼쪽)으로 끌어당기고 있으며, 트럭 안의 병사도 현우의 결박된 팔을 잡고 있음.",
    "built_space": "배경 왼쪽에 성당과 예수상이 위치하며, 오른쪽 중경에 트럭 적재함이 대각선으로 배치됨.",
    "entities": "현우의 외모와 복장은 참조와 일치하며 병사들은 하회탈을 썼으나, 지시된 바와 달리 예수상에는 하회탈이 없음.",
    "hard_violations": [],
    "physics": "현우는 오른발을 땅에 딛고 있으며 양쪽 병사들의 손에 의해 물리적으로 자연스럽게 지탱됨."
   },
   {
    "label": "B",
    "direction": "현우가 트럭 밖으로 쏟아지듯 떨어지고 있으며, 트럭 안의 병사가 그의 결박을 붙잡고 있음.",
    "built_space": "배경 왼쪽에 성당과 예수상이 위치하고, 오른쪽 중경에 트럭 적재함 가장자리가 대각선으로 보임.",
    "entities": "현우의 외모는 일치하고 예수상에도 하회탈이 씌워져 있으나, 현우 뒤에 정체불명의 하반신이 존재함.",
    "hard_violations": [
     "[gemini-pro] 현우의 다리 뒤편 지면에 상체가 없는 정체불명의 다리 한 쌍(extra bodies)이 서 있음.",
     "[gpt-high] 철창 안에 인출 동작과 무관한 추가 인물들이 등장하여, 숏 텍스트가 지정하지 않은 사람을 넣지 말라는 제한을 위반합니다."
    ],
    "physics": "현우는 공중에 뜬 상태로 트럭 안 병사의 그립에 의해 상체만 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "예수상에 하회탈이 누락된 점은 아쉽지만, 현우를 트럭 밖으로 끌어당기는 동작과 전반적인 구도를 물리적 오류 없이 안정적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "예수상에 하회탈을 씌운 디테일은 좋으나, 현우의 하반신 뒤로 상체가 없는 다리 한 쌍이 생성되는 치명적인 신체 렌더링 오류가 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지상의 병사가 현우를 트럭 밖(왼쪽)으로 끌어당기고 있으며, 트럭 안의 병사도 현우의 결박된 팔을 잡고 있음.",
        "built_space": "배경 왼쪽에 성당과 예수상이 위치하며, 오른쪽 중경에 트럭 적재함이 대각선으로 배치됨.",
        "entities": "현우의 외모와 복장은 참조와 일치하며 병사들은 하회탈을 썼으나, 지시된 바와 달리 예수상에는 하회탈이 없음.",
        "hard_violations": [],
        "physics": "현우는 오른발을 땅에 딛고 있으며 양쪽 병사들의 손에 의해 물리적으로 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우가 트럭 밖으로 쏟아지듯 떨어지고 있으며, 트럭 안의 병사가 그의 결박을 붙잡고 있음.",
        "built_space": "배경 왼쪽에 성당과 예수상이 위치하고, 오른쪽 중경에 트럭 적재함 가장자리가 대각선으로 보임.",
        "entities": "현우의 외모는 일치하고 예수상에도 하회탈이 씌워져 있으나, 현우 뒤에 정체불명의 하반신이 존재함.",
        "hard_violations": [
         "현우의 다리 뒤편 지면에 상체가 없는 정체불명의 다리 한 쌍(extra bodies)이 서 있음."
        ],
        "physics": "현우는 공중에 뜬 상태로 트럭 안 병사의 그립에 의해 상체만 지탱되고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "예수상에 하회탈이 누락된 점은 아쉽지만, 현우를 트럭 밖으로 끌어당기는 동작과 전반적인 구도를 물리적 오류 없이 안정적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "예수상에 하회탈을 씌운 디테일은 좋으나, 현우의 하반신 뒤로 상체가 없는 다리 한 쌍이 생성되는 치명적인 신체 렌더링 오류가 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "지상의 병사가 현우를 트럭 밖(왼쪽)으로 끌어당기고 있으며, 트럭 안의 병사도 현우의 결박된 팔을 잡고 있음.",
        "built_space": "배경 왼쪽에 성당과 예수상이 위치하며, 오른쪽 중경에 트럭 적재함이 대각선으로 배치됨.",
        "entities": "현우의 외모와 복장은 참조와 일치하며 병사들은 하회탈을 썼으나, 지시된 바와 달리 예수상에는 하회탈이 없음.",
        "hard_violations": [],
        "physics": "현우는 오른발을 땅에 딛고 있으며 양쪽 병사들의 손에 의해 물리적으로 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우가 트럭 밖으로 쏟아지듯 떨어지고 있으며, 트럭 안의 병사가 그의 결박을 붙잡고 있음.",
        "built_space": "배경 왼쪽에 성당과 예수상이 위치하고, 오른쪽 중경에 트럭 적재함 가장자리가 대각선으로 보임.",
        "entities": "현우의 외모는 일치하고 예수상에도 하회탈이 씌워져 있으나, 현우 뒤에 정체불명의 하반신이 존재함.",
        "hard_violations": [
         "현우의 다리 뒤편 지면에 상체가 없는 정체불명의 다리 한 쌍(extra bodies)이 서 있음."
        ],
        "physics": "현우는 공중에 뜬 상태로 트럭 안 병사의 그립에 의해 상체만 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "추가 인물들이 등장하는 중대 위반이 있고, 전신 위주의 넓은 구도와 적재함 안에서 뒤로 잡는 동작이 병사들의 상체 중심 인출 장면에 맞지 않습니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 하회탈 병사의 상체와 현우를 밖으로 당기는 힘, 우측 적재함 경계를 미디엄 구도로 구현했으나 예수상의 하회탈과 치료된 부상은 명확하지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 머리와 상체를 화면 왼쪽 아래, 트럭 바깥으로 숙이고 있습니다. 그러나 주된 병사는 적재함 안에서 현우의 뒤로 묶인 손목을 잡아, 바깥으로 끌어당기기보다 뒤에서 붙들거나 내보내는 동작으로 보입니다. 병사의 얼굴은 현우 쪽을 향하며, 바깥에서 당기는 병사는 보이지 않습니다.",
        "built_space": "오른쪽에 철창 적재함 하나와 열린 출입구 하나, 아래로 내려진 후면 판 하나가 보입니다. 왼쪽 배경에는 석조 성당 하나, 계단과 받침대 위의 두 팔을 벌린 예수상 하나가 있습니다. 장소의 재료와 주요 구조는 참조와 유사하지만 성당과 동상이 화면에서 크게 차지합니다. 현우의 신발까지 포함한 넓은 구도로, 요청한 병사들의 굳은 상체 중심 미디엄 숏과 다릅니다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 회색 셔츠, 녹색 작업 바지, 갈색 허리띠가 참조와 대체로 맞습니다. 얼굴에 상처가 있고 손목은 뒤로 묶여 있지만 치료 흔적은 뚜렷하지 않습니다. 현우를 잡는 하회탈 병사 외에 철창 안에 맨얼굴 인물, 하회탈을 쓴 결박 인물, 오른쪽 끝에 잘린 인물이 더 보입니다. 이들은 숏 텍스트의 인출 동작에 참여하지 않는 추가 인물입니다. 예수상에는 하회탈이 분명히 보입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "철창 안에 인출 동작과 무관한 추가 인물들이 등장하여, 숏 텍스트가 지정하지 않은 사람을 넣지 말라는 제한을 위반합니다."
        ],
        "physics": "현우의 두 발은 지면에서 떨어져 있지만 병사의 두 손이 묶인 손목을 붙잡고 있으며, 바로 뒤에 내려온 출발 지점인 적재함 문턱이 있습니다. 따라서 근거 없는 공중 부유라기보다 문턱에서 떨어지는 몸을 뒤에서 붙든 순간으로 읽힙니다. 병사의 한쪽 부츠는 적재함 경계에 놓여 있습니다. 다만 이 지지와 힘의 방향은 바깥에서 거칠게 끌어당기는 지정 동작과 다릅니다."
       },
       {
        "label": "B",
        "direction": "왼쪽 병사는 현우의 팔을 붙잡고 자기 몸 쪽, 즉 트럭 바깥으로 당기고 있습니다. 오른쪽 병사는 적재함 경계에서 현우의 반대쪽 팔을 잡고 몸을 앞으로 숙입니다. 현우의 머리와 몸통은 오른쪽 트럭에서 왼쪽 앞쪽으로 쏠리고 시선은 아래를 향합니다. 두 병사의 얼굴 방향도 현우와 붙잡은 팔을 향해 있어 인출의 대상과 힘의 방향이 명확합니다.",
        "built_space": "철창 트럭 하나의 적재함 경계가 화면 중간 오른쪽을 차지하고, 금속 측판이 오른쪽으로 이어지는 사선을 만듭니다. 왼쪽 병사는 트럭 밖, 오른쪽 병사는 적재함 가장자리, 현우는 그 경계를 넘어오는 위치입니다. 배경에는 종탑이 있는 석조 성당 하나, 중앙 계단, 받침대 위 예수상 하나와 목제 난간이 보입니다. 참조 장소의 주요 배치와 재료가 유지되며 배경 크기도 인물과 경쟁하지 않습니다. 인물의 상체와 허벅지 일부를 담는 구도로 A보다 요청한 미디엄 숏에 가깝습니다.",
        "entities": "현우의 앳된 동아시아계 얼굴, 헝클어진 검은 머리, 회색 셔츠, 녹색 바지와 갈색 허리띠가 참조와 잘 맞습니다. 볼에 상처가 있으며 팔 주변에 결박용 밧줄이 보이지만 치료된 부상의 구체적인 흔적은 확인하기 어렵습니다. 하회탈을 쓴 병사 두 명만 현우와 함께 등장하며, 가려진 얼굴로 병사의 나이나 민족성을 확인할 수는 없습니다. 배경 예수상은 두 팔을 벌리고 있지만 작은 얼굴에서 하회탈 착용은 분명하지 않습니다. 신발 속 카드는 노출되지 않으며 읽을 수 있는 글자도 없습니다.",
        "hard_violations": [],
        "physics": "현우의 양팔에는 병사들의 손과 팔이 접촉하고 있어 앞으로 쏠리는 상체를 붙잡는 힘이 보입니다. 오른쪽 병사는 굽힌 하체를 적재함 가장자리에 두고 몸을 지탱하며, 왼쪽 병사는 바깥에서 몸을 뒤로 기울여 당깁니다. 현우의 한쪽 다리는 뒤로 굽혀져 내려오는 동작을 보이고 다른 다리의 발은 화면 밖입니다. 보이는 접촉과 적재함 문턱이 인출 동작의 지지와 출발점을 제공하므로, 근거 없이 떠 있는 자세는 아닙니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "추가 인물들이 등장하는 중대 위반이 있고, 전신 위주의 넓은 구도와 적재함 안에서 뒤로 잡는 동작이 병사들의 상체 중심 인출 장면에 맞지 않습니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 하회탈 병사의 상체와 현우를 밖으로 당기는 힘, 우측 적재함 경계를 미디엄 구도로 구현했으나 예수상의 하회탈과 치료된 부상은 명확하지 않습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 머리와 상체를 화면 왼쪽 아래, 트럭 바깥으로 숙이고 있습니다. 그러나 주된 병사는 적재함 안에서 현우의 뒤로 묶인 손목을 잡아, 바깥으로 끌어당기기보다 뒤에서 붙들거나 내보내는 동작으로 보입니다. 병사의 얼굴은 현우 쪽을 향하며, 바깥에서 당기는 병사는 보이지 않습니다.",
        "built_space": "오른쪽에 철창 적재함 하나와 열린 출입구 하나, 아래로 내려진 후면 판 하나가 보입니다. 왼쪽 배경에는 석조 성당 하나, 계단과 받침대 위의 두 팔을 벌린 예수상 하나가 있습니다. 장소의 재료와 주요 구조는 참조와 유사하지만 성당과 동상이 화면에서 크게 차지합니다. 현우의 신발까지 포함한 넓은 구도로, 요청한 병사들의 굳은 상체 중심 미디엄 숏과 다릅니다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 회색 셔츠, 녹색 작업 바지, 갈색 허리띠가 참조와 대체로 맞습니다. 얼굴에 상처가 있고 손목은 뒤로 묶여 있지만 치료 흔적은 뚜렷하지 않습니다. 현우를 잡는 하회탈 병사 외에 철창 안에 맨얼굴 인물, 하회탈을 쓴 결박 인물, 오른쪽 끝에 잘린 인물이 더 보입니다. 이들은 숏 텍스트의 인출 동작에 참여하지 않는 추가 인물입니다. 예수상에는 하회탈이 분명히 보입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "철창 안에 인출 동작과 무관한 추가 인물들이 등장하여, 숏 텍스트가 지정하지 않은 사람을 넣지 말라는 제한을 위반합니다."
        ],
        "physics": "현우의 두 발은 지면에서 떨어져 있지만 병사의 두 손이 묶인 손목을 붙잡고 있으며, 바로 뒤에 내려온 출발 지점인 적재함 문턱이 있습니다. 따라서 근거 없는 공중 부유라기보다 문턱에서 떨어지는 몸을 뒤에서 붙든 순간으로 읽힙니다. 병사의 한쪽 부츠는 적재함 경계에 놓여 있습니다. 다만 이 지지와 힘의 방향은 바깥에서 거칠게 끌어당기는 지정 동작과 다릅니다."
       },
       {
        "label": "A",
        "direction": "왼쪽 병사는 현우의 팔을 붙잡고 자기 몸 쪽, 즉 트럭 바깥으로 당기고 있습니다. 오른쪽 병사는 적재함 경계에서 현우의 반대쪽 팔을 잡고 몸을 앞으로 숙입니다. 현우의 머리와 몸통은 오른쪽 트럭에서 왼쪽 앞쪽으로 쏠리고 시선은 아래를 향합니다. 두 병사의 얼굴 방향도 현우와 붙잡은 팔을 향해 있어 인출의 대상과 힘의 방향이 명확합니다.",
        "built_space": "철창 트럭 하나의 적재함 경계가 화면 중간 오른쪽을 차지하고, 금속 측판이 오른쪽으로 이어지는 사선을 만듭니다. 왼쪽 병사는 트럭 밖, 오른쪽 병사는 적재함 가장자리, 현우는 그 경계를 넘어오는 위치입니다. 배경에는 종탑이 있는 석조 성당 하나, 중앙 계단, 받침대 위 예수상 하나와 목제 난간이 보입니다. 참조 장소의 주요 배치와 재료가 유지되며 배경 크기도 인물과 경쟁하지 않습니다. 인물의 상체와 허벅지 일부를 담는 구도로 A보다 요청한 미디엄 숏에 가깝습니다.",
        "entities": "현우의 앳된 동아시아계 얼굴, 헝클어진 검은 머리, 회색 셔츠, 녹색 바지와 갈색 허리띠가 참조와 잘 맞습니다. 볼에 상처가 있으며 팔 주변에 결박용 밧줄이 보이지만 치료된 부상의 구체적인 흔적은 확인하기 어렵습니다. 하회탈을 쓴 병사 두 명만 현우와 함께 등장하며, 가려진 얼굴로 병사의 나이나 민족성을 확인할 수는 없습니다. 배경 예수상은 두 팔을 벌리고 있지만 작은 얼굴에서 하회탈 착용은 분명하지 않습니다. 신발 속 카드는 노출되지 않으며 읽을 수 있는 글자도 없습니다.",
        "hard_violations": [],
        "physics": "현우의 양팔에는 병사들의 손과 팔이 접촉하고 있어 앞으로 쏠리는 상체를 붙잡는 힘이 보입니다. 오른쪽 병사는 굽힌 하체를 적재함 가장자리에 두고 몸을 지탱하며, 왼쪽 병사는 바깥에서 몸을 뒤로 기울여 당깁니다. 현우의 한쪽 다리는 뒤로 굽혀져 내려오는 동작을 보이고 다른 다리의 발은 화면 밖입니다. 보이는 접촉과 적재함 문턱이 인출 동작의 지지와 출발점을 제공하므로, 근거 없이 떠 있는 자세는 아닙니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 현우의 다리 뒤편 지면에 상체가 없는 정체불명의 다리 한 쌍(extra bodies)이 서 있음.",
     "[gpt-high] 철창 안에 인출 동작과 무관한 추가 인물들이 등장하여, 숏 텍스트가 지정하지 않은 사람을 넣지 말라는 제한을 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "예수상에 하회탈이 누락된 점은 아쉽지만, 현우를 트럭 밖으로 끌어당기는 동작과 전반적인 구도를 물리적 오류 없이 안정적으로 구현함."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "예수상에 하회탈을 씌운 디테일은 좋으나, 현우의 하반신 뒤로 상체가 없는 다리 한 쌍이 생성되는 치명적인 신체 렌더링 오류가 발생함.  ★위반: [gemini-pro] 현우의 다리 뒤편 지면에 상체가 없는 정체불명의 다리 한 쌍(extra bodies)이 서 있음. / [gpt-high] 철창 안에 인출 동작과 무관한 추가 인물들이 등장하여, 숏 텍스트가 지정하지 않은 사람을 넣지 말라는 제한을 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_church_arena_forecourt_18a654.png",
    "asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-1ccc-7ae6-957f-9dfe6d964cfb",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5__bgfirst_bg.png",
   "bg_asset_id": "5771d1f0-905b-4ad6-80b0-1c684d75f479",
   "bg_record_key": "S57sh5::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "church_arena_forecourt",
   "groupbg_asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S57sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:38:25.856295+00:00",
  "fingerprint": "62bb4e3ef54175df262a626fe38a226b4919c501f8c6596140993f7c2836b8ce",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S57sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S57sh5_sel.png",
  "source_sha256": "65085154a3239a81e512561a32dcf8b1a71709d0917afe117cb9598c42bb5ca0",
  "file": "S57sh5_cine.png",
  "staged_sha256": "4d22157bbd927a7bfa55bd07e8b005d03dcd86fbcb00716f50d251578babf8ba",
  "latency_ms": 10626
 },
 "S57sh7::signage": {
  "fp": "d62bf2bea1640978",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S57sh7": {
  "input_fingerprint": "00c867d2f878f114",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양팔을 벌린 거대한 예수 동상의 얼굴에 기괴한 가면이 씌워져 있는 섬뜩한 광경.\n\nLOCATION (lock): Outside the village church, at the large outstretched-arm religious statue fitted with a mask. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Jesus statue (Arms outstretched, with a Hahoe mask covering its face) — Seen from below at a lateral three-quarter angle, with both arms legible; used as Primary architectural subject at the end of the tilt; Hahoe mask on the statue (Covering the statue's face) — Its face is visible obliquely above the camera rather than symmetrically head-on; used as Final point of attention within the wider statue composition; Church-like exterior (Surrounding the statue near the entrance) — Exterior portions remain visible around the upward view; used as Provides architectural scale and negative space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and restrained tonal contrast let the mask's incongruity carry the unease without introducing an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Jesus statue has both arms extended and a Hahoe mask over its face. The transport truck and barred cage remain outside the church-like building.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양팔을 벌린 거대한 예수 동상의 얼굴에 기괴한 가면이 씌워져 있는 섬뜩한 광경.\n\nLOCATION (lock): Outside the village church, at the large outstretched-arm religious statue fitted with a mask. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Jesus statue (Arms outstretched, with a Hahoe mask covering its face) — Seen from below at a lateral three-quarter angle, with both arms legible; used as Primary architectural subject at the end of the tilt; Hahoe mask on the statue (Covering the statue's face) — Its face is visible obliquely above the camera rather than symmetrically head-on; used as Final point of attention within the wider statue composition; Church-like exterior (Surrounding the statue near the entrance) — Exterior portions remain visible around the upward view; used as Provides architectural scale and negative space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and restrained tonal contrast let the mask's incongruity carry the unease without introducing an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Jesus statue has both arms extended and a Hahoe mask over its face. The transport truck and barred cage remain outside the church-like building.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양팔을 벌린 거대한 예수 동상의 얼굴에 기괴한 가면이 씌워져 있는 섬뜩한 광경.\n\nLOCATION (lock): Outside the village church, at the large outstretched-arm religious statue fitted with a mask. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Jesus statue (Arms outstretched, with a Hahoe mask covering its face) — Seen from below at a lateral three-quarter angle, with both arms legible; used as Primary architectural subject at the end of the tilt; Hahoe mask on the statue (Covering the statue's face) — Its face is visible obliquely above the camera rather than symmetrically head-on; used as Final point of attention within the wider statue composition; Church-like exterior (Surrounding the statue near the entrance) — Exterior portions remain visible around the upward view; used as Provides architectural scale and negative space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and restrained tonal contrast let the mask's incongruity carry the unease without introducing an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Jesus statue has both arms extended and a Hahoe mask over its face. The transport truck and barred cage remain outside the church-like building.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 동상을 위에서 아래로 내려다보고 있으며, 동상의 양팔은 옆으로 뻗어 있고 시선은 아래를 향함.",
    "built_space": "동상 뒤로 석조 교회 건물이 있으며, 화면 우측 하단에 철창이 달린 트럭이 일부 보이나 시점이 잘못됨.",
    "entities": "하회탈을 쓴 예수 동상, 교회 건물, 트럭 짐칸이 보이며, 사람은 등장하지 않음.",
    "hard_violations": [
     "[gemini-pro] 지시된 카메라 앵글('Seen from below', 'upward view')과 정반대로 위에서 내려다보는 하이 앵글을 적용함."
    ],
    "physics": "동상과 건물, 트럭 모두 지면에 안정적으로 배치되어 중력에 어긋남이 없음."
   },
   {
    "label": "B",
    "direction": "카메라는 아래에서 동상을 비스듬히 올려다보고 있으며, 동상의 양팔은 넓게 벌려져 있음.",
    "built_space": "레퍼런스와 일치하는 석조 교회 정면, 십자가 종탑, 계단, 목조 횃불 거치대가 동상 뒤로 배치되어 있으며 우측에 철창 트럭이 있음.",
    "entities": "양팔을 벌린 예수 동상, 얼굴을 덮은 하회탈, 석조 교회, 철창 트럭이 프롬프트와 레퍼런스대로 묘사되었으며 사람은 없음.",
    "hard_violations": [],
    "physics": "동상과 트럭 등 모든 구조물이 지면 위에 물리적으로 올바르게 안착되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 로우 앵글(Seen from below) 카메라 구도를 완벽히 구현했으며, 레퍼런스의 배경(교회, 트럭, 계단 등)을 사람 없이 정확하게 재현하여 프롬프트를 충실히 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시한 '아래에서 올려다보는 앵글(Seen from below)'을 완전히 위반하고 위에서 내려다보는 부감 앵글을 사용하여 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 동상을 위에서 아래로 내려다보고 있으며, 동상의 양팔은 옆으로 뻗어 있고 시선은 아래를 향함.",
        "built_space": "동상 뒤로 석조 교회 건물이 있으며, 화면 우측 하단에 철창이 달린 트럭이 일부 보이나 시점이 잘못됨.",
        "entities": "하회탈을 쓴 예수 동상, 교회 건물, 트럭 짐칸이 보이며, 사람은 등장하지 않음.",
        "hard_violations": [
         "지시된 카메라 앵글('Seen from below', 'upward view')과 정반대로 위에서 내려다보는 하이 앵글을 적용함."
        ],
        "physics": "동상과 건물, 트럭 모두 지면에 안정적으로 배치되어 중력에 어긋남이 없음."
       },
       {
        "label": "B",
        "direction": "카메라는 아래에서 동상을 비스듬히 올려다보고 있으며, 동상의 양팔은 넓게 벌려져 있음.",
        "built_space": "레퍼런스와 일치하는 석조 교회 정면, 십자가 종탑, 계단, 목조 횃불 거치대가 동상 뒤로 배치되어 있으며 우측에 철창 트럭이 있음.",
        "entities": "양팔을 벌린 예수 동상, 얼굴을 덮은 하회탈, 석조 교회, 철창 트럭이 프롬프트와 레퍼런스대로 묘사되었으며 사람은 없음.",
        "hard_violations": [],
        "physics": "동상과 트럭 등 모든 구조물이 지면 위에 물리적으로 올바르게 안착되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 로우 앵글(Seen from below) 카메라 구도를 완벽히 구현했으며, 레퍼런스의 배경(교회, 트럭, 계단 등)을 사람 없이 정확하게 재현하여 프롬프트를 충실히 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시한 '아래에서 올려다보는 앵글(Seen from below)'을 완전히 위반하고 위에서 내려다보는 부감 앵글을 사용하여 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 동상을 위에서 아래로 내려다보고 있으며, 동상의 양팔은 옆으로 뻗어 있고 시선은 아래를 향함.",
        "built_space": "동상 뒤로 석조 교회 건물이 있으며, 화면 우측 하단에 철창이 달린 트럭이 일부 보이나 시점이 잘못됨.",
        "entities": "하회탈을 쓴 예수 동상, 교회 건물, 트럭 짐칸이 보이며, 사람은 등장하지 않음.",
        "hard_violations": [
         "지시된 카메라 앵글('Seen from below', 'upward view')과 정반대로 위에서 내려다보는 하이 앵글을 적용함."
        ],
        "physics": "동상과 건물, 트럭 모두 지면에 안정적으로 배치되어 중력에 어긋남이 없음."
       },
       {
        "label": "B",
        "direction": "카메라는 아래에서 동상을 비스듬히 올려다보고 있으며, 동상의 양팔은 넓게 벌려져 있음.",
        "built_space": "레퍼런스와 일치하는 석조 교회 정면, 십자가 종탑, 계단, 목조 횃불 거치대가 동상 뒤로 배치되어 있으며 우측에 철창 트럭이 있음.",
        "entities": "양팔을 벌린 예수 동상, 얼굴을 덮은 하회탈, 석조 교회, 철창 트럭이 프롬프트와 레퍼런스대로 묘사되었으며 사람은 없음.",
        "hard_violations": [],
        "physics": "동상과 트럭 등 모든 구조물이 지면 위에 물리적으로 올바르게 안착되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "아래에서 비스듬히 올려다본 가면과 양팔, 석조 교회 외관이 핵심 구도를 충족하지만, 동상 하부가 잘려 와이드 숏의 여유는 부족합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "가면을 쓴 동상과 장소의 재질은 맞지만, 머리와 지면을 내려다보는 기울어진 하이앵글이 지정된 아래쪽 사선 시점을 정면으로 어깁니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "동상의 두 팔이 화면 좌우로 뻗고 손바닥은 앞쪽과 위쪽으로 열려 있습니다. 가면의 얼굴은 화면 오른쪽으로 조금 돌아가 있으며, 카메라보다 높은 위치에서 비스듬히 보입니다. 특정 인물이나 물체를 겨냥하는 동작은 없습니다.",
        "built_space": "동상 뒤 왼쪽에 석조 교회 한 채, 종탑 하나와 꼭대기 십자가 하나, 상부 첨두창 하나가 보입니다. 중앙 뒤에는 교회로 올라가는 돌계단이 있고 왼쪽에는 목재 난간과 화로 세 개가 보입니다. 오른쪽 아래에는 운송 트럭 한 대와 적재함의 철창이 일부 보입니다. 참고의 석재와 계단, 낡은 외부 시설은 이어지지만 동상이 프레임을 크게 채워 받침대와 입구의 정확한 관계는 확인하기 어렵습니다.",
        "entities": "긴 머리와 주름진 로브를 가진 석조 예수상 하나가 있으며 양팔이 모두 보입니다. 갈색의 입체적인 하회탈이 얼굴을 덮고 머리 둘레의 끈으로 고정되어 있습니다. 교회, 트럭, 철창이 보이고 살아 있는 사람은 없습니다. 읽을 수 있는 글자도 없습니다. 햇빛과 낡은 석재 표면은 참고의 낮 풍경과 부합합니다.",
        "hard_violations": [],
        "physics": "동상의 팔과 손은 몸통에 연결된 조각 구조이며 공중에 따로 떠 있는 부분은 없습니다. 하부 지지대는 화면 밖이지만 몸통이 하단으로 이어져 부유하는 형태로 보이지 않습니다. 가면은 머리를 두른 밧줄로 지지됩니다. 트럭은 지면 쪽에 놓여 있고 철창은 적재함에 연결되어 있습니다."
       },
       {
        "label": "B",
        "direction": "동상의 양팔과 열린 손바닥이 좌우로 뻗어 있습니다. 가면은 아래쪽을 향하며, 카메라는 동상 머리 윗면과 어깨 위를 내려다봅니다. 따라서 가면을 카메라 위에서 비스듬히 올려다보라는 방향 조건과 반대입니다.",
        "built_space": "석조 교회 전면에 하부 첨두창 두 개와 중앙 출입구 일부가 보이고, 동상 뒤에는 돌계단이 있습니다. 동상은 사각 받침대 하나 위에 서 있으며 오른쪽 아래에는 철창 적재함 하나가 보입니다. 종탑과 십자가는 프레임 밖입니다. 석조 벽과 기와 건물은 장소의 분위기를 잇지만, 넓게 드러난 지면과 받침대 윗면이 높은 카메라 위치를 명확히 보여줍니다.",
        "entities": "양팔을 벌린 석조 예수상 하나와 얼굴을 덮는 갈색 하회탈이 보입니다. 가면에는 좁은 눈, 돌출된 코와 벌어진 입이 입체적으로 표현되어 있습니다. 교회와 철창은 확인되지만 트럭의 운전실과 바퀴는 화면에 없어 차량 전체는 확인할 수 없습니다. 살아 있는 사람과 읽을 수 있는 글자는 없습니다. 낮의 자연광과 마모된 석재 표현은 조건에 맞습니다.",
        "hard_violations": [],
        "physics": "동상의 두 발은 받침대 윗면에 닿아 있고 받침대는 바닥 구조로 이어집니다. 팔은 로브와 몸통에 연결된 조각입니다. 가면에는 머리 옆으로 이어지는 고정 끈이 보입니다. 철창은 적재함 바닥 구조에 고정되어 있으며, 지지 없이 떠 있는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "아래에서 비스듬히 올려다본 가면과 양팔, 석조 교회 외관이 핵심 구도를 충족하지만, 동상 하부가 잘려 와이드 숏의 여유는 부족합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "가면을 쓴 동상과 장소의 재질은 맞지만, 머리와 지면을 내려다보는 기울어진 하이앵글이 지정된 아래쪽 사선 시점을 정면으로 어깁니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "동상의 두 팔이 화면 좌우로 뻗고 손바닥은 앞쪽과 위쪽으로 열려 있습니다. 가면의 얼굴은 화면 오른쪽으로 조금 돌아가 있으며, 카메라보다 높은 위치에서 비스듬히 보입니다. 특정 인물이나 물체를 겨냥하는 동작은 없습니다.",
        "built_space": "동상 뒤 왼쪽에 석조 교회 한 채, 종탑 하나와 꼭대기 십자가 하나, 상부 첨두창 하나가 보입니다. 중앙 뒤에는 교회로 올라가는 돌계단이 있고 왼쪽에는 목재 난간과 화로 세 개가 보입니다. 오른쪽 아래에는 운송 트럭 한 대와 적재함의 철창이 일부 보입니다. 참고의 석재와 계단, 낡은 외부 시설은 이어지지만 동상이 프레임을 크게 채워 받침대와 입구의 정확한 관계는 확인하기 어렵습니다.",
        "entities": "긴 머리와 주름진 로브를 가진 석조 예수상 하나가 있으며 양팔이 모두 보입니다. 갈색의 입체적인 하회탈이 얼굴을 덮고 머리 둘레의 끈으로 고정되어 있습니다. 교회, 트럭, 철창이 보이고 살아 있는 사람은 없습니다. 읽을 수 있는 글자도 없습니다. 햇빛과 낡은 석재 표면은 참고의 낮 풍경과 부합합니다.",
        "hard_violations": [],
        "physics": "동상의 팔과 손은 몸통에 연결된 조각 구조이며 공중에 따로 떠 있는 부분은 없습니다. 하부 지지대는 화면 밖이지만 몸통이 하단으로 이어져 부유하는 형태로 보이지 않습니다. 가면은 머리를 두른 밧줄로 지지됩니다. 트럭은 지면 쪽에 놓여 있고 철창은 적재함에 연결되어 있습니다."
       },
       {
        "label": "A",
        "direction": "동상의 양팔과 열린 손바닥이 좌우로 뻗어 있습니다. 가면은 아래쪽을 향하며, 카메라는 동상 머리 윗면과 어깨 위를 내려다봅니다. 따라서 가면을 카메라 위에서 비스듬히 올려다보라는 방향 조건과 반대입니다.",
        "built_space": "석조 교회 전면에 하부 첨두창 두 개와 중앙 출입구 일부가 보이고, 동상 뒤에는 돌계단이 있습니다. 동상은 사각 받침대 하나 위에 서 있으며 오른쪽 아래에는 철창 적재함 하나가 보입니다. 종탑과 십자가는 프레임 밖입니다. 석조 벽과 기와 건물은 장소의 분위기를 잇지만, 넓게 드러난 지면과 받침대 윗면이 높은 카메라 위치를 명확히 보여줍니다.",
        "entities": "양팔을 벌린 석조 예수상 하나와 얼굴을 덮는 갈색 하회탈이 보입니다. 가면에는 좁은 눈, 돌출된 코와 벌어진 입이 입체적으로 표현되어 있습니다. 교회와 철창은 확인되지만 트럭의 운전실과 바퀴는 화면에 없어 차량 전체는 확인할 수 없습니다. 살아 있는 사람과 읽을 수 있는 글자는 없습니다. 낮의 자연광과 마모된 석재 표현은 조건에 맞습니다.",
        "hard_violations": [],
        "physics": "동상의 두 발은 받침대 윗면에 닿아 있고 받침대는 바닥 구조로 이어집니다. 팔은 로브와 몸통에 연결된 조각입니다. 가면에는 머리 옆으로 이어지는 고정 끈이 보입니다. 철창은 적재함 바닥 구조에 고정되어 있으며, 지지 없이 떠 있는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.929,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.679,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시된 카메라 앵글('Seen from below', 'upward view')과 정반대로 위에서 내려다보는 하이 앵글을 적용함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 679
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시된 로우 앵글(Seen from below) 카메라 구도를 완벽히 구현했으며, 레퍼런스의 배경(교회, 트럭, 계단 등)을 사람 없이 정확하게 재현하여 프롬프트를 충실히 따랐습니다."
   },
   {
    "label": "A",
    "score": 679,
    "verdict_ko": "프롬프트에서 명시한 '아래에서 올려다보는 앵글(Seen from below)'을 완전히 위반하고 위에서 내려다보는 부감 앵글을 사용하여 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 지시된 카메라 앵글('Seen from below', 'upward view')과 정반대로 위에서 내려다보는 하이 앵글을 적용함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5_sel.png",
    "asset_id": "76330711-4803-4cb9-91d0-4ae955e482dd",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-218f-7b11-a19f-7b0a6e4d720a",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S57sh5"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S57sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:39:20.435537+00:00",
  "fingerprint": "d505b538669281a4032b8e9ec5b8bfb3b754589de86eb118b9099bb85b933165",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S57sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S57sh7_sel.png",
  "source_sha256": "ddb72ba94b0c9a61ae47256ad5273777db1a8507ef880f1507d06e8b86001861",
  "file": "S57sh7_cine.png",
  "staged_sha256": "e881ed6abac834e00deacef94e958a1b9450e48ca755926d00c7c72be2fab3df",
  "latency_ms": 10975
 },
 "S58sh5::signage": {
  "fp": "7e52833d11ea9ba5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::ae7a6cc9c201b720": {
  "subjects": [],
  "subject_text": "익산 마을 성당 예배당과 단상\n훼손된 고딕 양식 예배당. 높은 기둥과 2층 발코니, 앞쪽 단상이 있으며 스테인드글라스를 통과한 빛이 내부로 스며든다.",
  "identity": "canonical",
  "scope_id": "L220",
  "scope_role": "location_interior",
  "scope_sha": "73535dd8e86b426d"
 },
 "S58sh5::bgfirst_bg": {
  "input_fingerprint": "cadd1dfc5e244007",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5__bgfirst_bg.png",
  "asset_id": "31941b00-404b-4a16-90fe-5d0f9ded9490",
  "input_asset_ids": [
   "6d80eb92-6d35-4bd5-bfec-92b0a707c02a",
   "b25ec118-d5a5-4461-bcc9-1623494e1fcd"
  ]
 },
 "S58sh5": {
  "input_fingerprint": "7f00752dc09b76f8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Daylight filters through stained glass into the damaged church. Charlie is already confined in a steel cage behind the barred entrance, retaining the oversized straw hat, boots and colorful raincoat from his capture. 백산: He has a massive, obese build and wears regal clothing and an intact golden Hahoe mask. His pre-existing radiation-disfigured face remains concealed beneath it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Daylight filters through stained glass into the damaged church. Charlie is already confined in a steel cage behind the barred entrance, retaining the oversized straw hat, boots and colorful raincoat from his capture. 백산: He has a massive, obese build and wears regal clothing and an intact golden Hahoe mask. His pre-existing radiation-disfigured face remains concealed beneath it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Daylight filters through stained glass into the damaged church. Charlie is already confined in a steel cage behind the barred entrance, retaining the oversized straw hat, boots and colorful raincoat from his capture. 백산: He has a massive, obese build and wears regal clothing and an intact golden Hahoe mask. His pre-existing radiation-disfigured face remains concealed beneath it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5__bgfirst_bg.png",
     "asset_id": "31941b00-404b-4a16-90fe-5d0f9ded9490",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S58sh5.png",
     "asset_id": "6d80eb92-6d35-4bd5-bfec-92b0a707c02a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1083564>",
     "asset_id": "9a307034-865c-47d1-bf2c-6230188e67c7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L220B01.png",
     "asset_id": "b25ec118-d5a5-4461-bcc9-1623494e1fcd",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1083564>",
     "asset_id": "9a307034-865c-47d1-bf2c-6230188e67c7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "백산의 시선과 몸의 방향은 우측을 향하고 있으며, 우측 하단의 계단을 향해 걷고 있습니다.",
    "built_space": "카메라가 예배당 정면 제단을 향하고 있으나, 우측 전경에 알 수 없는 또 다른 단상 계단이 생성되어 공간이 왜곡되었습니다. 좌측에는 원본에 없는 여러 개의 철장이 중복 생성되었습니다.",
    "entities": "백산은 거대한 체형과 황금빛 하회탈을 착용했으나 레퍼런스의 모자가 누락되었습니다. 엎드린 군중이 존재하나, 철장 안의 인물은 찰리의 지정된 복장(밀짚모자, 우비 등)을 갖추지 않았습니다.",
    "hard_violations": [
     "[gemini-pro] 공간의 왜곡 및 없는 구조물 생성: 정면 제단이 보이면서 우측에 또 다른 단상 계단이 나타남",
     "[gemini-pro] 구조물 중복: 좌측에 철장(cage)이 여러 개로 복제됨",
     "[gpt-high] 백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다.",
     "[gpt-high] 참고 장소의 중앙 계단식 단상을 뒤에 남겨둔 채 오른쪽 전경에 별도의 계단식 단상을 추가하여 이동 목적지가 되는 고정 구조물을 새로 만들었다."
    ],
    "physics": "백산의 왼쪽 발이 바닥을 지탱하고 있으며, 오른쪽 발은 공중에 들린 mid-action 걷기 자세를 자연스럽게 유지하고 있습니다."
   },
   {
    "label": "B",
    "direction": "백산의 시선과 몸은 우측 단상을 향해 정확히 이동하고 있습니다.",
    "built_space": "카메라가 예배당을 측면(lateral)에서 바라보는 구도를 정확히 취하여, 좌측에 창문과 군중, 우측에 단상 계단이 배치되었습니다. 배경에 하나의 철장이 올바르게 위치합니다.",
    "entities": "백산은 레퍼런스와 일치하는 모자와 화려한 의복, 황금빛 하회탈을 착용하고 있습니다. 고개 숙인 군중이 배치되었고, 배경의 철장 안에는 밀짚모자와 화려한 우비, 장화를 신은 찰리가 정확히 묘사되었습니다.",
    "hard_violations": [
     "[gpt-high] 백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다."
    ],
    "physics": "백산의 왼쪽 발이 바닥에 닿아 체중을 지탱하고, 오른쪽 발이 공중에 들려 걷는 순간의 물리적 동작이 자연스럽게 표현되었습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 측면 앵글과 공간 구조를 정확히 구현했으며, 철장 속 찰리의 의상과 백산의 레퍼런스 복장 및 동작 지침을 충실히 반영한 훌륭한 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 측면(lateral) 구도가 아닌 정면 구도를 취하여 공간 구조를 왜곡했고, 참조 이미지에 없는 구조물과 중복된 철장이 생성되어 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백산의 시선과 몸의 방향은 우측을 향하고 있으며, 우측 하단의 계단을 향해 걷고 있습니다.",
        "built_space": "카메라가 예배당 정면 제단을 향하고 있으나, 우측 전경에 알 수 없는 또 다른 단상 계단이 생성되어 공간이 왜곡되었습니다. 좌측에는 원본에 없는 여러 개의 철장이 중복 생성되었습니다.",
        "entities": "백산은 거대한 체형과 황금빛 하회탈을 착용했으나 레퍼런스의 모자가 누락되었습니다. 엎드린 군중이 존재하나, 철장 안의 인물은 찰리의 지정된 복장(밀짚모자, 우비 등)을 갖추지 않았습니다.",
        "hard_violations": [
         "공간의 왜곡 및 없는 구조물 생성: 정면 제단이 보이면서 우측에 또 다른 단상 계단이 나타남",
         "구조물 중복: 좌측에 철장(cage)이 여러 개로 복제됨"
        ],
        "physics": "백산의 왼쪽 발이 바닥을 지탱하고 있으며, 오른쪽 발은 공중에 들린 mid-action 걷기 자세를 자연스럽게 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "백산의 시선과 몸은 우측 단상을 향해 정확히 이동하고 있습니다.",
        "built_space": "카메라가 예배당을 측면(lateral)에서 바라보는 구도를 정확히 취하여, 좌측에 창문과 군중, 우측에 단상 계단이 배치되었습니다. 배경에 하나의 철장이 올바르게 위치합니다.",
        "entities": "백산은 레퍼런스와 일치하는 모자와 화려한 의복, 황금빛 하회탈을 착용하고 있습니다. 고개 숙인 군중이 배치되었고, 배경의 철장 안에는 밀짚모자와 화려한 우비, 장화를 신은 찰리가 정확히 묘사되었습니다.",
        "hard_violations": [],
        "physics": "백산의 왼쪽 발이 바닥에 닿아 체중을 지탱하고, 오른쪽 발이 공중에 들려 걷는 순간의 물리적 동작이 자연스럽게 표현되었습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 측면 앵글과 공간 구조를 정확히 구현했으며, 철장 속 찰리의 의상과 백산의 레퍼런스 복장 및 동작 지침을 충실히 반영한 훌륭한 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 측면(lateral) 구도가 아닌 정면 구도를 취하여 공간 구조를 왜곡했고, 참조 이미지에 없는 구조물과 중복된 철장이 생성되어 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "백산의 시선과 몸의 방향은 우측을 향하고 있으며, 우측 하단의 계단을 향해 걷고 있습니다.",
        "built_space": "카메라가 예배당 정면 제단을 향하고 있으나, 우측 전경에 알 수 없는 또 다른 단상 계단이 생성되어 공간이 왜곡되었습니다. 좌측에는 원본에 없는 여러 개의 철장이 중복 생성되었습니다.",
        "entities": "백산은 거대한 체형과 황금빛 하회탈을 착용했으나 레퍼런스의 모자가 누락되었습니다. 엎드린 군중이 존재하나, 철장 안의 인물은 찰리의 지정된 복장(밀짚모자, 우비 등)을 갖추지 않았습니다.",
        "hard_violations": [
         "공간의 왜곡 및 없는 구조물 생성: 정면 제단이 보이면서 우측에 또 다른 단상 계단이 나타남",
         "구조물 중복: 좌측에 철장(cage)이 여러 개로 복제됨"
        ],
        "physics": "백산의 왼쪽 발이 바닥을 지탱하고 있으며, 오른쪽 발은 공중에 들린 mid-action 걷기 자세를 자연스럽게 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "백산의 시선과 몸은 우측 단상을 향해 정확히 이동하고 있습니다.",
        "built_space": "카메라가 예배당을 측면(lateral)에서 바라보는 구도를 정확히 취하여, 좌측에 창문과 군중, 우측에 단상 계단이 배치되었습니다. 배경에 하나의 철장이 올바르게 위치합니다.",
        "entities": "백산은 레퍼런스와 일치하는 모자와 화려한 의복, 황금빛 하회탈을 착용하고 있습니다. 고개 숙인 군중이 배치되었고, 배경의 철장 안에는 밀짚모자와 화려한 우비, 장화를 신은 찰리가 정확히 묘사되었습니다.",
        "hard_violations": [],
        "physics": "백산의 왼쪽 발이 바닥에 닿아 체중을 지탱하고, 오른쪽 발이 공중에 들려 걷는 순간의 물리적 동작이 자연스럽게 표현되었습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "오른쪽 단상을 향한 전신 보행과 측면 와이드 구도, 참고 의상은 더 충실하지만, 금지된 군중과 수감자를 추가했고 육중한 비만 체격이 부족하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "비만 체격과 지지발이 있는 동작은 맞지만, 금지된 추가 인물들과 별도 단상 구조가 있으며 정면 배경과 변경된 의상이 장소·구도 충실도를 떨어뜨린다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백산의 가면과 몸통, 공중에 든 앞발이 모두 화면 오른쪽 계단식 단상을 향한다. 가면은 비스듬한 측면으로 보이며 보행 목적지가 명확하다. 왼쪽 군중은 몸과 고개를 아래로 숙이고, 철창 안 인물은 통로 쪽을 바라본다. 무기나 조준 대상은 없다.",
        "built_space": "훼손된 회벽, 한쪽 벽을 따라 이어진 상층 철제 난간, 위아래 각각 세 구획의 스테인드글라스 창이 보인다. 오른쪽에는 하나의 단상으로 오르는 계단과 측면 가장자리가 있고, 그 왼쪽 뒤에는 철제 우리 하나가 있다. 백산은 계단 앞 평평한 통로에 서 있다. 왼쪽에는 여러 목재 장의자와 군중이 배치되어 있다. 교회 측면을 배경으로 삼은 구도와 재료는 참고 장소에 가깝지만, 우리는 참고 사진의 출입 철창과의 관계가 분명하지 않다.",
        "entities": "백산 외에 열 명 이상의 성인 군중과 우리 안 인물 한 명이 보인다. 이는 백산만 허용한 인물 제한과 충돌한다. 백산은 성인 남성으로 보이며 얼굴은 온전한 황금빛 가면으로 가려져 한국인 정체성이나 얼굴 일치는 직접 확인할 수 없다. 가면은 웃는 탈 형태이나 하회탈 특유의 조형은 다소 약하다. 관모, 금장식 자주색 상의, 어깨 장식, 녹색 바지, 검은 장화와 옆가방은 인물 참고에 가깝다. 다만 체격은 건장한 정도로, 요구한 거대한 비만형은 아니다. 우리 안 인물에게는 넓은 밀짚모자와 다채로운 외투, 노란 장화가 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다."
        ],
        "physics": "뒤쪽 장화의 밑창이 바닥에 닿아 체중을 지지하고, 앞쪽 다리는 오른쪽으로 뻗어 발이 공중에 떠 있다. 다음 발을 디딜 바닥도 보여 실제 보행 중간 자세로 성립한다. 군중은 발로 서서 허리를 숙이고, 우리 안 인물은 우리 바닥에 서 있다. 가방은 어깨끈으로 지지되며, 우리와 장의자는 바닥에 놓여 있다. 근거 없이 공중에 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "백산의 얼굴과 들어 올린 앞발은 오른쪽 전경의 낮은 계단을 향한다. 그러나 십자가와 의자가 있는 본래의 중심 단상은 인물 뒤쪽에 정면으로 보여, 이동 방향과 중심 단상의 방향이 갈라진다. 가면은 삼사분면보다 옆모습에 가깝다. 왼쪽 군중은 바닥을 향해 고개를 숙이고 있다.",
        "built_space": "양쪽 상층 난간, 손상된 회벽, 여러 스테인드글라스 창, 후벽 십자가 하나와 높은 등받이 의자 하나가 보인다. 후방에는 참고 장소와 유사한 넓은 중앙 계단식 단상이 있고, 오른쪽 전경에는 별개의 낮은 계단식 단상이 추가되어 있다. 백산은 이 전경 단상으로 접근한다. 왼쪽에는 큰 출입 철창과 우리 하나가 있으며 군중이 통로 바닥을 차지한다. 중앙 단상을 정면으로 드러내는 배경은 요청한 측면 접근 구도보다 정면 교회 구도에 가깝다.",
        "entities": "육중하고 배가 나온 성인 남성 백산은 요구한 비만 체격에 가깝고, 금속성의 온전한 웃는 가면이 얼굴을 가린다. 가려진 얼굴의 민족성이나 참고 얼굴 일치는 확인할 수 없다. 금장식 자주색 왕실 의상과 검은 장화는 있지만, 참고의 관모 대신 머리띠가 보이고 짧은 장식 상의와 녹색 바지는 긴 자주색 겉옷으로 바뀌었다. 백산 외에 열 명 이상의 군중과 우리 안 인물이 추가되어 있다. 우리 안 인물에게 넓은 모자는 보이지만 요구된 다채로운 우비와 장화는 뚜렷하게 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다.",
         "참고 장소의 중앙 계단식 단상을 뒤에 남겨둔 채 오른쪽 전경에 별도의 계단식 단상을 추가하여 이동 목적지가 되는 고정 구조물을 새로 만들었다."
        ],
        "physics": "뒤쪽 장화가 바닥에 확실히 닿고 그 위로 몸의 무게가 실려 있다. 앞다리는 무릎을 들어 오른쪽 계단 쪽으로 뻗으며, 디딜 계단도 보여 지지와 도착점이 있는 동작이다. 긴 옷자락은 몸에 연결되어 아래로 늘어지고 일부가 보행에 따라 벌어진다. 군중은 무릎과 손을 바닥에 대어 몸을 지지한다. 우리도 바닥에 놓여 있으며, 명백히 무지지 상태로 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "오른쪽 단상을 향한 전신 보행과 측면 와이드 구도, 참고 의상은 더 충실하지만, 금지된 군중과 수감자를 추가했고 육중한 비만 체격이 부족하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "비만 체격과 지지발이 있는 동작은 맞지만, 금지된 추가 인물들과 별도 단상 구조가 있으며 정면 배경과 변경된 의상이 장소·구도 충실도를 떨어뜨린다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "백산의 가면과 몸통, 공중에 든 앞발이 모두 화면 오른쪽 계단식 단상을 향한다. 가면은 비스듬한 측면으로 보이며 보행 목적지가 명확하다. 왼쪽 군중은 몸과 고개를 아래로 숙이고, 철창 안 인물은 통로 쪽을 바라본다. 무기나 조준 대상은 없다.",
        "built_space": "훼손된 회벽, 한쪽 벽을 따라 이어진 상층 철제 난간, 위아래 각각 세 구획의 스테인드글라스 창이 보인다. 오른쪽에는 하나의 단상으로 오르는 계단과 측면 가장자리가 있고, 그 왼쪽 뒤에는 철제 우리 하나가 있다. 백산은 계단 앞 평평한 통로에 서 있다. 왼쪽에는 여러 목재 장의자와 군중이 배치되어 있다. 교회 측면을 배경으로 삼은 구도와 재료는 참고 장소에 가깝지만, 우리는 참고 사진의 출입 철창과의 관계가 분명하지 않다.",
        "entities": "백산 외에 열 명 이상의 성인 군중과 우리 안 인물 한 명이 보인다. 이는 백산만 허용한 인물 제한과 충돌한다. 백산은 성인 남성으로 보이며 얼굴은 온전한 황금빛 가면으로 가려져 한국인 정체성이나 얼굴 일치는 직접 확인할 수 없다. 가면은 웃는 탈 형태이나 하회탈 특유의 조형은 다소 약하다. 관모, 금장식 자주색 상의, 어깨 장식, 녹색 바지, 검은 장화와 옆가방은 인물 참고에 가깝다. 다만 체격은 건장한 정도로, 요구한 거대한 비만형은 아니다. 우리 안 인물에게는 넓은 밀짚모자와 다채로운 외투, 노란 장화가 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다."
        ],
        "physics": "뒤쪽 장화의 밑창이 바닥에 닿아 체중을 지지하고, 앞쪽 다리는 오른쪽으로 뻗어 발이 공중에 떠 있다. 다음 발을 디딜 바닥도 보여 실제 보행 중간 자세로 성립한다. 군중은 발로 서서 허리를 숙이고, 우리 안 인물은 우리 바닥에 서 있다. 가방은 어깨끈으로 지지되며, 우리와 장의자는 바닥에 놓여 있다. 근거 없이 공중에 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "백산의 얼굴과 들어 올린 앞발은 오른쪽 전경의 낮은 계단을 향한다. 그러나 십자가와 의자가 있는 본래의 중심 단상은 인물 뒤쪽에 정면으로 보여, 이동 방향과 중심 단상의 방향이 갈라진다. 가면은 삼사분면보다 옆모습에 가깝다. 왼쪽 군중은 바닥을 향해 고개를 숙이고 있다.",
        "built_space": "양쪽 상층 난간, 손상된 회벽, 여러 스테인드글라스 창, 후벽 십자가 하나와 높은 등받이 의자 하나가 보인다. 후방에는 참고 장소와 유사한 넓은 중앙 계단식 단상이 있고, 오른쪽 전경에는 별개의 낮은 계단식 단상이 추가되어 있다. 백산은 이 전경 단상으로 접근한다. 왼쪽에는 큰 출입 철창과 우리 하나가 있으며 군중이 통로 바닥을 차지한다. 중앙 단상을 정면으로 드러내는 배경은 요청한 측면 접근 구도보다 정면 교회 구도에 가깝다.",
        "entities": "육중하고 배가 나온 성인 남성 백산은 요구한 비만 체격에 가깝고, 금속성의 온전한 웃는 가면이 얼굴을 가린다. 가려진 얼굴의 민족성이나 참고 얼굴 일치는 확인할 수 없다. 금장식 자주색 왕실 의상과 검은 장화는 있지만, 참고의 관모 대신 머리띠가 보이고 짧은 장식 상의와 녹색 바지는 긴 자주색 겉옷으로 바뀌었다. 백산 외에 열 명 이상의 군중과 우리 안 인물이 추가되어 있다. 우리 안 인물에게 넓은 모자는 보이지만 요구된 다채로운 우비와 장화는 뚜렷하게 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다.",
         "참고 장소의 중앙 계단식 단상을 뒤에 남겨둔 채 오른쪽 전경에 별도의 계단식 단상을 추가하여 이동 목적지가 되는 고정 구조물을 새로 만들었다."
        ],
        "physics": "뒤쪽 장화가 바닥에 확실히 닿고 그 위로 몸의 무게가 실려 있다. 앞다리는 무릎을 들어 오른쪽 계단 쪽으로 뻗으며, 디딜 계단도 보여 지지와 도착점이 있는 동작이다. 긴 옷자락은 몸에 연결되어 아래로 늘어지고 일부가 보행에 따라 벌어진다. 군중은 무릎과 손을 바닥에 대어 몸을 지지한다. 우리도 바닥에 놓여 있으며, 명백히 무지지 상태로 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.875,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.625,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 공간의 왜곡 및 없는 구조물 생성: 정면 제단이 보이면서 우측에 또 다른 단상 계단이 나타남",
     "[gemini-pro] 구조물 중복: 좌측에 철장(cage)이 여러 개로 복제됨",
     "[gpt-high] 백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다.",
     "[gpt-high] 참고 장소의 중앙 계단식 단상을 뒤에 남겨둔 채 오른쪽 전경에 별도의 계단식 단상을 추가하여 이동 목적지가 되는 고정 구조물을 새로 만들었다."
    ],
    "B": [
     "[gpt-high] 백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 625
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지정된 측면 앵글과 공간 구조를 정확히 구현했으며, 철장 속 찰리의 의상과 백산의 레퍼런스 복장 및 동작 지침을 충실히 반영한 훌륭한 결과물입니다.  ★위반: [gpt-high] 백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다."
   },
   {
    "label": "A",
    "score": 625,
    "verdict_ko": "지정된 측면(lateral) 구도가 아닌 정면 구도를 취하여 공간 구조를 왜곡했고, 참조 이미지에 없는 구조물과 중복된 철장이 생성되어 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 공간의 왜곡 및 없는 구조물 생성: 정면 제단이 보이면서 우측에 또 다른 단상 계단이 나타남 / [gemini-pro] 구조물 중복: 좌측에 철장(cage)이 여러 개로 복제됨 / [gpt-high] 백산만 등장할 수 있다는 명시적 제한을 어기고 다수의 군중과 철제 우리 안 인물을 추가했다. / [gpt-high] 참고 장소의 중앙 계단식 단상을 뒤에 남겨둔 채 오른쪽 전경에 별도의 계단식 단상을 추가하여 이동 목적지가 되는 고정 구조물을 새로 만들었다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L220B01.png",
    "asset_id": "b25ec118-d5a5-4461-bcc9-1623494e1fcd",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1083564>",
    "asset_id": "9a307034-865c-47d1-bf2c-6230188e67c7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-2337-75fb-92b2-f07df9ddf63b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5__bgfirst_bg.png",
   "bg_asset_id": "31941b00-404b-4a16-90fe-5d0f9ded9490",
   "bg_record_key": "S58sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S58sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:40:39.076317+00:00",
  "fingerprint": "8ac07b46425c1c7495a13f71ab56855328312c3066898cbe8554fa990564380d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S58sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S58sh5_sel.png",
  "source_sha256": "57246cdb940a785ad5a5cf361ad67fab2d7036c99c72717a36d82ac3bc6139f4",
  "file": "S58sh5_cine.png",
  "staged_sha256": "bb402a9f1eeb2fca8bf900c360b469c92313a6dc534fe6d6a4ca4073343e6d6b",
  "latency_ms": 9873
 },
 "S58sh15::signage": {
  "fp": "b1d466408b45ee77",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S58sh15": {
  "input_fingerprint": "9622adef4223c8be",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 백산의 손짓에 맞춰 거칠게 열린 철창문 안에서 모습을 드러낸 강철 케이지 속 찰리의 낡은 금속 전신.\n\nLOCATION (lock): At a barred side opening adjoining the village church's platform, where a steel cage appears in stained-glass daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Barred gate (Opening far enough to reveal 찰리's full body) — Seen obliquely from the prisoners' side, with its displaced section beside the opening; used as Revealing edge that no longer cuts through 찰리's silhouette; Steel cage (Containing 찰리 behind the open gate) — Front and side bars are visible in the oblique view; used as Establishes a second enclosure around the revealed figure; Church interior around the gate (Damaged) — The approach floor and architecture around the opening remain visible; used as Separates the cage from the foreground and preserves spatial scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the church's established stained-glass daylight with controlled metal highlights and no separate illumination invented for the cage.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The barred entrance is open, revealing Charlie's steel cage inside the damaged, stained-glass-lit church. Charlie retains the oversized straw hat, rubber boots and colorful raincoat worn at capture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 백산의 손짓에 맞춰 거칠게 열린 철창문 안에서 모습을 드러낸 강철 케이지 속 찰리의 낡은 금속 전신.\n\nLOCATION (lock): At a barred side opening adjoining the village church's platform, where a steel cage appears in stained-glass daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Barred gate (Opening far enough to reveal 찰리's full body) — Seen obliquely from the prisoners' side, with its displaced section beside the opening; used as Revealing edge that no longer cuts through 찰리's silhouette; Steel cage (Containing 찰리 behind the open gate) — Front and side bars are visible in the oblique view; used as Establishes a second enclosure around the revealed figure; Church interior around the gate (Damaged) — The approach floor and architecture around the opening remain visible; used as Separates the cage from the foreground and preserves spatial scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the church's established stained-glass daylight with controlled metal highlights and no separate illumination invented for the cage.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The barred entrance is open, revealing Charlie's steel cage inside the damaged, stained-glass-lit church. Charlie retains the oversized straw hat, rubber boots and colorful raincoat worn at capture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 백산의 손짓에 맞춰 거칠게 열린 철창문 안에서 모습을 드러낸 강철 케이지 속 찰리의 낡은 금속 전신.\n\nLOCATION (lock): At a barred side opening adjoining the village church's platform, where a steel cage appears in stained-glass daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Barred gate (Opening far enough to reveal 찰리's full body) — Seen obliquely from the prisoners' side, with its displaced section beside the opening; used as Revealing edge that no longer cuts through 찰리's silhouette; Steel cage (Containing 찰리 behind the open gate) — Front and side bars are visible in the oblique view; used as Establishes a second enclosure around the revealed figure; Church interior around the gate (Damaged) — The approach floor and architecture around the opening remain visible; used as Separates the cage from the foreground and preserves spatial scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the church's established stained-glass daylight with controlled metal highlights and no separate illumination invented for the cage.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The barred entrance is open, revealing Charlie's steel cage inside the damaged, stained-glass-lit church. Charlie retains the oversized straw hat, rubber boots and colorful raincoat worn at capture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 카메라 쪽을 향해 서 있음.",
    "built_space": "이전 샷 레퍼런스의 카메라 앵글과 배경 구조를 그대로 복사함. 떨어져 나온 철창이 화면 중앙에 뜬금없이 배치되어 공간적 맥락이 부자연스러움.",
    "entities": "찰리의 캐릭터 디자인은 대체로 일치하나, 프롬프트에서 등장하지 말아야 한다고 강하게 지시한 이전 샷의 '기도하는 사람들'이 좌측 배경에 그대로 나타남.",
    "hard_violations": [
     "[gemini-pro] invented people (등장이 금지된 이전 샷의 군중을 그대로 포함)",
     "[gemini-pro] physically impossible staging (떨어져 나온 철창이 아무런 물리적 지지 없이 부자연스러운 각도로 바닥에 서 있음)",
     "[gpt-high] 이전 장면의 군중 약 13명을 그대로 추가하여, 찰리 외 인물의 등장 금지와 이전 사진 인물의 재사용 금지를 위반했다."
    ],
    "physics": "중앙에 뜯겨져 나온 철창문이 바닥에 닿아 있긴 하나, 그 각도와 무게중심을 볼 때 스스로 지지되어 있을 수 없는 물리적으로 불가능한 형태로 세워져 있음."
   },
   {
    "label": "B",
    "direction": "찰리가 카메라 정면을 향해 시선을 두고 서 있음.",
    "built_space": "프롬프트가 요구한 대로 수감자 측에서 비스듬히 바라보는 뷰(oblique view)를 채택함. 전경에 열린 철창문이 있고, 그 너머 성당 공간 안에 찰리를 가둔 강철 케이지가 올바르게 배치되어 공간의 깊이감을 형성함.",
    "entities": "프롬프트에 지정된 유일한 인물인 찰리가 밀짚모자, 화려한 우비, 고무장화, 로봇 외형 등의 특징을 모두 갖춘 채 등장하며, 금지된 추가 인물은 없음.",
    "hard_violations": [],
    "physics": "찰리가 케이지 바닥에 두 발로 안정적으로 서 있으며, 열린 철창문과 케이지 모두 바닥에 자연스럽게 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "수감자 측에서 열린 철창을 통해 강철 케이지 안의 찰리를 바라보는 복잡한 공간 구도를 지시대로 정확히 구현했으며, 불필요한 인물을 완벽히 배제했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "이전 샷의 인물들을 절대 포함하지 말라는 명시적 지시를 무시하고 기도하는 사람들을 그대로 배치했으며, 레퍼런스의 카메라 앵글을 그대로 복사하는 치명적인 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리가 카메라 정면을 향해 시선을 두고 서 있음.",
        "built_space": "프롬프트가 요구한 대로 수감자 측에서 비스듬히 바라보는 뷰(oblique view)를 채택함. 전경에 열린 철창문이 있고, 그 너머 성당 공간 안에 찰리를 가둔 강철 케이지가 올바르게 배치되어 공간의 깊이감을 형성함.",
        "entities": "프롬프트에 지정된 유일한 인물인 찰리가 밀짚모자, 화려한 우비, 고무장화, 로봇 외형 등의 특징을 모두 갖춘 채 등장하며, 금지된 추가 인물은 없음.",
        "hard_violations": [],
        "physics": "찰리가 케이지 바닥에 두 발로 안정적으로 서 있으며, 열린 철창문과 케이지 모두 바닥에 자연스럽게 지지되어 있음."
       },
       {
        "label": "A",
        "direction": "찰리가 카메라 쪽을 향해 서 있음.",
        "built_space": "이전 샷 레퍼런스의 카메라 앵글과 배경 구조를 그대로 복사함. 떨어져 나온 철창이 화면 중앙에 뜬금없이 배치되어 공간적 맥락이 부자연스러움.",
        "entities": "찰리의 캐릭터 디자인은 대체로 일치하나, 프롬프트에서 등장하지 말아야 한다고 강하게 지시한 이전 샷의 '기도하는 사람들'이 좌측 배경에 그대로 나타남.",
        "hard_violations": [
         "invented people (등장이 금지된 이전 샷의 군중을 그대로 포함)",
         "physically impossible staging (떨어져 나온 철창이 아무런 물리적 지지 없이 부자연스러운 각도로 바닥에 서 있음)"
        ],
        "physics": "중앙에 뜯겨져 나온 철창문이 바닥에 닿아 있긴 하나, 그 각도와 무게중심을 볼 때 스스로 지지되어 있을 수 없는 물리적으로 불가능한 형태로 세워져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "수감자 측에서 열린 철창을 통해 강철 케이지 안의 찰리를 바라보는 복잡한 공간 구도를 지시대로 정확히 구현했으며, 불필요한 인물을 완벽히 배제했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "이전 샷의 인물들을 절대 포함하지 말라는 명시적 지시를 무시하고 기도하는 사람들을 그대로 배치했으며, 레퍼런스의 카메라 앵글을 그대로 복사하는 치명적인 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리가 카메라 정면을 향해 시선을 두고 서 있음.",
        "built_space": "프롬프트가 요구한 대로 수감자 측에서 비스듬히 바라보는 뷰(oblique view)를 채택함. 전경에 열린 철창문이 있고, 그 너머 성당 공간 안에 찰리를 가둔 강철 케이지가 올바르게 배치되어 공간의 깊이감을 형성함.",
        "entities": "프롬프트에 지정된 유일한 인물인 찰리가 밀짚모자, 화려한 우비, 고무장화, 로봇 외형 등의 특징을 모두 갖춘 채 등장하며, 금지된 추가 인물은 없음.",
        "hard_violations": [],
        "physics": "찰리가 케이지 바닥에 두 발로 안정적으로 서 있으며, 열린 철창문과 케이지 모두 바닥에 자연스럽게 지지되어 있음."
       },
       {
        "label": "A",
        "direction": "찰리가 카메라 쪽을 향해 서 있음.",
        "built_space": "이전 샷 레퍼런스의 카메라 앵글과 배경 구조를 그대로 복사함. 떨어져 나온 철창이 화면 중앙에 뜬금없이 배치되어 공간적 맥락이 부자연스러움.",
        "entities": "찰리의 캐릭터 디자인은 대체로 일치하나, 프롬프트에서 등장하지 말아야 한다고 강하게 지시한 이전 샷의 '기도하는 사람들'이 좌측 배경에 그대로 나타남.",
        "hard_violations": [
         "invented people (등장이 금지된 이전 샷의 군중을 그대로 포함)",
         "physically impossible staging (떨어져 나온 철창이 아무런 물리적 지지 없이 부자연스러운 각도로 바닥에 서 있음)"
        ],
        "physics": "중앙에 뜯겨져 나온 철창문이 바닥에 닿아 있긴 하나, 그 각도와 무게중심을 볼 때 스스로 지지되어 있을 수 없는 물리적으로 불가능한 형태로 세워져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "열린 바깥 철창문과 내부 케이지를 분리하고 찰리의 전신과 접근 바닥을 담았으며, 다만 요구된 사선 시점보다 정면성이 강하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 이전 장면의 군중을 그대로 등장시켰고, 이전 구도까지 답습하며 찰리를 가두는 두 번째 케이지의 전면 철창도 구현하지 못했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 몸통은 열린 입구를 지나 카메라 쪽을 향한다. 양손은 아래로 내려와 있으며 특정 대상을 가리키지 않는다. 바깥 철창문 한 짝은 화면 왼쪽으로 열려 찰리의 외곽선을 가리지 않는다. 백산이나 손짓은 화면에 없으며, 허용되지 않은 인물을 추가하지 않았다.",
        "built_space": "전경에는 중앙 통로 양옆의 고정 철창 두 구역과 왼쪽으로 열린 문짝 하나가 있고, 그 뒤에 독립된 케이지 하나가 있다. 케이지의 전면과 오른쪽 측면 철창, 상부와 바닥이 보여 이중 수용 구조가 성립한다. 찰리는 그 내부에 서 있다. 왼쪽의 단상 계단 한 벌, 상층 난간, 뚜렷한 상층 스테인드글라스 창 두 개와 하층 오른쪽 창 하나, 왼쪽 벽등 하나, 오른쪽 의자 일부가 보인다. 손상된 회벽과 접근 바닥은 유지되지만, 기준 사진과 달라진 카메라 방향 때문에 고정 시설의 정확한 연결은 확인하기 어렵다. 와이드 구도는 충족하나 케이지를 보는 각도는 거의 정면이다.",
        "entities": "등장 개체는 찰리 하나뿐이다. 흰 분절형 마스크 얼굴, 샌드 베이지 장갑판, 넓은 어깨와 긴 기계 팔이 캐릭터 참조와 부합한다. 큰 밀짚모자, 다색 무늬 우비, 짙은 바지와 녹색 고무장화도 유지된다. 명시된 비인간 기계 캐릭터이므로 인간의 나이·성별·민족성 조건을 적용할 대상은 아니다. 철창문과 별도 강철 케이지가 모두 있으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 두 장화는 케이지 바닥에 닿아 몸을 지지한다. 케이지는 하부 프레임과 바퀴로 교회 바닥에 놓여 있다. 열린 문짝은 왼쪽 문틀의 경첩으로 지지되고, 모자는 머리에 얹혀 있으며 우비는 어깨에서 내려온다. 떠 있는 신체나 물체는 없다. 문이 이미 열린 순간은 표현되지만 거칠게 열리는 운동 자체는 뚜렷하지 않다."
       },
       {
        "label": "B",
        "direction": "찰리는 몸과 얼굴을 거의 정면으로 카메라 쪽에 향하고 양팔을 아래로 내린다. 왼쪽 군중은 고개와 상체를 바닥 쪽으로 숙이고 있다. 철창문은 찰리 왼쪽으로 비켜나 전신을 드러내지만, 찰리 앞을 닫아 두는 별도 케이지 전면은 보이지 않는다.",
        "built_space": "오른쪽 단상 계단 한 벌, 상층 난간, 상층 스테인드글라스 창 세 개, 하층의 뚜렷한 창 두 개, 오른쪽 높은 벽등과 낮은 벽등, 왼쪽 나무 의자들이 기준 사진과 거의 같은 배치로 보인다. 촬영 구도도 이전 사진을 거의 그대로 따른다. 찰리는 오른쪽 개구부의 금속 바닥판 위에 서 있고, 왼쪽에 열린 철창문 한 짝이 있다. 뒤쪽 구조의 천장과 측면 철창은 보이지만 찰리를 둘러싼 두 번째 폐쇄 케이지의 전면 철창은 없다. 따라서 열린 입구 너머 별도 케이지를 사선으로 보여 주는 공간 관계가 부족하다.",
        "entities": "찰리의 밀짚모자, 흰 기계 얼굴, 베이지 장갑판, 육중한 긴 팔, 다색 우비와 녹색 장화는 참조에 대체로 맞는다. 그러나 왼쪽에는 이전 사진의 절하는 성인 군중 약 13명이 그대로 등장한다. 이들은 찰리만 허용한 등장 개체 조건과 이전 사진 인물의 재사용 금지를 명백히 어긴다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이전 장면의 군중 약 13명을 그대로 추가하여, 찰리 외 인물의 등장 금지와 이전 사진 인물의 재사용 금지를 위반했다."
        ],
        "physics": "찰리의 두 장화는 개구부 안 바닥판에 닿아 체중을 지지한다. 왼쪽으로 기울어진 문짝은 하단 모서리가 교회 바닥에 닿아 있으며, 연결부는 선명하지 않지만 공중에 떠 있지는 않다. 군중도 발을 바닥에 둔 채 허리를 굽히고 있다. 모자와 우비는 머리와 어깨에 지지되며, 명백히 지지 없는 물체나 불가능한 신체 자세는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "열린 바깥 철창문과 내부 케이지를 분리하고 찰리의 전신과 접근 바닥을 담았으며, 다만 요구된 사선 시점보다 정면성이 강하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 이전 장면의 군중을 그대로 등장시켰고, 이전 구도까지 답습하며 찰리를 가두는 두 번째 케이지의 전면 철창도 구현하지 못했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 몸통은 열린 입구를 지나 카메라 쪽을 향한다. 양손은 아래로 내려와 있으며 특정 대상을 가리키지 않는다. 바깥 철창문 한 짝은 화면 왼쪽으로 열려 찰리의 외곽선을 가리지 않는다. 백산이나 손짓은 화면에 없으며, 허용되지 않은 인물을 추가하지 않았다.",
        "built_space": "전경에는 중앙 통로 양옆의 고정 철창 두 구역과 왼쪽으로 열린 문짝 하나가 있고, 그 뒤에 독립된 케이지 하나가 있다. 케이지의 전면과 오른쪽 측면 철창, 상부와 바닥이 보여 이중 수용 구조가 성립한다. 찰리는 그 내부에 서 있다. 왼쪽의 단상 계단 한 벌, 상층 난간, 뚜렷한 상층 스테인드글라스 창 두 개와 하층 오른쪽 창 하나, 왼쪽 벽등 하나, 오른쪽 의자 일부가 보인다. 손상된 회벽과 접근 바닥은 유지되지만, 기준 사진과 달라진 카메라 방향 때문에 고정 시설의 정확한 연결은 확인하기 어렵다. 와이드 구도는 충족하나 케이지를 보는 각도는 거의 정면이다.",
        "entities": "등장 개체는 찰리 하나뿐이다. 흰 분절형 마스크 얼굴, 샌드 베이지 장갑판, 넓은 어깨와 긴 기계 팔이 캐릭터 참조와 부합한다. 큰 밀짚모자, 다색 무늬 우비, 짙은 바지와 녹색 고무장화도 유지된다. 명시된 비인간 기계 캐릭터이므로 인간의 나이·성별·민족성 조건을 적용할 대상은 아니다. 철창문과 별도 강철 케이지가 모두 있으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 두 장화는 케이지 바닥에 닿아 몸을 지지한다. 케이지는 하부 프레임과 바퀴로 교회 바닥에 놓여 있다. 열린 문짝은 왼쪽 문틀의 경첩으로 지지되고, 모자는 머리에 얹혀 있으며 우비는 어깨에서 내려온다. 떠 있는 신체나 물체는 없다. 문이 이미 열린 순간은 표현되지만 거칠게 열리는 운동 자체는 뚜렷하지 않다."
       },
       {
        "label": "A",
        "direction": "찰리는 몸과 얼굴을 거의 정면으로 카메라 쪽에 향하고 양팔을 아래로 내린다. 왼쪽 군중은 고개와 상체를 바닥 쪽으로 숙이고 있다. 철창문은 찰리 왼쪽으로 비켜나 전신을 드러내지만, 찰리 앞을 닫아 두는 별도 케이지 전면은 보이지 않는다.",
        "built_space": "오른쪽 단상 계단 한 벌, 상층 난간, 상층 스테인드글라스 창 세 개, 하층의 뚜렷한 창 두 개, 오른쪽 높은 벽등과 낮은 벽등, 왼쪽 나무 의자들이 기준 사진과 거의 같은 배치로 보인다. 촬영 구도도 이전 사진을 거의 그대로 따른다. 찰리는 오른쪽 개구부의 금속 바닥판 위에 서 있고, 왼쪽에 열린 철창문 한 짝이 있다. 뒤쪽 구조의 천장과 측면 철창은 보이지만 찰리를 둘러싼 두 번째 폐쇄 케이지의 전면 철창은 없다. 따라서 열린 입구 너머 별도 케이지를 사선으로 보여 주는 공간 관계가 부족하다.",
        "entities": "찰리의 밀짚모자, 흰 기계 얼굴, 베이지 장갑판, 육중한 긴 팔, 다색 우비와 녹색 장화는 참조에 대체로 맞는다. 그러나 왼쪽에는 이전 사진의 절하는 성인 군중 약 13명이 그대로 등장한다. 이들은 찰리만 허용한 등장 개체 조건과 이전 사진 인물의 재사용 금지를 명백히 어긴다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이전 장면의 군중 약 13명을 그대로 추가하여, 찰리 외 인물의 등장 금지와 이전 사진 인물의 재사용 금지를 위반했다."
        ],
        "physics": "찰리의 두 장화는 개구부 안 바닥판에 닿아 체중을 지지한다. 왼쪽으로 기울어진 문짝은 하단 모서리가 교회 바닥에 닿아 있으며, 연결부는 선명하지 않지만 공중에 떠 있지는 않다. 군중도 발을 바닥에 둔 채 허리를 굽히고 있다. 모자와 우비는 머리와 어깨에 지지되며, 명백히 지지 없는 물체나 불가능한 신체 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.472,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.222,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people (등장이 금지된 이전 샷의 군중을 그대로 포함)",
     "[gemini-pro] physically impossible staging (떨어져 나온 철창이 아무런 물리적 지지 없이 부자연스러운 각도로 바닥에 서 있음)",
     "[gpt-high] 이전 장면의 군중 약 13명을 그대로 추가하여, 찰리 외 인물의 등장 금지와 이전 사진 인물의 재사용 금지를 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 222
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "수감자 측에서 열린 철창을 통해 강철 케이지 안의 찰리를 바라보는 복잡한 공간 구도를 지시대로 정확히 구현했으며, 불필요한 인물을 완벽히 배제했습니다."
   },
   {
    "label": "A",
    "score": 222,
    "verdict_ko": "이전 샷의 인물들을 절대 포함하지 말라는 명시적 지시를 무시하고 기도하는 사람들을 그대로 배치했으며, 레퍼런스의 카메라 앵글을 그대로 복사하는 치명적인 오류를 범했습니다.  ★위반: [gemini-pro] invented people (등장이 금지된 이전 샷의 군중을 그대로 포함) / [gemini-pro] physically impossible staging (떨어져 나온 철창이 아무런 물리적 지지 없이 부자연스러운 각도로 바닥에 서 있음) / [gpt-high] 이전 장면의 군중 약 13명을 그대로 추가하여, 찰리 외 인물의 등장 금지와 이전 사진 인물의 재사용 금지를 위반했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5_sel.png",
    "asset_id": "ab7dea1a-0cd2-4a2a-aa9a-e72739e0610b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-268e-7e17-b382-5a63507b8bdb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S58sh5"
  }
 },
 "S58sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:41:49.719199+00:00",
  "fingerprint": "acf79d381aedc7480b5621910784a745bfbf549ff4f91bd68bee35503399623a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S58sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S58sh15_sel.png",
  "source_sha256": "150b3743823bbc5caec165ebd8fcf9384f1ba7f4b91119c469c5bcb74b6ceeca",
  "file": "S58sh15_cine.png",
  "staged_sha256": "262b8570d2380dbf1d848ed901a1ecdd9818a17360ce43cf204783b8f09319f3",
  "latency_ms": 10231
 },
 "S58sh25::signage": {
  "fp": "c011b4c6e8eb8b30",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S58sh25": {
  "input_fingerprint": "85a792d709b80f16",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 포박당해 몸이 뒤로 쏠린 mid-action 상태로, 찰리가 갇힌 철창문을 향해 절박하게 한 손을 뻗고 입을 크게 벌려 오열하듯 고정된 앰버의 상체.\n\nLOCATION (lock): On the church floor near the platform and barred cage entrance, under daylight filtering through stained glass. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: closed barred gate fragment in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Barred gate (Closed between 앰버 and 찰리) — Only a narrow oblique section is visible at the right edge; used as Marks the direction of the unreachable destination without obscuring her hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established stained-glass daylight and gentle facial tonal separation, without adding a new emotional spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the steel cage, barred doorway, damaged church surfaces, and stained-glass daylight. Exclude the earlier open-door state; the barred door has now closed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's steel cage door is being closed again, with Charlie restrained inside and still retaining his capture outfit. The damaged church remains lit through its stained-glass windows. 앰버: She is being bound again inside the church.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 포박당해 몸이 뒤로 쏠린 mid-action 상태로, 찰리가 갇힌 철창문을 향해 절박하게 한 손을 뻗고 입을 크게 벌려 오열하듯 고정된 앰버의 상체.\n\nLOCATION (lock): On the church floor near the platform and barred cage entrance, under daylight filtering through stained glass. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: closed barred gate fragment in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Barred gate (Closed between 앰버 and 찰리) — Only a narrow oblique section is visible at the right edge; used as Marks the direction of the unreachable destination without obscuring her hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established stained-glass daylight and gentle facial tonal separation, without adding a new emotional spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the steel cage, barred doorway, damaged church surfaces, and stained-glass daylight. Exclude the earlier open-door state; the barred door has now closed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's steel cage door is being closed again, with Charlie restrained inside and still retaining his capture outfit. The damaged church remains lit through its stained-glass windows. 앰버: She is being bound again inside the church.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 포박당해 몸이 뒤로 쏠린 mid-action 상태로, 찰리가 갇힌 철창문을 향해 절박하게 한 손을 뻗고 입을 크게 벌려 오열하듯 고정된 앰버의 상체.\n\nLOCATION (lock): On the church floor near the platform and barred cage entrance, under daylight filtering through stained glass. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: closed barred gate fragment in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Barred gate (Closed between 앰버 and 찰리) — Only a narrow oblique section is visible at the right edge; used as Marks the direction of the unreachable destination without obscuring her hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established stained-glass daylight and gentle facial tonal separation, without adding a new emotional spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the steel cage, barred doorway, damaged church surfaces, and stained-glass daylight. Exclude the earlier open-door state; the barred door has now closed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's steel cage door is being closed again, with Charlie restrained inside and still retaining his capture outfit. The damaged church remains lit through its stained-glass windows. 앰버: She is being bound again inside the church.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버의 시선과 펼친 손은 화면 오른쪽 철창 및 그 안의 인물을 향한다. 손이 철창에 가리지 않는 점과 도달하지 못한 상대를 향하는 방향은 맞는다. 다만 몸통은 뒤로 끌려가기보다 손을 따라 앞으로 기울어 있다.",
    "built_space": "오른쪽 절반가량을 큰 철제 우리와 닫힌 것으로 보이는 문이 차지한다. 왼쪽에는 큰 스테인드글라스 창 한 곳과 라디에이터 하나, 뒤쪽 위에는 작은 창 일부가 보인다. 낡은 벽과 낮빛은 참고 장소와 유사하지만, 철창은 요구한 오른쪽 가장자리의 좁은 배경 조각이 아니라 주요 공간으로 확대되어 있다. 앰버는 우리 밖에 있으며 허벅지까지 보여 클로즈업보다 넓다.",
    "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 남색 티셔츠, 갈색 작업용 멜빵바지, 머리 위 방독면과 공구 벨트 등 참고 인물의 특징을 대체로 유지한다. 한국계 백인 혼혈이라는 세부 정체성은 외관만으로 확정할 수 없다. 뒤로 둔 손목 부근에는 밧줄이 보인다. 그러나 우리 안에 이전 스틸의 모자·가면·갑옷·화려한 목도리 차림 인물까지 등장해 인물 제외 지시를 직접 어긴다. 읽을 수 있는 글자는 보이지 않는다.",
    "hard_violations": [
     "이번 화면에 허용되지 않은 이전 스틸의 인물을 철창 안에 추가하고 그 얼굴 가리개와 의상까지 그대로 재사용했다."
    ],
    "physics": "뻗은 팔은 어깨에 자연스럽게 연결되고, 공구는 벨트에 걸려 있다. 밧줄은 뒤쪽 손목 부근에 접촉하지만 끝의 고정점이나 후방으로 당기는 힘은 보이지 않는다. 하체가 화면 아래로 이어져 발의 접지는 판단할 수 없으며, 몸이 공중에 뜬 증거는 없다. 가능한 전방 기울기이지만 포박으로 몸이 뒤로 쏠리는 동작은 아니다."
   },
   {
    "label": "A",
    "direction": "앰버의 시선과 뻗은 손은 카메라 쪽을 겸한 화면 오른쪽 앞을 향한다. 오른쪽 철창의 배경 부분을 향해 손을 뻗는 것이 아니라 열린 출입구 바깥으로 손을 내미는 모습이다. 다른 손은 문 가장자리 부근에 닿아 있고 몸도 앞으로 기운다.",
    "built_space": "전경의 철창이 화면 양쪽과 위아래를 넓게 가로지르며, 중앙 오른쪽에는 사각 잠금판 하나가 보인다. 왼쪽 고정 철창과 잠금판이 달린 문 사이의 개구부에 앰버가 서서 몸을 내민다. 뒤에는 왼쪽 계단 한 줄, 상층 난간, 스테인드글라스 창들과 손상된 벽이 있어 장소의 재료와 일부 구조는 이어진다. 그러나 문이 닫힌 배경의 좁은 단편이어야 한다는 배치와 다르며, 프레임도 허벅지까지 포함한다.",
    "entities": "보이는 인물은 앰버 한 명뿐이다. 금발의 어린 여자아이, 둥근 얼굴, 남색 티셔츠, 갈색 멜빵바지, 머리 위 방독면과 공구 벨트는 참고 이미지와 대체로 맞는다. 혼혈 정체성 자체는 외관만으로 확인할 수 없다. 입을 크게 벌린 울음 표정은 있으나 밧줄이나 다른 포박 장치는 보이지 않는다. 찰리나 이전 스틸 인물은 등장하지 않으며 읽을 수 있는 글자도 없다.",
    "hard_violations": [
     "닫혀 있어야 할 철창문에 사람이 몸을 내밀 수 있는 열린 통로를 만들고 앰버를 그 개구부에 배치해, 문이 두 사람 사이를 막는 필수 공간 관계를 깨뜨렸다."
    ],
    "physics": "한 손은 앞으로 뻗고 다른 손은 문 가장자리에 접촉해 전방으로 기울어진 자세 자체는 물리적으로 가능하다. 공구는 벨트와 주머니에 지지되어 있다. 발은 프레임 밖이므로 접지를 확인할 수 없지만 부유하는 모습은 아니다. 포박이나 후방 견인이 보이지 않아 지정된 뒤로 쏠리는 동작을 뒷받침하지 못한다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "철창 쪽으로 뻗는 손과 오열은 맞지만, 제외하도록 명시한 이전 장면의 인물을 재등장시키고 상체 클로즈업을 허벅지까지 넓혔으며 몸도 뒤가 아닌 앞으로 쏠린다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "앰버만 등장시킨 점에서 상대적으로 낫지만, 닫힌 문 대신 열린 출입구로 몸과 손을 내밀고 있으며 포박·후방 쏠림·상체 클로즈업을 구현하지 못했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선과 펼친 손은 화면 오른쪽 철창 및 그 안의 인물을 향한다. 손이 철창에 가리지 않는 점과 도달하지 못한 상대를 향하는 방향은 맞는다. 다만 몸통은 뒤로 끌려가기보다 손을 따라 앞으로 기울어 있다.",
        "built_space": "오른쪽 절반가량을 큰 철제 우리와 닫힌 것으로 보이는 문이 차지한다. 왼쪽에는 큰 스테인드글라스 창 한 곳과 라디에이터 하나, 뒤쪽 위에는 작은 창 일부가 보인다. 낡은 벽과 낮빛은 참고 장소와 유사하지만, 철창은 요구한 오른쪽 가장자리의 좁은 배경 조각이 아니라 주요 공간으로 확대되어 있다. 앰버는 우리 밖에 있으며 허벅지까지 보여 클로즈업보다 넓다.",
        "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 남색 티셔츠, 갈색 작업용 멜빵바지, 머리 위 방독면과 공구 벨트 등 참고 인물의 특징을 대체로 유지한다. 한국계 백인 혼혈이라는 세부 정체성은 외관만으로 확정할 수 없다. 뒤로 둔 손목 부근에는 밧줄이 보인다. 그러나 우리 안에 이전 스틸의 모자·가면·갑옷·화려한 목도리 차림 인물까지 등장해 인물 제외 지시를 직접 어긴다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이번 화면에 허용되지 않은 이전 스틸의 인물을 철창 안에 추가하고 그 얼굴 가리개와 의상까지 그대로 재사용했다."
        ],
        "physics": "뻗은 팔은 어깨에 자연스럽게 연결되고, 공구는 벨트에 걸려 있다. 밧줄은 뒤쪽 손목 부근에 접촉하지만 끝의 고정점이나 후방으로 당기는 힘은 보이지 않는다. 하체가 화면 아래로 이어져 발의 접지는 판단할 수 없으며, 몸이 공중에 뜬 증거는 없다. 가능한 전방 기울기이지만 포박으로 몸이 뒤로 쏠리는 동작은 아니다."
       },
       {
        "label": "B",
        "direction": "앰버의 시선과 뻗은 손은 카메라 쪽을 겸한 화면 오른쪽 앞을 향한다. 오른쪽 철창의 배경 부분을 향해 손을 뻗는 것이 아니라 열린 출입구 바깥으로 손을 내미는 모습이다. 다른 손은 문 가장자리 부근에 닿아 있고 몸도 앞으로 기운다.",
        "built_space": "전경의 철창이 화면 양쪽과 위아래를 넓게 가로지르며, 중앙 오른쪽에는 사각 잠금판 하나가 보인다. 왼쪽 고정 철창과 잠금판이 달린 문 사이의 개구부에 앰버가 서서 몸을 내민다. 뒤에는 왼쪽 계단 한 줄, 상층 난간, 스테인드글라스 창들과 손상된 벽이 있어 장소의 재료와 일부 구조는 이어진다. 그러나 문이 닫힌 배경의 좁은 단편이어야 한다는 배치와 다르며, 프레임도 허벅지까지 포함한다.",
        "entities": "보이는 인물은 앰버 한 명뿐이다. 금발의 어린 여자아이, 둥근 얼굴, 남색 티셔츠, 갈색 멜빵바지, 머리 위 방독면과 공구 벨트는 참고 이미지와 대체로 맞는다. 혼혈 정체성 자체는 외관만으로 확인할 수 없다. 입을 크게 벌린 울음 표정은 있으나 밧줄이나 다른 포박 장치는 보이지 않는다. 찰리나 이전 스틸 인물은 등장하지 않으며 읽을 수 있는 글자도 없다.",
        "hard_violations": [
         "닫혀 있어야 할 철창문에 사람이 몸을 내밀 수 있는 열린 통로를 만들고 앰버를 그 개구부에 배치해, 문이 두 사람 사이를 막는 필수 공간 관계를 깨뜨렸다."
        ],
        "physics": "한 손은 앞으로 뻗고 다른 손은 문 가장자리에 접촉해 전방으로 기울어진 자세 자체는 물리적으로 가능하다. 공구는 벨트와 주머니에 지지되어 있다. 발은 프레임 밖이므로 접지를 확인할 수 없지만 부유하는 모습은 아니다. 포박이나 후방 견인이 보이지 않아 지정된 뒤로 쏠리는 동작을 뒷받침하지 못한다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "철창 쪽으로 뻗는 손과 오열은 맞지만, 제외하도록 명시한 이전 장면의 인물을 재등장시키고 상체 클로즈업을 허벅지까지 넓혔으며 몸도 뒤가 아닌 앞으로 쏠린다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "앰버만 등장시킨 점에서 상대적으로 낫지만, 닫힌 문 대신 열린 출입구로 몸과 손을 내밀고 있으며 포박·후방 쏠림·상체 클로즈업을 구현하지 못했다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 시선과 펼친 손은 화면 오른쪽 철창 및 그 안의 인물을 향한다. 손이 철창에 가리지 않는 점과 도달하지 못한 상대를 향하는 방향은 맞는다. 다만 몸통은 뒤로 끌려가기보다 손을 따라 앞으로 기울어 있다.",
        "built_space": "오른쪽 절반가량을 큰 철제 우리와 닫힌 것으로 보이는 문이 차지한다. 왼쪽에는 큰 스테인드글라스 창 한 곳과 라디에이터 하나, 뒤쪽 위에는 작은 창 일부가 보인다. 낡은 벽과 낮빛은 참고 장소와 유사하지만, 철창은 요구한 오른쪽 가장자리의 좁은 배경 조각이 아니라 주요 공간으로 확대되어 있다. 앰버는 우리 밖에 있으며 허벅지까지 보여 클로즈업보다 넓다.",
        "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 남색 티셔츠, 갈색 작업용 멜빵바지, 머리 위 방독면과 공구 벨트 등 참고 인물의 특징을 대체로 유지한다. 한국계 백인 혼혈이라는 세부 정체성은 외관만으로 확정할 수 없다. 뒤로 둔 손목 부근에는 밧줄이 보인다. 그러나 우리 안에 이전 스틸의 모자·가면·갑옷·화려한 목도리 차림 인물까지 등장해 인물 제외 지시를 직접 어긴다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이번 화면에 허용되지 않은 이전 스틸의 인물을 철창 안에 추가하고 그 얼굴 가리개와 의상까지 그대로 재사용했다."
        ],
        "physics": "뻗은 팔은 어깨에 자연스럽게 연결되고, 공구는 벨트에 걸려 있다. 밧줄은 뒤쪽 손목 부근에 접촉하지만 끝의 고정점이나 후방으로 당기는 힘은 보이지 않는다. 하체가 화면 아래로 이어져 발의 접지는 판단할 수 없으며, 몸이 공중에 뜬 증거는 없다. 가능한 전방 기울기이지만 포박으로 몸이 뒤로 쏠리는 동작은 아니다."
       },
       {
        "label": "A",
        "direction": "앰버의 시선과 뻗은 손은 카메라 쪽을 겸한 화면 오른쪽 앞을 향한다. 오른쪽 철창의 배경 부분을 향해 손을 뻗는 것이 아니라 열린 출입구 바깥으로 손을 내미는 모습이다. 다른 손은 문 가장자리 부근에 닿아 있고 몸도 앞으로 기운다.",
        "built_space": "전경의 철창이 화면 양쪽과 위아래를 넓게 가로지르며, 중앙 오른쪽에는 사각 잠금판 하나가 보인다. 왼쪽 고정 철창과 잠금판이 달린 문 사이의 개구부에 앰버가 서서 몸을 내민다. 뒤에는 왼쪽 계단 한 줄, 상층 난간, 스테인드글라스 창들과 손상된 벽이 있어 장소의 재료와 일부 구조는 이어진다. 그러나 문이 닫힌 배경의 좁은 단편이어야 한다는 배치와 다르며, 프레임도 허벅지까지 포함한다.",
        "entities": "보이는 인물은 앰버 한 명뿐이다. 금발의 어린 여자아이, 둥근 얼굴, 남색 티셔츠, 갈색 멜빵바지, 머리 위 방독면과 공구 벨트는 참고 이미지와 대체로 맞는다. 혼혈 정체성 자체는 외관만으로 확인할 수 없다. 입을 크게 벌린 울음 표정은 있으나 밧줄이나 다른 포박 장치는 보이지 않는다. 찰리나 이전 스틸 인물은 등장하지 않으며 읽을 수 있는 글자도 없다.",
        "hard_violations": [
         "닫혀 있어야 할 철창문에 사람이 몸을 내밀 수 있는 열린 통로를 만들고 앰버를 그 개구부에 배치해, 문이 두 사람 사이를 막는 필수 공간 관계를 깨뜨렸다."
        ],
        "physics": "한 손은 앞으로 뻗고 다른 손은 문 가장자리에 접촉해 전방으로 기울어진 자세 자체는 물리적으로 가능하다. 공구는 벨트와 주머니에 지지되어 있다. 발은 프레임 밖이므로 접지를 확인할 수 없지만 부유하는 모습은 아니다. 포박이나 후방 견인이 보이지 않아 지정된 뒤로 쏠리는 동작을 뒷받침하지 못한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 2,
   "A": 3
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2,
    "verdict_ko": "철창 쪽으로 뻗는 손과 오열은 맞지만, 제외하도록 명시한 이전 장면의 인물을 재등장시키고 상체 클로즈업을 허벅지까지 넓혔으며 몸도 뒤가 아닌 앞으로 쏠린다."
   },
   {
    "label": "A",
    "score": 3,
    "verdict_ko": "앰버만 등장시킨 점에서 상대적으로 낫지만, 닫힌 문 대신 열린 출입구로 몸과 손을 내밀고 있으며 포박·후방 쏠림·상체 클로즈업을 구현하지 못했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh15_sel.png",
    "asset_id": "ec141498-a070-4329-92ed-0a0afc64cb2e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-284a-718b-b14c-5777567288da",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S58sh15"
  }
 },
 "S58sh25::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:43:09.152602+00:00",
  "fingerprint": "b069270831ee10e3ceafafb232c2a06af929a9671a908b5cbe82dfeedb6ea8eb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S58sh25_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S58sh25_sel.png",
  "source_sha256": "0c0e0df742a16b81812bf07a84b99202e58773ae006784beb18a0485d7b7564d",
  "file": "S58sh25_cine.png",
  "staged_sha256": "8499051d86a3e516c08b34587b8cb81092ac50875dac6acb79ebe03316cdb9d8",
  "latency_ms": 9951
 },
 "S59sh10::signage": {
  "fp": "4ce5d70473e7a9fd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::59af04c98c5db919": {
  "subjects": [],
  "subject_text": "익산 마을 지하감옥 감방\n굵은 쇠창살로 막힌 어두운 지하 감방. 바닥에는 건초더미와 고인 물이 있으며, 작은 창으로 제한적인 빛이 들어온다.",
  "identity": "canonical",
  "scope_id": "L222",
  "scope_role": "location_interior",
  "scope_sha": "e6618845247b342f"
 },
 "S59sh10::bgfirst_bg": {
  "input_fingerprint": "42c97c67a1138a9c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10__bgfirst_bg.png",
  "asset_id": "9ab1887b-2da7-40ca-a980-d56e92af21ea",
  "input_asset_ids": [
   "0b9b8bdc-0f3c-40f7-a516-90ff65fd60d8",
   "47d011f7-2e24-447d-9f02-310f84be3871"
  ]
 },
 "S59sh10": {
  "input_fingerprint": "7096400995ba3972",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cell remains locked, with hay on the floor and pooled water inside; a handful of hay has already been wetted. A potato is being offered inside the cell. 현우: He is confined at the bars after the unsuccessful attempt to force them, retaining his recent bindings and accumulated injuries. The contact card remains hidden in his shoe. 수빈: Her face is scratched and dirty, and pre-existing radiation damage to her torso remains covered by her T-shirt. She holds out a potato.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수빈 right now, so 수빈's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수빈: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cell remains locked, with hay on the floor and pooled water inside; a handful of hay has already been wetted. A potato is being offered inside the cell. 현우: He is confined at the bars after the unsuccessful attempt to force them, retaining his recent bindings and accumulated injuries. The contact card remains hidden in his shoe. 수빈: Her face is scratched and dirty, and pre-existing radiation damage to her torso remains covered by her T-shirt. She holds out a potato.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수빈 right now, so 수빈's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수빈: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cell remains locked, with hay on the floor and pooled water inside; a handful of hay has already been wetted. A potato is being offered inside the cell. 현우: He is confined at the bars after the unsuccessful attempt to force them, retaining his recent bindings and accumulated injuries. The contact card remains hidden in his shoe. 수빈: Her face is scratched and dirty, and pre-existing radiation damage to her torso remains covered by her T-shirt. She holds out a potato.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수빈 right now, so 수빈's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수빈: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10__bgfirst_bg.png",
     "asset_id": "9ab1887b-2da7-40ca-a980-d56e92af21ea",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S59sh10.png",
     "asset_id": "0b9b8bdc-0f3c-40f7-a516-90ff65fd60d8",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:779285>",
     "asset_id": "6e320331-1fae-44e2-b73c-832ea9dfbc28",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B02.png",
     "asset_id": "47d011f7-2e24-447d-9f02-310f84be3871",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:779285>",
     "asset_id": "6e320331-1fae-44e2-b73c-832ea9dfbc28",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "수빈은 현우를 향해 감자를 쥔 손을 뻗고 있으며, 현우는 내밀어진 감자를 내려다보고 있습니다.",
    "built_space": "지하 감옥 내부로 바닥에 짚과 물웅덩이가 있습니다. 철창은 배경과 우측에 위치하며, 두 인물 모두 철창의 동일한 내부 공간에 자리 잡고 있습니다.",
    "entities": "수빈(얼굴 상처, 어두운 색 옷, 단발머리)과 현우(얼굴 상처, 팔의 붕대, 헝클어진 머리), 그리고 둥근 감자 모두 프롬프트 및 레퍼런스와 일치합니다.",
    "hard_violations": [],
    "physics": "수빈은 바닥에 쪼그려 앉고 현우는 주저앉아 체중을 지지하고 있으며, 수빈의 손바닥이 감자를 안정적으로 받치고 있습니다."
   },
   {
    "label": "B",
    "direction": "수빈의 팔이 철창 너머의 현우를 향해 감자를 내밀고 있으며, 현우는 감자를 응시하고 있습니다.",
    "built_space": "지하 감옥 환경이며 바닥에 짚이 보이나, 철창이 카메라/수빈과 현우 사이를 가로막고 있어 두 인물의 공간이 분리되어 있습니다.",
    "entities": "현우(얼굴 상처, 앳된 얼굴)와 수빈의 팔(어두운 색 소매), 둥근 감자가 보입니다.",
    "hard_violations": [
     "[gemini-pro] 두 인물이 철창을 사이에 두고 분리되어 있어 '둘 다 같은 내부 공간에 머문다(both people remain on the same interior side)'는 명시적 위치 지시를 위반함",
     "[gpt-high] 현우와 수빈을 철창의 서로 반대쪽에 배치하여, 둘 다 같은 감방 내부에 있어야 한다는 명시적 공간 조건을 위반한다."
    ],
    "physics": "손가락이 감자를 쥐고 허공에 떠 있으며, 현우는 철창 뒤에서 자세를 유지하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "클로즈업 지시를 어기고 넓은 화각으로 연출되었으나, 두 인물이 철창 안의 같은 공간에 위치해야 한다는 공간 지시와 인물들의 디테일을 성실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화각은 클로즈업에 더 가깝지만, 두 인물이 같은 내부 공간에 있어야 한다는 지시를 어기고 철창을 사이에 두고 분리된 치명적인 공간 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수빈은 현우를 향해 감자를 쥔 손을 뻗고 있으며, 현우는 내밀어진 감자를 내려다보고 있습니다.",
        "built_space": "지하 감옥 내부로 바닥에 짚과 물웅덩이가 있습니다. 철창은 배경과 우측에 위치하며, 두 인물 모두 철창의 동일한 내부 공간에 자리 잡고 있습니다.",
        "entities": "수빈(얼굴 상처, 어두운 색 옷, 단발머리)과 현우(얼굴 상처, 팔의 붕대, 헝클어진 머리), 그리고 둥근 감자 모두 프롬프트 및 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "수빈은 바닥에 쪼그려 앉고 현우는 주저앉아 체중을 지지하고 있으며, 수빈의 손바닥이 감자를 안정적으로 받치고 있습니다."
       },
       {
        "label": "B",
        "direction": "수빈의 팔이 철창 너머의 현우를 향해 감자를 내밀고 있으며, 현우는 감자를 응시하고 있습니다.",
        "built_space": "지하 감옥 환경이며 바닥에 짚이 보이나, 철창이 카메라/수빈과 현우 사이를 가로막고 있어 두 인물의 공간이 분리되어 있습니다.",
        "entities": "현우(얼굴 상처, 앳된 얼굴)와 수빈의 팔(어두운 색 소매), 둥근 감자가 보입니다.",
        "hard_violations": [
         "두 인물이 철창을 사이에 두고 분리되어 있어 '둘 다 같은 내부 공간에 머문다(both people remain on the same interior side)'는 명시적 위치 지시를 위반함"
        ],
        "physics": "손가락이 감자를 쥐고 허공에 떠 있으며, 현우는 철창 뒤에서 자세를 유지하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "클로즈업 지시를 어기고 넓은 화각으로 연출되었으나, 두 인물이 철창 안의 같은 공간에 위치해야 한다는 공간 지시와 인물들의 디테일을 성실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화각은 클로즈업에 더 가깝지만, 두 인물이 같은 내부 공간에 있어야 한다는 지시를 어기고 철창을 사이에 두고 분리된 치명적인 공간 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수빈은 현우를 향해 감자를 쥔 손을 뻗고 있으며, 현우는 내밀어진 감자를 내려다보고 있습니다.",
        "built_space": "지하 감옥 내부로 바닥에 짚과 물웅덩이가 있습니다. 철창은 배경과 우측에 위치하며, 두 인물 모두 철창의 동일한 내부 공간에 자리 잡고 있습니다.",
        "entities": "수빈(얼굴 상처, 어두운 색 옷, 단발머리)과 현우(얼굴 상처, 팔의 붕대, 헝클어진 머리), 그리고 둥근 감자 모두 프롬프트 및 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "수빈은 바닥에 쪼그려 앉고 현우는 주저앉아 체중을 지지하고 있으며, 수빈의 손바닥이 감자를 안정적으로 받치고 있습니다."
       },
       {
        "label": "B",
        "direction": "수빈의 팔이 철창 너머의 현우를 향해 감자를 내밀고 있으며, 현우는 감자를 응시하고 있습니다.",
        "built_space": "지하 감옥 환경이며 바닥에 짚이 보이나, 철창이 카메라/수빈과 현우 사이를 가로막고 있어 두 인물의 공간이 분리되어 있습니다.",
        "entities": "현우(얼굴 상처, 앳된 얼굴)와 수빈의 팔(어두운 색 소매), 둥근 감자가 보입니다.",
        "hard_violations": [
         "두 인물이 철창을 사이에 두고 분리되어 있어 '둘 다 같은 내부 공간에 머문다(both people remain on the same interior side)'는 명시적 위치 지시를 위반함"
        ],
        "physics": "손가락이 감자를 쥐고 허공에 떠 있으며, 현우는 철창 뒤에서 자세를 유지하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "감자를 내민 손의 클로즈업은 잘 맞지만, 철창이 두 사람 사이를 갈라 같은 감방 내부에 있어야 한다는 핵심 배치를 위반한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 사람이 철창 안쪽에서 아직 건네받지 않은 감자를 사이에 둔 관계는 정확하지만, 손의 클로즈업 대신 두 사람의 상반신과 감방을 넓게 보여준다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수빈의 팔은 화면 오른쪽에서 왼쪽 현우의 얼굴 앞으로 뻗어 있고, 손에 든 감자가 현우를 향한다. 현우는 오른쪽의 감자와 손 쪽을 바라본다. 수빈의 얼굴은 화면 밖이며 현우가 감자를 받는 동작은 없다.",
        "built_space": "화면 왼쪽 전경의 굵은 문틀과 잠금장치 일부, 중앙에서 뒤로 이어지는 약 십여 개의 수직 철봉과 가로 보강대가 보인다. 오른쪽에는 거친 콘크리트 벽과 바닥의 짚이 있다. 그러나 철봉과 가로대가 현우의 얼굴·몸 앞을 가리고 수빈의 팔은 반대편 공간에 있어, 두 사람이 철창을 사이에 둔 배치로 읽힌다. 철창은 파손되지 않았다.",
        "entities": "감자는 하나이며 둥글고 흙 묻은 실제 감자의 질감이다. 현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 상의가 참고 이미지에 가깝고 얼굴에 상처가 있다. 수빈은 손·팔·몸통 일부만 보이며 어두운 긴소매 상의와 올리브색 바지가 참고 의상에 가깝다. 손의 크기와 체형은 젊은 여성의 것으로 무리 없으나 얼굴 정체성은 확인할 수 없다. 결박과 신발 속 카드는 화면 밖이다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [
         "현우와 수빈을 철창의 서로 반대쪽에 배치하여, 둘 다 같은 감방 내부에 있어야 한다는 명시적 공간 조건을 위반한다."
        ],
        "physics": "감자는 수빈의 손바닥과 아래쪽 손가락에 받쳐지고 엄지가 옆을 고정한다. 손목과 팔은 오른쪽 몸통까지 자연스럽게 연결되어 있어 허공에 멈춘 제시 동작이 가능하다. 현우의 하체 지지는 화면 밖이므로 판단할 수 없지만, 상체가 떠 있다는 증거는 없다."
       },
       {
        "label": "B",
        "direction": "수빈은 왼쪽에서 오른쪽 현우를 향해 손바닥 위의 감자를 내민다. 수빈의 시선은 현우의 얼굴을 향하고 현우는 아래의 감자를 바라본다. 현우의 손이 감자에 닿지 않아 아직 받지 않은 순간이 분명하다.",
        "built_space": "오른쪽에 하나의 연속된 철창 면과 잠금판·빗장 한 세트가 보이며, 수직 철봉은 약 스무 개가 원근을 따라 이어진다. 현우는 그 철창 바로 안쪽, 수빈은 같은 실내의 왼쪽에 있다. 거친 콘크리트 벽과 천장, 뒤쪽 짚더미 한 곳, 바닥의 여러 물웅덩이가 보인다. 철창은 온전하고 두 사람 사이를 가로막지 않는다. 뒤쪽 벽의 가는 채광 개구부는 제공된 장소 사진에서 확인되지 않는 차이다. 물 표면의 밝은 반사에는 명백한 광학적 모순이 없다.",
        "entities": "젊은 동아시아계 여성과 남성 각 한 명만 보인다. 수빈의 검은 단발과 날렵한 머리 끝, 긁히고 더러운 얼굴은 조건에 부합하지만 참고의 긴소매 대신 반소매를 입었다. 현우는 앳된 얼굴과 헝클어진 검은 머리, 누적된 얼굴·팔 상처와 팔의 결박을 보인다. 현우의 상의는 참고의 남색보다 회색에 가깝다. 감자는 하나이고 실제 감자처럼 보이지만 둥근 형태보다는 타원형이다. 짚과 고인 물은 보이나 이미 적신 짚 한 줌은 구별되지 않는다. 신발 속 카드는 화면 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "감자는 수빈의 오목한 손바닥과 굽힌 손가락에 안정적으로 받쳐져 있다. 팔꿈치에서 손목으로 이어지는 자세도 내민 채 멈추는 동작으로 가능하다. 두 사람은 바닥 가까이 앉아 무릎을 굽힌 자세이며 하체 일부는 프레임 밖이다. 보이는 범위에서 부유하거나 지지 없이 매달린 신체·물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "감자를 내민 손의 클로즈업은 잘 맞지만, 철창이 두 사람 사이를 갈라 같은 감방 내부에 있어야 한다는 핵심 배치를 위반한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "두 사람이 철창 안쪽에서 아직 건네받지 않은 감자를 사이에 둔 관계는 정확하지만, 손의 클로즈업 대신 두 사람의 상반신과 감방을 넓게 보여준다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "수빈의 팔은 화면 오른쪽에서 왼쪽 현우의 얼굴 앞으로 뻗어 있고, 손에 든 감자가 현우를 향한다. 현우는 오른쪽의 감자와 손 쪽을 바라본다. 수빈의 얼굴은 화면 밖이며 현우가 감자를 받는 동작은 없다.",
        "built_space": "화면 왼쪽 전경의 굵은 문틀과 잠금장치 일부, 중앙에서 뒤로 이어지는 약 십여 개의 수직 철봉과 가로 보강대가 보인다. 오른쪽에는 거친 콘크리트 벽과 바닥의 짚이 있다. 그러나 철봉과 가로대가 현우의 얼굴·몸 앞을 가리고 수빈의 팔은 반대편 공간에 있어, 두 사람이 철창을 사이에 둔 배치로 읽힌다. 철창은 파손되지 않았다.",
        "entities": "감자는 하나이며 둥글고 흙 묻은 실제 감자의 질감이다. 현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 상의가 참고 이미지에 가깝고 얼굴에 상처가 있다. 수빈은 손·팔·몸통 일부만 보이며 어두운 긴소매 상의와 올리브색 바지가 참고 의상에 가깝다. 손의 크기와 체형은 젊은 여성의 것으로 무리 없으나 얼굴 정체성은 확인할 수 없다. 결박과 신발 속 카드는 화면 밖이다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [
         "현우와 수빈을 철창의 서로 반대쪽에 배치하여, 둘 다 같은 감방 내부에 있어야 한다는 명시적 공간 조건을 위반한다."
        ],
        "physics": "감자는 수빈의 손바닥과 아래쪽 손가락에 받쳐지고 엄지가 옆을 고정한다. 손목과 팔은 오른쪽 몸통까지 자연스럽게 연결되어 있어 허공에 멈춘 제시 동작이 가능하다. 현우의 하체 지지는 화면 밖이므로 판단할 수 없지만, 상체가 떠 있다는 증거는 없다."
       },
       {
        "label": "A",
        "direction": "수빈은 왼쪽에서 오른쪽 현우를 향해 손바닥 위의 감자를 내민다. 수빈의 시선은 현우의 얼굴을 향하고 현우는 아래의 감자를 바라본다. 현우의 손이 감자에 닿지 않아 아직 받지 않은 순간이 분명하다.",
        "built_space": "오른쪽에 하나의 연속된 철창 면과 잠금판·빗장 한 세트가 보이며, 수직 철봉은 약 스무 개가 원근을 따라 이어진다. 현우는 그 철창 바로 안쪽, 수빈은 같은 실내의 왼쪽에 있다. 거친 콘크리트 벽과 천장, 뒤쪽 짚더미 한 곳, 바닥의 여러 물웅덩이가 보인다. 철창은 온전하고 두 사람 사이를 가로막지 않는다. 뒤쪽 벽의 가는 채광 개구부는 제공된 장소 사진에서 확인되지 않는 차이다. 물 표면의 밝은 반사에는 명백한 광학적 모순이 없다.",
        "entities": "젊은 동아시아계 여성과 남성 각 한 명만 보인다. 수빈의 검은 단발과 날렵한 머리 끝, 긁히고 더러운 얼굴은 조건에 부합하지만 참고의 긴소매 대신 반소매를 입었다. 현우는 앳된 얼굴과 헝클어진 검은 머리, 누적된 얼굴·팔 상처와 팔의 결박을 보인다. 현우의 상의는 참고의 남색보다 회색에 가깝다. 감자는 하나이고 실제 감자처럼 보이지만 둥근 형태보다는 타원형이다. 짚과 고인 물은 보이나 이미 적신 짚 한 줌은 구별되지 않는다. 신발 속 카드는 화면 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "감자는 수빈의 오목한 손바닥과 굽힌 손가락에 안정적으로 받쳐져 있다. 팔꿈치에서 손목으로 이어지는 자세도 내민 채 멈추는 동작으로 가능하다. 두 사람은 바닥 가까이 앉아 무릎을 굽힌 자세이며 하체 일부는 프레임 밖이다. 보이는 범위에서 부유하거나 지지 없이 매달린 신체·물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.786
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.536
   },
   "violations": {
    "B": [
     "[gemini-pro] 두 인물이 철창을 사이에 두고 분리되어 있어 '둘 다 같은 내부 공간에 머문다(both people remain on the same interior side)'는 명시적 위치 지시를 위반함",
     "[gpt-high] 현우와 수빈을 철창의 서로 반대쪽에 배치하여, 둘 다 같은 감방 내부에 있어야 한다는 명시적 공간 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 536
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 지시를 어기고 넓은 화각으로 연출되었으나, 두 인물이 철창 안의 같은 공간에 위치해야 한다는 공간 지시와 인물들의 디테일을 성실히 구현했습니다."
   },
   {
    "label": "B",
    "score": 536,
    "verdict_ko": "화각은 클로즈업에 더 가깝지만, 두 인물이 같은 내부 공간에 있어야 한다는 지시를 어기고 철창을 사이에 두고 분리된 치명적인 공간 오류를 범했습니다.  ★위반: [gemini-pro] 두 인물이 철창을 사이에 두고 분리되어 있어 '둘 다 같은 내부 공간에 머문다(both people remain on the same interior side)'는 명시적 위치 지시를 위반함 / [gpt-high] 현우와 수빈을 철창의 서로 반대쪽에 배치하여, 둘 다 같은 감방 내부에 있어야 한다는 명시적 공간 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B02.png",
    "asset_id": "47d011f7-2e24-447d-9f02-310f84be3871",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:779285>",
    "asset_id": "6e320331-1fae-44e2-b73c-832ea9dfbc28",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-29ea-76d2-b489-01c8ddc185cc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10__bgfirst_bg.png",
   "bg_asset_id": "9ab1887b-2da7-40ca-a980-d56e92af21ea",
   "bg_record_key": "S59sh10::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C01"
  ]
 },
 "S59sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:44:20.654831+00:00",
  "fingerprint": "8891f199903e89fa0a44e037e30d81d5fc5d87d7adeb2db40214f0e6fab173db",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S59sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S59sh10_sel.png",
  "source_sha256": "7ca15f37abe83c5e809b086e06224ebfd45936f92b33fb0fc59815bfcbda7740",
  "file": "S59sh10_cine.png",
  "staged_sha256": "141353dc42b6f5c29f417de253ccc95f092f85f42a6575c46d9ec2c11c825108",
  "latency_ms": 9979
 },
 "S59sh18::signage": {
  "fp": "62bd332904ce1413",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S59sh18": {
  "input_fingerprint": "59c0eb7e33a049b9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수빈의 배 위에 선명하게 드러난 검붉고 흉측한 방사능 피폭 흉터 클로즈업.\n\nLOCATION (lock): Inside the sparsely furnished underground prison cell, in subdued dungeon light near the other prisoners. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 수빈의 티셔츠 (Lifted to expose the abdominal radiation injuries) — The raised hem crosses the upper edge of the obliquely viewed torso; used as Retains the revealing action within the close-up rather than isolating the injury as an abstract texture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the cell, keeping the dark-red scars legible without sensational contrast or added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground cell remains locked, with hay on the floor and pooled water inside. 수빈: Her face is scarred and dirty, and her raised T-shirt exposes radiation-damaged skin across her abdomen.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수빈의 배 위에 선명하게 드러난 검붉고 흉측한 방사능 피폭 흉터 클로즈업.\n\nLOCATION (lock): Inside the sparsely furnished underground prison cell, in subdued dungeon light near the other prisoners. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 수빈의 티셔츠 (Lifted to expose the abdominal radiation injuries) — The raised hem crosses the upper edge of the obliquely viewed torso; used as Retains the revealing action within the close-up rather than isolating the injury as an abstract texture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the cell, keeping the dark-red scars legible without sensational contrast or added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground cell remains locked, with hay on the floor and pooled water inside. 수빈: Her face is scarred and dirty, and her raised T-shirt exposes radiation-damaged skin across her abdomen.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수빈의 배 위에 선명하게 드러난 검붉고 흉측한 방사능 피폭 흉터 클로즈업.\n\nLOCATION (lock): Inside the sparsely furnished underground prison cell, in subdued dungeon light near the other prisoners. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 수빈의 티셔츠 (Lifted to expose the abdominal radiation injuries) — The raised hem crosses the upper edge of the obliquely viewed torso; used as Retains the revealing action within the close-up rather than isolating the injury as an abstract texture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the cell, keeping the dark-red scars legible without sensational contrast or added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground cell remains locked, with hay on the floor and pooled water inside. 수빈: Her face is scarred and dirty, and her raised T-shirt exposes radiation-damaged skin across her abdomen.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물의 두 손이 짙은 회색 티셔츠 밑단을 위로 당겨 올리는 동작에 집중되어 있음.",
    "built_space": "어두운 지하 감옥 내부. 우측에 철창 구조물이 보이고, 바닥에는 지푸라기와 물이 고인 웅덩이가 있어 레퍼런스와 공간적 일치도가 높음.",
    "entities": "수빈의 몸통(어두운 회색 반팔 티셔츠, 올리브색 바지 및 전술 벨트), 복부 중앙에 매우 흉측하고 입체적인 검붉은 방사능 피폭 흉터가 명확히 보임.",
    "hard_violations": [],
    "physics": "두 손이 티셔츠 밑단을 확실하게 움켜쥐고 들어 올리는 힘과 무게감이 자연스럽게 연출됨."
   },
   {
    "label": "B",
    "direction": "한 손은 티셔츠를 위로 당기고, 다른 한 손은 복부 아래쪽 바지 벨트 근처에 자연스럽게 놓여 있음.",
    "built_space": "감옥의 콘크리트 벽면과 바닥의 지푸라기 일부가 보이나, 철창 등 공간의 뚜렷한 특징이 화면에 거의 나타나지 않음.",
    "entities": "수빈의 몸통(어두운 회색 티셔츠), 복부에 흉터가 존재하나 화상이나 방사능 피폭보다는 긁힌 흉터 모양에 가까움.",
    "hard_violations": [],
    "physics": "위쪽 손이 티셔츠를 잡아당기고 아래쪽 손이 복부에 안착하여 지지하고 있으나, 위쪽 손가락의 관절 묘사가 다소 뭉개져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트의 핵심인 '검붉고 흉측한 방사능 피폭 흉터'를 매우 사실적이고 강렬하게 묘사했으며, 이전 샷의 감옥 배경(철창, 물웅덩이 등)을 완벽하게 유지하여 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "복부의 흉터가 방사능 피폭보다는 일반적인 자상이나 긁힌 상처에 가깝게 표현되었고, 감옥 배경의 디테일이 A에 비해 상대적으로 부족합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물의 두 손이 짙은 회색 티셔츠 밑단을 위로 당겨 올리는 동작에 집중되어 있음.",
        "built_space": "어두운 지하 감옥 내부. 우측에 철창 구조물이 보이고, 바닥에는 지푸라기와 물이 고인 웅덩이가 있어 레퍼런스와 공간적 일치도가 높음.",
        "entities": "수빈의 몸통(어두운 회색 반팔 티셔츠, 올리브색 바지 및 전술 벨트), 복부 중앙에 매우 흉측하고 입체적인 검붉은 방사능 피폭 흉터가 명확히 보임.",
        "hard_violations": [],
        "physics": "두 손이 티셔츠 밑단을 확실하게 움켜쥐고 들어 올리는 힘과 무게감이 자연스럽게 연출됨."
       },
       {
        "label": "B",
        "direction": "한 손은 티셔츠를 위로 당기고, 다른 한 손은 복부 아래쪽 바지 벨트 근처에 자연스럽게 놓여 있음.",
        "built_space": "감옥의 콘크리트 벽면과 바닥의 지푸라기 일부가 보이나, 철창 등 공간의 뚜렷한 특징이 화면에 거의 나타나지 않음.",
        "entities": "수빈의 몸통(어두운 회색 티셔츠), 복부에 흉터가 존재하나 화상이나 방사능 피폭보다는 긁힌 흉터 모양에 가까움.",
        "hard_violations": [],
        "physics": "위쪽 손이 티셔츠를 잡아당기고 아래쪽 손이 복부에 안착하여 지지하고 있으나, 위쪽 손가락의 관절 묘사가 다소 뭉개져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트의 핵심인 '검붉고 흉측한 방사능 피폭 흉터'를 매우 사실적이고 강렬하게 묘사했으며, 이전 샷의 감옥 배경(철창, 물웅덩이 등)을 완벽하게 유지하여 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "복부의 흉터가 방사능 피폭보다는 일반적인 자상이나 긁힌 상처에 가깝게 표현되었고, 감옥 배경의 디테일이 A에 비해 상대적으로 부족합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물의 두 손이 짙은 회색 티셔츠 밑단을 위로 당겨 올리는 동작에 집중되어 있음.",
        "built_space": "어두운 지하 감옥 내부. 우측에 철창 구조물이 보이고, 바닥에는 지푸라기와 물이 고인 웅덩이가 있어 레퍼런스와 공간적 일치도가 높음.",
        "entities": "수빈의 몸통(어두운 회색 반팔 티셔츠, 올리브색 바지 및 전술 벨트), 복부 중앙에 매우 흉측하고 입체적인 검붉은 방사능 피폭 흉터가 명확히 보임.",
        "hard_violations": [],
        "physics": "두 손이 티셔츠 밑단을 확실하게 움켜쥐고 들어 올리는 힘과 무게감이 자연스럽게 연출됨."
       },
       {
        "label": "B",
        "direction": "한 손은 티셔츠를 위로 당기고, 다른 한 손은 복부 아래쪽 바지 벨트 근처에 자연스럽게 놓여 있음.",
        "built_space": "감옥의 콘크리트 벽면과 바닥의 지푸라기 일부가 보이나, 철창 등 공간의 뚜렷한 특징이 화면에 거의 나타나지 않음.",
        "entities": "수빈의 몸통(어두운 회색 티셔츠), 복부에 흉터가 존재하나 화상이나 방사능 피폭보다는 긁힌 흉터 모양에 가까움.",
        "hard_violations": [],
        "physics": "위쪽 손이 티셔츠를 잡아당기고 아래쪽 손이 복부에 안착하여 지지하고 있으나, 위쪽 손가락의 관절 묘사가 다소 뭉개져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "비스듬한 복부와 검붉은 흉터가 화면을 크게 차지하고 들어 올린 티셔츠가 상단을 가로질러, 요구된 상처 중심 클로즈업을 더 충실하게 구현한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "흉측한 검붉은 흉터와 감방의 재질은 잘 구현했지만, 양팔과 배경의 비중이 커서 A보다 복부 클로즈업의 집중도가 낮다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 눈은 화면 밖이라 시선은 보이지 않는다. 화면 왼쪽 위의 손은 티셔츠를 위로 당기고, 오른쪽 아래 손은 허리선에 닿아 있다. 복부가 카메라에 비스듬히 드러나며 상처를 가리는 동작은 없다. 무기나 이동 물체는 없다.",
        "built_space": "왼쪽 배경에 거칠고 얼룩진 콘크리트 벽면 하나가 보이고, 오른쪽 아래에는 흐릿한 짚이 보인다. 철창과 잠금장치, 고인 물은 이 좁은 구도에서 확인되지 않는다. 보이는 벽의 재질과 낡은 상태는 이전 장면과 어울리며, 중복된 시설이나 불가능한 반사는 없다.",
        "entities": "한 사람의 복부와 두 팔만 보인다. 얼굴과 머리카락이 제외되어 수빈의 정확한 얼굴, 단발머리, 나이와 출신은 확인할 수 없다. 노출된 체형은 참조와 크게 충돌하지 않는다. 짙은 회색 반팔 티셔츠와 올리브색 바지는 이전 장면의 옷차림에 부합한다. 복부에는 검붉고 갈라지듯 이어진 융기성 흉터가 있으며 피부의 굴곡과 조명을 따른다. 방사능이라는 원인 자체는 외형만으로 검증할 수 없다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "위쪽 손이 티셔츠를 직접 잡아 들어 올리고 있어 옷의 지지가 분명하다. 천의 주름은 잡아당기는 방향을 따른다. 아래쪽 손은 배와 허리선에 접촉한다. 몸통과 팔의 연결은 자연스럽고, 하체의 지지점은 클로즈업 밖에 있다. 공중에 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "얼굴이 잘려 시선은 확인되지 않는다. 양손이 티셔츠의 양쪽을 잡아 위로 들어 올려 복부를 노출한다. 몸통은 약간 비스듬하지만 A보다 정면에 가깝게 보이며, 상처는 카메라 쪽으로 드러난다. 무기나 이동 물체는 없다.",
        "built_space": "왼쪽 콘크리트 벽면과 뒤쪽 벽면, 오른쪽 철창 구획 하나가 보인다. 바닥에는 벽 가까이 모인 짚과 물웅덩이가 있다. 철창은 수직 봉과 가로 보강대로 구성되며, 잠금장치는 구도 밖이다. 배경의 배치와 재질은 이전 감방 장면에 부합한다. 물 표면의 밝은 반사도 오른쪽에서 들어오는 빛과 모순되지 않는다.",
        "entities": "한 사람의 몸통과 두 팔, 양손이 보이고 다른 인물은 없다. 얼굴과 머리카락이 없어 수빈의 얼굴 동일성이나 정확한 나이·출신은 판단할 수 없다. 짙은 회색 반팔 티셔츠와 올리브색 바지는 이전 장면과 일치한다. 복부에는 검붉은 조직과 두껍게 수축한 흉터가 넓게 분포하며, 피부에 붙인 평면 무늬가 아니라 실제 굴곡으로 보인다. 손과 팔에도 상처와 때가 보인다. 방사능 손상이라는 원인은 영상만으로 확정할 수 없으며, 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "양손이 옷감을 움켜쥐어 들어 올리고 있으며 손의 접촉과 천의 당김이 연결된다. 팔꿈치를 굽힌 자세로 충분히 가능한 동작이다. 바지는 허리에 걸쳐 있고, 몸통은 화면 아래의 하체로 이어진다. 발과 바닥의 접촉은 잘려 있지만 부유를 나타내는 요소는 없고, 지지 없이 떠 있는 물체도 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "비스듬한 복부와 검붉은 흉터가 화면을 크게 차지하고 들어 올린 티셔츠가 상단을 가로질러, 요구된 상처 중심 클로즈업을 더 충실하게 구현한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "흉측한 검붉은 흉터와 감방의 재질은 잘 구현했지만, 양팔과 배경의 비중이 커서 A보다 복부 클로즈업의 집중도가 낮다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 눈은 화면 밖이라 시선은 보이지 않는다. 화면 왼쪽 위의 손은 티셔츠를 위로 당기고, 오른쪽 아래 손은 허리선에 닿아 있다. 복부가 카메라에 비스듬히 드러나며 상처를 가리는 동작은 없다. 무기나 이동 물체는 없다.",
        "built_space": "왼쪽 배경에 거칠고 얼룩진 콘크리트 벽면 하나가 보이고, 오른쪽 아래에는 흐릿한 짚이 보인다. 철창과 잠금장치, 고인 물은 이 좁은 구도에서 확인되지 않는다. 보이는 벽의 재질과 낡은 상태는 이전 장면과 어울리며, 중복된 시설이나 불가능한 반사는 없다.",
        "entities": "한 사람의 복부와 두 팔만 보인다. 얼굴과 머리카락이 제외되어 수빈의 정확한 얼굴, 단발머리, 나이와 출신은 확인할 수 없다. 노출된 체형은 참조와 크게 충돌하지 않는다. 짙은 회색 반팔 티셔츠와 올리브색 바지는 이전 장면의 옷차림에 부합한다. 복부에는 검붉고 갈라지듯 이어진 융기성 흉터가 있으며 피부의 굴곡과 조명을 따른다. 방사능이라는 원인 자체는 외형만으로 검증할 수 없다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "위쪽 손이 티셔츠를 직접 잡아 들어 올리고 있어 옷의 지지가 분명하다. 천의 주름은 잡아당기는 방향을 따른다. 아래쪽 손은 배와 허리선에 접촉한다. 몸통과 팔의 연결은 자연스럽고, 하체의 지지점은 클로즈업 밖에 있다. 공중에 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "얼굴이 잘려 시선은 확인되지 않는다. 양손이 티셔츠의 양쪽을 잡아 위로 들어 올려 복부를 노출한다. 몸통은 약간 비스듬하지만 A보다 정면에 가깝게 보이며, 상처는 카메라 쪽으로 드러난다. 무기나 이동 물체는 없다.",
        "built_space": "왼쪽 콘크리트 벽면과 뒤쪽 벽면, 오른쪽 철창 구획 하나가 보인다. 바닥에는 벽 가까이 모인 짚과 물웅덩이가 있다. 철창은 수직 봉과 가로 보강대로 구성되며, 잠금장치는 구도 밖이다. 배경의 배치와 재질은 이전 감방 장면에 부합한다. 물 표면의 밝은 반사도 오른쪽에서 들어오는 빛과 모순되지 않는다.",
        "entities": "한 사람의 몸통과 두 팔, 양손이 보이고 다른 인물은 없다. 얼굴과 머리카락이 없어 수빈의 얼굴 동일성이나 정확한 나이·출신은 판단할 수 없다. 짙은 회색 반팔 티셔츠와 올리브색 바지는 이전 장면과 일치한다. 복부에는 검붉은 조직과 두껍게 수축한 흉터가 넓게 분포하며, 피부에 붙인 평면 무늬가 아니라 실제 굴곡으로 보인다. 손과 팔에도 상처와 때가 보인다. 방사능 손상이라는 원인은 영상만으로 확정할 수 없으며, 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "양손이 옷감을 움켜쥐어 들어 올리고 있으며 손의 접촉과 천의 당김이 연결된다. 팔꿈치를 굽힌 자세로 충분히 가능한 동작이다. 바지는 허리에 걸쳐 있고, 몸통은 화면 아래의 하체로 이어진다. 발과 바닥의 접촉은 잘려 있지만 부유를 나타내는 요소는 없고, 지지 없이 떠 있는 물체도 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.7
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.7
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1700
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "프롬프트의 핵심인 '검붉고 흉측한 방사능 피폭 흉터'를 매우 사실적이고 강렬하게 묘사했으며, 이전 샷의 감옥 배경(철창, 물웅덩이 등)을 완벽하게 유지하여 훌륭하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1700,
    "verdict_ko": "복부의 흉터가 방사능 피폭보다는 일반적인 자상이나 긁힌 상처에 가깝게 표현되었고, 감옥 배경의 디테일이 A에 비해 상대적으로 부족합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10_sel.png",
    "asset_id": "4aefda3b-2871-4c34-b4ac-0bc07ffa68b4",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:779285>",
    "asset_id": "6e320331-1fae-44e2-b73c-832ea9dfbc28",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-2d3f-75e8-85c6-fbdd00ca5e91",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S59sh10"
  }
 },
 "S59sh18::cine": {
  "applied": false,
  "attempted_at": "2026-09-19T18:45:28.714704+00:00",
  "fingerprint": "3cc9d9e4139e1310b14f0cb16774e2361ad281288134fa1167feea302252c9f9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S59sh18_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S59sh18_sel.png",
  "source_sha256": "3440298fbed40e27061050e5fe60f9b68d3d7f14255b26d8187cf97afbff013b",
  "error": "RuntimeError: moderation blocked: {\"code\":\"imagine:content-moderated\",\"error\":\"Generated image rejected by content moderation.\",\"usage\":{\"cost_in_usd_ticks\":220000000}}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S59sh36::signage": {
  "fp": "d852448c9d03c41c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S59sh36": {
  "input_fingerprint": "a14dfd6ec0ba32e4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The woodcut-style flashback depicts Baeksan with an increasingly heavy body and a gold Hahoe mask; the village water has been concentrated in a tower as part of the same illustrated sequence.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The woodcut-style flashback depicts Baeksan with an increasingly heavy body and a gold Hahoe mask; the village water has been concentrated in a tower as part of the same illustrated sequence.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The woodcut-style flashback depicts Baeksan with an increasingly heavy body and a gold Hahoe mask; the village water has been concentrated in a tower as part of the same illustrated sequence.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh36__bgfirst_bg.png",
     "asset_id": "c373b774-4997-4590-9fd6-e9ba287393af",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S59sh36.png",
     "asset_id": "f5ca6eee-4fd8-46c4-909f-e7354ed6208f",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 백산의 지배를 묘사한 목판화: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1581563>",
     "asset_id": "e7bb9df0-db9f-4c6d-9c8a-088105423f0f",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B01.png",
     "asset_id": "dfd97579-913b-4622-86a5-0d6cd32f2ea6",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 백산의 지배를 묘사한 목판화: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1581563>",
     "asset_id": "e7bb9df0-db9f-4c6d-9c8a-088105423f0f",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "전경의 군중은 단상 위의 백산을 향해 시선을 던지고, 백산은 아래의 군중을 내려다봄.",
    "built_space": "지정된 나무 벽면 로케이션을 무시하고, 임의의 야외 목조 단상과 2개의 깃대가 세워진 3D 공간을 연출함.",
    "entities": "금빛 하회탈을 쓴 육중한 백산과 다양한 뒷모습을 한 전경의 군중이 묘사됨.",
    "hard_violations": [],
    "physics": "인물들은 바닥과 단상에 정상적으로 서 있으며, 깃발은 깃대에 매달려 있음."
   },
   {
    "label": "B",
    "direction": "그림 내부의 군중들이 중앙에 있는 백산을 향해 시선을 모으고 있음.",
    "built_space": "제시된 로케이션(거친 나무 벽면) 위에 그림이 그려진 사각 종이가 붙어 있는 형태임.",
    "entities": "종이 그림 속에 금빛 하회탈을 쓴 백산과 여러 군중의 모습이 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨",
     "[gpt-high] 판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
    ],
    "physics": "종이 포스터가 나무 벽면에 밀착되어 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "지정된 로케이션(나무 벽면)을 무시하고 야외 단상을 임의로 연출했으나, 텍스트 유출과 같은 치명적 위반이 없어 우선됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시문 텍스트가 이미지 하단에 캡션으로 그대로 노출되는 치명적인 텍스트 유출 위반(Hard Violation)이 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "전경의 군중은 단상 위의 백산을 향해 시선을 던지고, 백산은 아래의 군중을 내려다봄.",
        "built_space": "지정된 나무 벽면 로케이션을 무시하고, 임의의 야외 목조 단상과 2개의 깃대가 세워진 3D 공간을 연출함.",
        "entities": "금빛 하회탈을 쓴 육중한 백산과 다양한 뒷모습을 한 전경의 군중이 묘사됨.",
        "hard_violations": [],
        "physics": "인물들은 바닥과 단상에 정상적으로 서 있으며, 깃발은 깃대에 매달려 있음."
       },
       {
        "label": "B",
        "direction": "그림 내부의 군중들이 중앙에 있는 백산을 향해 시선을 모으고 있음.",
        "built_space": "제시된 로케이션(거친 나무 벽면) 위에 그림이 그려진 사각 종이가 붙어 있는 형태임.",
        "entities": "종이 그림 속에 금빛 하회탈을 쓴 백산과 여러 군중의 모습이 묘사됨.",
        "hard_violations": [
         "절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨"
        ],
        "physics": "종이 포스터가 나무 벽면에 밀착되어 지탱되고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "지정된 로케이션(나무 벽면)을 무시하고 야외 단상을 임의로 연출했으나, 텍스트 유출과 같은 치명적 위반이 없어 우선됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시문 텍스트가 이미지 하단에 캡션으로 그대로 노출되는 치명적인 텍스트 유출 위반(Hard Violation)이 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "전경의 군중은 단상 위의 백산을 향해 시선을 던지고, 백산은 아래의 군중을 내려다봄.",
        "built_space": "지정된 나무 벽면 로케이션을 무시하고, 임의의 야외 목조 단상과 2개의 깃대가 세워진 3D 공간을 연출함.",
        "entities": "금빛 하회탈을 쓴 육중한 백산과 다양한 뒷모습을 한 전경의 군중이 묘사됨.",
        "hard_violations": [],
        "physics": "인물들은 바닥과 단상에 정상적으로 서 있으며, 깃발은 깃대에 매달려 있음."
       },
       {
        "label": "B",
        "direction": "그림 내부의 군중들이 중앙에 있는 백산을 향해 시선을 모으고 있음.",
        "built_space": "제시된 로케이션(거친 나무 벽면) 위에 그림이 그려진 사각 종이가 붙어 있는 형태임.",
        "entities": "종이 그림 속에 금빛 하회탈을 쓴 백산과 여러 군중의 모습이 묘사됨.",
        "hard_violations": [
         "절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨"
        ],
        "physics": "종이 포스터가 나무 벽면에 밀착되어 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "하단에 지시문을 읽을 수 있는 자막으로 그대로 노출한 중대 위반이 있으며, 군중을 내려다보는 비스듬한 가면 대신 정면 가면과 판화 소품 자체를 보여준다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "낮은 시점의 와이드 구도와 높은 백산을 올려다보는 전경 군중의 관계는 더 정확하지만, 인물이 실사 인간이 아닌 평면 판화처럼 보이고 자연광 배경도 지정된 목판 명암 표현과 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "군중의 등과 좌우 옆얼굴은 중앙 백산을 향하며 일부는 고개를 올리고 있다. 반면 백산의 가면은 거의 정면으로 카메라를 향해 있어, 전경 군중을 향해 아래로 기울어진 조각 면이라는 지시는 약하게 구현됐다. 무기나 이동 동작은 없다.",
        "built_space": "거친 검은 목재 표면 위에 테두리가 있는 판화 한 장이 놓여 있다. 판화 내부에는 좌우의 여러 가옥과 왼쪽 뒤 원통형 물탱크 한 개가 보인다. 백산은 중앙 군중 뒤에 있지만 높이를 제공하는 단이나 지면은 보이지 않는다. 목재 바탕은 장소 참조와 가깝지만, 장면 안의 와이드 숏보다 판화 소품을 내려다보는 촬영에 가깝다.",
        "entities": "배가 크게 나온 성인 남성 백산 한 명과 다수의 군중이 판화로 표현됐다. 백산은 얼굴 전체를 덮는 온전한 금빛 웃는 탈과 어두운 겹옷을 착용한다. 탈 때문에 참조 얼굴의 동일성은 확인할 수 없으며, 가려진 얼굴을 노출하지는 않았다. 군중은 등과 옆얼굴, 옷 주름이 구별되지만 실사 인간은 아니다. 물탱크는 보이나 물 자체는 보이지 않는다. 하단에는 한국어 지시문이 명확하게 읽힌다.",
        "hard_violations": [
         "판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
        ],
        "physics": "판화 종이는 목재 표면에 밀착해 지지되는 것으로 보인다. 그림 속 백산의 하체와 군중의 발은 가려져 있어 발 접촉은 확인할 수 없지만, 공중에 떠 있는 자세는 아니다. 탈은 얼굴에 착용돼 있고 옷은 몸을 따라 내려간다. 다만 실제 배우의 체중과 자세를 촬영한 장면이 아니라 인쇄된 그림이다."
       },
       {
        "label": "B",
        "direction": "백산은 머리와 가면을 화면 오른쪽 아래 군중 쪽으로 돌리고 숙이고 있다. 가면의 앞면은 아래에서 비스듬히 보인다. 전경 군중은 중앙 높은 백산을 향해 몸을 모으고 고개를 올리며, 머리 기울기와 어깨 높이도 서로 다르다. 무기나 이동 동작은 없다.",
        "built_space": "중앙에 높은 목제 단 한 개가 있고, 앞쪽에는 판재와 가로 난간, 세 개의 굵은 기둥이 보인다. 단 양쪽에는 검은 천을 단 장대가 하나씩, 총 두 개 있다. 백산은 단 위, 군중은 그 아래 전경에 배치돼 높이 관계가 명확하다. 카메라는 군중 높이보다 낮은 곳에서 올려다본다. 뒤로 하늘과 산, 일부 지붕이 보이며, 참조의 거친 검은 목재 재질은 단에 부분적으로 반영됐지만 같은 목재 표면 중심의 공간은 아니다.",
        "entities": "백산은 배가 나온 성인 남성으로 표현됐고, 온전한 금빛 하회탈과 어두운 긴 겹옷을 착용한다. 얼굴은 가려져 참조 인물의 얼굴 동일성을 확인할 수 없다. 머리는 참조의 짧은 머리 대신 상투 형태다. 아래 군중은 각각 구별되는 머리와 등, 일부 옆얼굴을 갖지만 모두 굵은 판화 선으로 그려져 실사 인물로 읽히지 않는다. 물탱크는 보이지 않으며, 글자나 로고도 보이지 않는다.",
        "hard_violations": [],
        "physics": "백산의 발은 단 앞벽과 옷자락에 가려져 있지만 몸 아래에 지지할 수 있는 단이 있어 부유로 보이지 않는다. 군중의 하체는 프레임 밖이거나 서로 가려져 있으며 서 있는 상체 배치는 가능하다. 두 천은 각각 장대의 가로대에 매달려 지지된다. 다만 백산과 군중의 평면적인 판화 표현이 실제 목재 구조 및 하늘과 분리돼 보여, 신체와 의복의 물리적 실재감은 부족하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "하단에 지시문을 읽을 수 있는 자막으로 그대로 노출한 중대 위반이 있으며, 군중을 내려다보는 비스듬한 가면 대신 정면 가면과 판화 소품 자체를 보여준다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "낮은 시점의 와이드 구도와 높은 백산을 올려다보는 전경 군중의 관계는 더 정확하지만, 인물이 실사 인간이 아닌 평면 판화처럼 보이고 자연광 배경도 지정된 목판 명암 표현과 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "군중의 등과 좌우 옆얼굴은 중앙 백산을 향하며 일부는 고개를 올리고 있다. 반면 백산의 가면은 거의 정면으로 카메라를 향해 있어, 전경 군중을 향해 아래로 기울어진 조각 면이라는 지시는 약하게 구현됐다. 무기나 이동 동작은 없다.",
        "built_space": "거친 검은 목재 표면 위에 테두리가 있는 판화 한 장이 놓여 있다. 판화 내부에는 좌우의 여러 가옥과 왼쪽 뒤 원통형 물탱크 한 개가 보인다. 백산은 중앙 군중 뒤에 있지만 높이를 제공하는 단이나 지면은 보이지 않는다. 목재 바탕은 장소 참조와 가깝지만, 장면 안의 와이드 숏보다 판화 소품을 내려다보는 촬영에 가깝다.",
        "entities": "배가 크게 나온 성인 남성 백산 한 명과 다수의 군중이 판화로 표현됐다. 백산은 얼굴 전체를 덮는 온전한 금빛 웃는 탈과 어두운 겹옷을 착용한다. 탈 때문에 참조 얼굴의 동일성은 확인할 수 없으며, 가려진 얼굴을 노출하지는 않았다. 군중은 등과 옆얼굴, 옷 주름이 구별되지만 실사 인간은 아니다. 물탱크는 보이나 물 자체는 보이지 않는다. 하단에는 한국어 지시문이 명확하게 읽힌다.",
        "hard_violations": [
         "판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
        ],
        "physics": "판화 종이는 목재 표면에 밀착해 지지되는 것으로 보인다. 그림 속 백산의 하체와 군중의 발은 가려져 있어 발 접촉은 확인할 수 없지만, 공중에 떠 있는 자세는 아니다. 탈은 얼굴에 착용돼 있고 옷은 몸을 따라 내려간다. 다만 실제 배우의 체중과 자세를 촬영한 장면이 아니라 인쇄된 그림이다."
       },
       {
        "label": "A",
        "direction": "백산은 머리와 가면을 화면 오른쪽 아래 군중 쪽으로 돌리고 숙이고 있다. 가면의 앞면은 아래에서 비스듬히 보인다. 전경 군중은 중앙 높은 백산을 향해 몸을 모으고 고개를 올리며, 머리 기울기와 어깨 높이도 서로 다르다. 무기나 이동 동작은 없다.",
        "built_space": "중앙에 높은 목제 단 한 개가 있고, 앞쪽에는 판재와 가로 난간, 세 개의 굵은 기둥이 보인다. 단 양쪽에는 검은 천을 단 장대가 하나씩, 총 두 개 있다. 백산은 단 위, 군중은 그 아래 전경에 배치돼 높이 관계가 명확하다. 카메라는 군중 높이보다 낮은 곳에서 올려다본다. 뒤로 하늘과 산, 일부 지붕이 보이며, 참조의 거친 검은 목재 재질은 단에 부분적으로 반영됐지만 같은 목재 표면 중심의 공간은 아니다.",
        "entities": "백산은 배가 나온 성인 남성으로 표현됐고, 온전한 금빛 하회탈과 어두운 긴 겹옷을 착용한다. 얼굴은 가려져 참조 인물의 얼굴 동일성을 확인할 수 없다. 머리는 참조의 짧은 머리 대신 상투 형태다. 아래 군중은 각각 구별되는 머리와 등, 일부 옆얼굴을 갖지만 모두 굵은 판화 선으로 그려져 실사 인물로 읽히지 않는다. 물탱크는 보이지 않으며, 글자나 로고도 보이지 않는다.",
        "hard_violations": [],
        "physics": "백산의 발은 단 앞벽과 옷자락에 가려져 있지만 몸 아래에 지지할 수 있는 단이 있어 부유로 보이지 않는다. 군중의 하체는 프레임 밖이거나 서로 가려져 있으며 서 있는 상체 배치는 가능하다. 두 천은 각각 장대의 가로대에 매달려 지지된다. 다만 백산과 군중의 평면적인 판화 표현이 실제 목재 구조 및 하늘과 분리돼 보여, 신체와 의복의 물리적 실재감은 부족하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨",
     "[gpt-high] 판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 로케이션(나무 벽면)을 무시하고 야외 단상을 임의로 연출했으나, 텍스트 유출과 같은 치명적 위반이 없어 우선됨."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "지시문 텍스트가 이미지 하단에 캡션으로 그대로 노출되는 치명적인 텍스트 유출 위반(Hard Violation)이 발생함.  ★위반: [gemini-pro] 절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨 / [gpt-high] 판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B01.png",
    "asset_id": "dfd97579-913b-4622-86a5-0d6cd32f2ea6",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1241731>",
    "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 백산의 지배를 묘사한 목판화: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:1581563>",
    "asset_id": "e7bb9df0-db9f-4c6d-9c8a-088105423f0f",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-2f04-7fc1-9f0d-dca44cf2a978",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh36__bgfirst_bg.png",
   "bg_asset_id": "c373b774-4997-4590-9fd6-e9ba287393af",
   "bg_record_key": "S59sh36::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S59sh36::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:45:24.751027+00:00",
  "fingerprint": "23d7fee06eb7062de4f91135ca2e492a9dd59f3961bd63dfe3ef9a44f4a61b0b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S59sh36_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S59sh36_sel.png",
  "source_sha256": "707e0641b89f8cbfc4182752f6d7a5cd9f9c153105eb5b43f0e35bfec1416ae5",
  "file": "S59sh36_cine.png",
  "staged_sha256": "e9ee06592bb1dab36b60fa954fb028ef73512631c027ee6b44d15856baee84e3",
  "latency_ms": 9771
 },
 "S60sh4::signage": {
  "fp": "d91c2cce8e6ded46",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S60sh4": {
  "input_fingerprint": "6eb4f0fef1cdaa93",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The arena stands in front of the church, lit by torches and late-day light, with Charlie's raised cage open. Charlie retains the worn Ubik chest logo and his straw hat, colorful raincoat and oversized boots, while the larger, heavily built B-200 has converted gun-hands.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The arena stands in front of the church, lit by torches and late-day light, with Charlie's raised cage open. Charlie retains the worn Ubik chest logo and his straw hat, colorful raincoat and oversized boots, while the larger, heavily built B-200 has converted gun-hands.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The arena stands in front of the church, lit by torches and late-day light, with Charlie's raised cage open. Charlie retains the worn Ubik chest logo and his straw hat, colorful raincoat and oversized boots, while the larger, heavily built B-200 has converted gun-hands.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh4__bgfirst_bg.png",
     "asset_id": "ab433a49-41f6-42f3-9c2c-d68dd6bf3b32",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S60sh4.png",
     "asset_id": "6a6601d1-910a-4eba-a636-eb4b443cdfb9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B01.png",
     "asset_id": "239110df-f90f-4ca2-b3ef-4a8f15c13fe9",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_village_church_arena_sel.png",
     "asset_id": "7cc6d446-071f-4801-886a-a504b66fe88c",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 등을 보이며 B-200을 향해 서 있고, B-200은 찰리를 마주보며 양손의 무기를 그에게 겨냥하고 있습니다.",
    "built_space": "카메라는 격투장 바닥에 위치하며 지시대로 왼쪽 가장자리에 케이지의 열린 입구가 배치되어 있습니다. 그러나 배경의 교회는 필수 지정된 'STRUCTURE LOOK'의 폐허 형태가 아니라 'LOCATION' 사진의 교회를 따르고 있습니다.",
    "entities": "찰리는 챙이 있는 모자와 코트를 입고 있으나, 가슴에 있어야 할 'Ubik' 로고가 등 뒤에 크고 선명하게 적혀 있습니다. B-200은 거대한 체구와 중장비 형태를 잘 유지하고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 규칙 위반 (찰리 코트 뒷면에 'Ubik'이 명확하게 읽힘)",
     "[gemini-pro] 지정된 구조물 레퍼런스 미반영 (폐허가 된 구조물 레퍼런스 대신 장소 사진의 교회를 렌더링함)",
     "[gpt-high] 찰리 등판에 'Ubik'이라는 판독 가능한 로고를 노출하여 읽을 수 있는 글자와 로고 금지 조건을 위반했습니다.",
     "[gpt-high] 등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 여러 인물을 추가했습니다."
    ],
    "physics": "찰리와 B-200 모두 격투장 바닥에 안정적으로 발을 딛고 서 있으며 물리적으로 어색한 부분은 없습니다."
   },
   {
    "label": "B",
    "direction": "찰리와 B-200이 측면 프로필 구도로 서로를 마주보고 있으며, B-200의 무기가 찰리 쪽을 향하고 있습니다.",
    "built_space": "카메라가 관중석 난간 뒤에 위치하여 전경을 가리고 있으며, 지정된 케이지 출발점 프레이밍을 어겼습니다. 배경의 교회는 레퍼런스에 전혀 존재하지 않는 양쪽의 거대한 사각형 탑이 추가된 형태로 창작되었습니다.",
    "entities": "찰리는 고릴라형 체구와 지정된 의상을 입고 측면을 보이고 있습니다. B-200은 레퍼런스에 부합하는 거대한 무장 형태를 띱니다. 관중석에는 다수의 인물들이 묘사되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 존재하지 않는 구조물 창작 (레퍼런스에 없는 거대한 양쪽 탑을 가진 교회를 임의로 렌더링함)",
     "[gemini-pro] 지정된 뷰포인트 위반 (케이지를 출발점으로 삼지 않고 관중석 난간 뒤에서 촬영된 구도로 렌더링함)",
     "[gpt-high] 등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 다수의 인물을 추가했습니다.",
     "[gpt-high] 하나의 오른쪽 종탑을 가진 고정 교회를 두 개의 큰 탑이 있는 벽돌 교회로 대체하여 구조의 수와 정체성을 바꾸었습니다."
    ],
    "physics": "두 로봇 모두 흙바닥에 안정적으로 지지되어 서 있으며 중력에 반하는 묘사는 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 프레이밍(왼쪽 끝 케이지 배치, 넓은 샷)은 정확히 구현했으나, 읽을 수 없는 텍스트 규칙을 위반하여 코트 뒷면에 'Ubik'을 명확히 노출했고, 고정 구조물 레퍼런스 대신 장소 사진의 교회를 그대로 렌더링한 치명적 결함이 있습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "카메라 시점을 관중석 뒤로 배치하여 프레이밍 지시를 완전히 어겼으며, 어떤 레퍼런스에도 없는 거대한 두 개의 사각형 탑을 가진 교회를 임의로 창작하여 명백한 실격 사유에 해당합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 등을 보이며 B-200을 향해 서 있고, B-200은 찰리를 마주보며 양손의 무기를 그에게 겨냥하고 있습니다.",
        "built_space": "카메라는 격투장 바닥에 위치하며 지시대로 왼쪽 가장자리에 케이지의 열린 입구가 배치되어 있습니다. 그러나 배경의 교회는 필수 지정된 'STRUCTURE LOOK'의 폐허 형태가 아니라 'LOCATION' 사진의 교회를 따르고 있습니다.",
        "entities": "찰리는 챙이 있는 모자와 코트를 입고 있으나, 가슴에 있어야 할 'Ubik' 로고가 등 뒤에 크고 선명하게 적혀 있습니다. B-200은 거대한 체구와 중장비 형태를 잘 유지하고 있습니다.",
        "hard_violations": [
         "화면 내 읽을 수 있는 텍스트 금지 규칙 위반 (찰리 코트 뒷면에 'Ubik'이 명확하게 읽힘)",
         "지정된 구조물 레퍼런스 미반영 (폐허가 된 구조물 레퍼런스 대신 장소 사진의 교회를 렌더링함)"
        ],
        "physics": "찰리와 B-200 모두 격투장 바닥에 안정적으로 발을 딛고 서 있으며 물리적으로 어색한 부분은 없습니다."
       },
       {
        "label": "B",
        "direction": "찰리와 B-200이 측면 프로필 구도로 서로를 마주보고 있으며, B-200의 무기가 찰리 쪽을 향하고 있습니다.",
        "built_space": "카메라가 관중석 난간 뒤에 위치하여 전경을 가리고 있으며, 지정된 케이지 출발점 프레이밍을 어겼습니다. 배경의 교회는 레퍼런스에 전혀 존재하지 않는 양쪽의 거대한 사각형 탑이 추가된 형태로 창작되었습니다.",
        "entities": "찰리는 고릴라형 체구와 지정된 의상을 입고 측면을 보이고 있습니다. B-200은 레퍼런스에 부합하는 거대한 무장 형태를 띱니다. 관중석에는 다수의 인물들이 묘사되어 있습니다.",
        "hard_violations": [
         "존재하지 않는 구조물 창작 (레퍼런스에 없는 거대한 양쪽 탑을 가진 교회를 임의로 렌더링함)",
         "지정된 뷰포인트 위반 (케이지를 출발점으로 삼지 않고 관중석 난간 뒤에서 촬영된 구도로 렌더링함)"
        ],
        "physics": "두 로봇 모두 흙바닥에 안정적으로 지지되어 서 있으며 중력에 반하는 묘사는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 프레이밍(왼쪽 끝 케이지 배치, 넓은 샷)은 정확히 구현했으나, 읽을 수 없는 텍스트 규칙을 위반하여 코트 뒷면에 'Ubik'을 명확히 노출했고, 고정 구조물 레퍼런스 대신 장소 사진의 교회를 그대로 렌더링한 치명적 결함이 있습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "카메라 시점을 관중석 뒤로 배치하여 프레이밍 지시를 완전히 어겼으며, 어떤 레퍼런스에도 없는 거대한 두 개의 사각형 탑을 가진 교회를 임의로 창작하여 명백한 실격 사유에 해당합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 등을 보이며 B-200을 향해 서 있고, B-200은 찰리를 마주보며 양손의 무기를 그에게 겨냥하고 있습니다.",
        "built_space": "카메라는 격투장 바닥에 위치하며 지시대로 왼쪽 가장자리에 케이지의 열린 입구가 배치되어 있습니다. 그러나 배경의 교회는 필수 지정된 'STRUCTURE LOOK'의 폐허 형태가 아니라 'LOCATION' 사진의 교회를 따르고 있습니다.",
        "entities": "찰리는 챙이 있는 모자와 코트를 입고 있으나, 가슴에 있어야 할 'Ubik' 로고가 등 뒤에 크고 선명하게 적혀 있습니다. B-200은 거대한 체구와 중장비 형태를 잘 유지하고 있습니다.",
        "hard_violations": [
         "화면 내 읽을 수 있는 텍스트 금지 규칙 위반 (찰리 코트 뒷면에 'Ubik'이 명확하게 읽힘)",
         "지정된 구조물 레퍼런스 미반영 (폐허가 된 구조물 레퍼런스 대신 장소 사진의 교회를 렌더링함)"
        ],
        "physics": "찰리와 B-200 모두 격투장 바닥에 안정적으로 발을 딛고 서 있으며 물리적으로 어색한 부분은 없습니다."
       },
       {
        "label": "B",
        "direction": "찰리와 B-200이 측면 프로필 구도로 서로를 마주보고 있으며, B-200의 무기가 찰리 쪽을 향하고 있습니다.",
        "built_space": "카메라가 관중석 난간 뒤에 위치하여 전경을 가리고 있으며, 지정된 케이지 출발점 프레이밍을 어겼습니다. 배경의 교회는 레퍼런스에 전혀 존재하지 않는 양쪽의 거대한 사각형 탑이 추가된 형태로 창작되었습니다.",
        "entities": "찰리는 고릴라형 체구와 지정된 의상을 입고 측면을 보이고 있습니다. B-200은 레퍼런스에 부합하는 거대한 무장 형태를 띱니다. 관중석에는 다수의 인물들이 묘사되어 있습니다.",
        "hard_violations": [
         "존재하지 않는 구조물 창작 (레퍼런스에 없는 거대한 양쪽 탑을 가진 교회를 임의로 렌더링함)",
         "지정된 뷰포인트 위반 (케이지를 출발점으로 삼지 않고 관중석 난간 뒤에서 촬영된 구도로 렌더링함)"
        ],
        "physics": "두 로봇 모두 흙바닥에 안정적으로 지지되어 서 있으며 중력에 반하는 묘사는 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "B-200의 압도적인 크기는 드러나지만, 두 로봇이 서로 대치하기보다 카메라 쪽으로 나란히 서 있고 케이지 배치와 고정 교회 구조도 어긋납니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "왼쪽 가장자리의 열린 케이지와 찰리 너머 B-200의 대치 구도는 더 정확하지만, 읽히는 등판 로고와 허용되지 않은 관중 때문에 최종 사용에는 부적합합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 얼굴과 가슴을 주로 카메라 쪽으로 보이며 B-200을 명확히 바라보지 않습니다. B-200은 몸을 카메라 쪽으로 열고 머리를 약간 화면 왼쪽으로 돌렸습니다. 화면 왼쪽으로 뻗은 포신들은 찰리 머리보다 높은 공간을 향하고, 반대쪽 포신들도 찰리에게 정확히 수렴하지 않습니다. 발사 순간은 아니지만, 서로 마주 선 대치 관계가 약합니다.",
        "built_space": "중앙 교회 한 채, 그 양쪽의 높은 탑 두 개, 중앙 출입문과 계단, 좌우 관람석, 왼쪽 케이지 한 개가 보입니다. 큰 화염대는 위쪽 양끝과 아래쪽 좌우에 총 네 개 보입니다. 교회는 구조 기준의 붕괴된 회색 석조 정면과 오른쪽 종탑 대신 온전한 붉은 벽돌 쌍탑 건물입니다. 케이지는 왼쪽 끝의 출입구 일부가 아니라 찰리 뒤에 상당 부분 노출됩니다. 전경 관람석 난간과 벤치가 들어와 지정된 케이지 출발점보다 관람석에서 바라본 구도로 읽힙니다.",
        "entities": "로봇은 두 대입니다. B-200의 짙은 회색 중장갑, 굵은 관절, 양팔의 다연장 포신은 기준과 대체로 맞습니다. 찰리는 베이지 장갑, 흰 마스크형 얼굴, 긴 팔과 짧은 다리로 기준 체형에 가깝지만 밀짚모자와 화려한 우비 대신 갈색 모자와 단색 외투를 입고 있습니다. 가슴의 마모된 지정 로고는 확인되지 않습니다. 양쪽 관람석에 다수의 사람이 추가되어 있습니다. 이들은 동아시아계로 보이는 성인 남녀가 중심이지만 개별 국적은 판별할 수 없습니다. 판독 가능한 문구는 보이지 않습니다.",
        "hard_violations": [
         "등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 다수의 인물을 추가했습니다.",
         "하나의 오른쪽 종탑을 가진 고정 교회를 두 개의 큰 탑이 있는 벽돌 교회로 대체하여 구조의 수와 정체성을 바꾸었습니다."
        ],
        "physics": "두 로봇 모두 넓은 발을 흙바닥에 붙이고 서 있으며, 굽힌 다리와 벌어진 발이 중량을 지탱합니다. 포신은 팔 관절에 연결되어 있고 외투와 모자는 각각 몸과 머리에 지지됩니다. 관중은 계단식 좌석에 앉거나 바닥에 서 있습니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "찰리는 등을 카메라에 보이고 머리와 몸을 앞쪽의 B-200으로 향합니다. B-200도 찰리 쪽으로 정면을 열어 서로 대치하는 관계가 분명합니다. 화면 왼쪽 포신은 찰리 쪽의 왼쪽 전방을 향하고, 오른쪽 포신은 카메라 방향으로 더 열려 있어 두 무장이 한 지점에 정확히 조준된 상태는 아닙니다. 다만 요구된 순간은 사격이 아니라 위협적으로 서 있는 순간입니다.",
        "built_space": "열린 케이지 출입구 한 곳의 가까운 문틀이 왼쪽 가장자리에 걸리고, 찰리는 그 앞의 경기장 바닥에 서 있습니다. B-200은 빈 흙바닥 간격을 두고 맞은편에 있어 지정된 공간 관계를 잘 보여줍니다. 중앙 교회 출입문 한 곳과 계단, 좌우 관람석, 네 곳의 불꽃이 보입니다. 교회의 회색 재료와 원형 창은 구조 기준에 가까워졌지만 정면은 기준보다 온전하고 장식 및 개구부 배치도 다릅니다. 찰리가 전경에서 화면상 더 커져 B-200의 실제 크기 우위는 A만큼 즉각적이지 않습니다.",
        "entities": "로봇 두 대가 보이며 B-200은 회색 중장갑과 양팔 포신을 갖췄습니다. 찰리는 밀짚모자와 베이지 장갑을 유지하지만 외투가 단색이고 다리가 길어 기준의 고릴라형 비례가 약해졌습니다. 얼굴과 가슴은 뒤돌아선 구도 때문에 보이지 않으므로 그 부분의 일치 여부는 확인할 수 없습니다. 대신 등판에 'Ubik'이라는 큰 글자가 선명하게 읽힙니다. 좌우 관람석에는 허용된 두 로봇 외에 여러 사람이 있으며, 대부분 동아시아계 성인처럼 보이나 국적은 확인할 수 없습니다.",
        "hard_violations": [
         "찰리 등판에 'Ubik'이라는 판독 가능한 로고를 노출하여 읽을 수 있는 글자와 로고 금지 조건을 위반했습니다.",
         "등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 여러 인물을 추가했습니다."
        ],
        "physics": "찰리의 두 발과 B-200의 넓게 벌린 두 발이 모두 흙바닥에 접촉해 체중을 지탱합니다. B-200의 무장은 팔에 연결되어 있고 찰리의 모자는 머리에, 외투는 어깨와 몸에 걸려 있습니다. 열린 케이지 문틀은 바닥과 구조체에 연결되어 있습니다. 관중도 좌석이나 계단에 지지되며, 공중에 뜬 물체나 설명되지 않는 도약은 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "B-200의 압도적인 크기는 드러나지만, 두 로봇이 서로 대치하기보다 카메라 쪽으로 나란히 서 있고 케이지 배치와 고정 교회 구조도 어긋납니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "왼쪽 가장자리의 열린 케이지와 찰리 너머 B-200의 대치 구도는 더 정확하지만, 읽히는 등판 로고와 허용되지 않은 관중 때문에 최종 사용에는 부적합합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 얼굴과 가슴을 주로 카메라 쪽으로 보이며 B-200을 명확히 바라보지 않습니다. B-200은 몸을 카메라 쪽으로 열고 머리를 약간 화면 왼쪽으로 돌렸습니다. 화면 왼쪽으로 뻗은 포신들은 찰리 머리보다 높은 공간을 향하고, 반대쪽 포신들도 찰리에게 정확히 수렴하지 않습니다. 발사 순간은 아니지만, 서로 마주 선 대치 관계가 약합니다.",
        "built_space": "중앙 교회 한 채, 그 양쪽의 높은 탑 두 개, 중앙 출입문과 계단, 좌우 관람석, 왼쪽 케이지 한 개가 보입니다. 큰 화염대는 위쪽 양끝과 아래쪽 좌우에 총 네 개 보입니다. 교회는 구조 기준의 붕괴된 회색 석조 정면과 오른쪽 종탑 대신 온전한 붉은 벽돌 쌍탑 건물입니다. 케이지는 왼쪽 끝의 출입구 일부가 아니라 찰리 뒤에 상당 부분 노출됩니다. 전경 관람석 난간과 벤치가 들어와 지정된 케이지 출발점보다 관람석에서 바라본 구도로 읽힙니다.",
        "entities": "로봇은 두 대입니다. B-200의 짙은 회색 중장갑, 굵은 관절, 양팔의 다연장 포신은 기준과 대체로 맞습니다. 찰리는 베이지 장갑, 흰 마스크형 얼굴, 긴 팔과 짧은 다리로 기준 체형에 가깝지만 밀짚모자와 화려한 우비 대신 갈색 모자와 단색 외투를 입고 있습니다. 가슴의 마모된 지정 로고는 확인되지 않습니다. 양쪽 관람석에 다수의 사람이 추가되어 있습니다. 이들은 동아시아계로 보이는 성인 남녀가 중심이지만 개별 국적은 판별할 수 없습니다. 판독 가능한 문구는 보이지 않습니다.",
        "hard_violations": [
         "등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 다수의 인물을 추가했습니다.",
         "하나의 오른쪽 종탑을 가진 고정 교회를 두 개의 큰 탑이 있는 벽돌 교회로 대체하여 구조의 수와 정체성을 바꾸었습니다."
        ],
        "physics": "두 로봇 모두 넓은 발을 흙바닥에 붙이고 서 있으며, 굽힌 다리와 벌어진 발이 중량을 지탱합니다. 포신은 팔 관절에 연결되어 있고 외투와 모자는 각각 몸과 머리에 지지됩니다. 관중은 계단식 좌석에 앉거나 바닥에 서 있습니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "찰리는 등을 카메라에 보이고 머리와 몸을 앞쪽의 B-200으로 향합니다. B-200도 찰리 쪽으로 정면을 열어 서로 대치하는 관계가 분명합니다. 화면 왼쪽 포신은 찰리 쪽의 왼쪽 전방을 향하고, 오른쪽 포신은 카메라 방향으로 더 열려 있어 두 무장이 한 지점에 정확히 조준된 상태는 아닙니다. 다만 요구된 순간은 사격이 아니라 위협적으로 서 있는 순간입니다.",
        "built_space": "열린 케이지 출입구 한 곳의 가까운 문틀이 왼쪽 가장자리에 걸리고, 찰리는 그 앞의 경기장 바닥에 서 있습니다. B-200은 빈 흙바닥 간격을 두고 맞은편에 있어 지정된 공간 관계를 잘 보여줍니다. 중앙 교회 출입문 한 곳과 계단, 좌우 관람석, 네 곳의 불꽃이 보입니다. 교회의 회색 재료와 원형 창은 구조 기준에 가까워졌지만 정면은 기준보다 온전하고 장식 및 개구부 배치도 다릅니다. 찰리가 전경에서 화면상 더 커져 B-200의 실제 크기 우위는 A만큼 즉각적이지 않습니다.",
        "entities": "로봇 두 대가 보이며 B-200은 회색 중장갑과 양팔 포신을 갖췄습니다. 찰리는 밀짚모자와 베이지 장갑을 유지하지만 외투가 단색이고 다리가 길어 기준의 고릴라형 비례가 약해졌습니다. 얼굴과 가슴은 뒤돌아선 구도 때문에 보이지 않으므로 그 부분의 일치 여부는 확인할 수 없습니다. 대신 등판에 'Ubik'이라는 큰 글자가 선명하게 읽힙니다. 좌우 관람석에는 허용된 두 로봇 외에 여러 사람이 있으며, 대부분 동아시아계 성인처럼 보이나 국적은 확인할 수 없습니다.",
        "hard_violations": [
         "찰리 등판에 'Ubik'이라는 판독 가능한 로고를 노출하여 읽을 수 있는 글자와 로고 금지 조건을 위반했습니다.",
         "등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 여러 인물을 추가했습니다."
        ],
        "physics": "찰리의 두 발과 B-200의 넓게 벌린 두 발이 모두 흙바닥에 접촉해 체중을 지탱합니다. B-200의 무장은 팔에 연결되어 있고 찰리의 모자는 머리에, 외투는 어깨와 몸에 걸려 있습니다. 열린 케이지 문틀은 바닥과 구조체에 연결되어 있습니다. 관중도 좌석이나 계단에 지지되며, 공중에 뜬 물체나 설명되지 않는 도약은 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 규칙 위반 (찰리 코트 뒷면에 'Ubik'이 명확하게 읽힘)",
     "[gemini-pro] 지정된 구조물 레퍼런스 미반영 (폐허가 된 구조물 레퍼런스 대신 장소 사진의 교회를 렌더링함)",
     "[gpt-high] 찰리 등판에 'Ubik'이라는 판독 가능한 로고를 노출하여 읽을 수 있는 글자와 로고 금지 조건을 위반했습니다.",
     "[gpt-high] 등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 여러 인물을 추가했습니다."
    ],
    "B": [
     "[gemini-pro] 존재하지 않는 구조물 창작 (레퍼런스에 없는 거대한 양쪽 탑을 가진 교회를 임의로 렌더링함)",
     "[gemini-pro] 지정된 뷰포인트 위반 (케이지를 출발점으로 삼지 않고 관중석 난간 뒤에서 촬영된 구도로 렌더링함)",
     "[gpt-high] 등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 다수의 인물을 추가했습니다.",
     "[gpt-high] 하나의 오른쪽 종탑을 가진 고정 교회를 두 개의 큰 탑이 있는 벽돌 교회로 대체하여 구조의 수와 정체성을 바꾸었습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 프레이밍(왼쪽 끝 케이지 배치, 넓은 샷)은 정확히 구현했으나, 읽을 수 없는 텍스트 규칙을 위반하여 코트 뒷면에 'Ubik'을 명확히 노출했고, 고정 구조물 레퍼런스 대신 장소 사진의 교회를 그대로 렌더링한 치명적 결함이 있습니다.  ★위반: [gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 규칙 위반 (찰리 코트 뒷면에 'Ubik'이 명확하게 읽힘) / [gemini-pro] 지정된 구조물 레퍼런스 미반영 (폐허가 된 구조물 레퍼런스 대신 장소 사진의 교회를 렌더링함) / [gpt-high] 찰리 등판에 'Ubik'이라는 판독 가능한 로고를 노출하여 읽을 수 있는 글자와 로고 금지 조건을 위반했습니다. / [gpt-high] 등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 여러 인물을 추가했습니다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "카메라 시점을 관중석 뒤로 배치하여 프레이밍 지시를 완전히 어겼으며, 어떤 레퍼런스에도 없는 거대한 두 개의 사각형 탑을 가진 교회를 임의로 창작하여 명백한 실격 사유에 해당합니다.  ★위반: [gemini-pro] 존재하지 않는 구조물 창작 (레퍼런스에 없는 거대한 양쪽 탑을 가진 교회를 임의로 렌더링함) / [gemini-pro] 지정된 뷰포인트 위반 (케이지를 출발점으로 삼지 않고 관중석 난간 뒤에서 촬영된 구도로 렌더링함) / [gpt-high] 등장 인물을 두 로봇으로 제한한 조건과 달리 관람석에 다수의 인물을 추가했습니다. / [gpt-high] 하나의 오른쪽 종탑을 가진 고정 교회를 두 개의 큰 탑이 있는 벽돌 교회로 대체하여 구조의 수와 정체성을 바꾸었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B01.png",
    "asset_id": "239110df-f90f-4ca2-b3ef-4a8f15c13fe9",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_village_church_arena_sel.png",
    "asset_id": "7cc6d446-071f-4801-886a-a504b66fe88c",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-3267-731d-8890-b608a883d20c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh4__bgfirst_bg.png",
   "bg_asset_id": "ab433a49-41f6-42f3-9c2c-d68dd6bf3b32",
   "bg_record_key": "S60sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready",
  "staged_characters_added": [
   "C06"
  ]
 },
 "S60sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:49:17.783310+00:00",
  "fingerprint": "f1241306a6bfc6405122f01788423fb4299985b7cac25b4e2351a757799bab67",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S60sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S60sh4_sel.png",
  "source_sha256": "b8dbe3ee42a4412a031c6a0b0d5d5bbbacad2eb0094d38c7f2793f30cca87988",
  "file": "S60sh4_cine.png",
  "staged_sha256": "4365296b6842ce72dc63dcfc28be6ff8f664593df54be07fe83a11f8b738a0b1",
  "latency_ms": 10045
 },
 "S60sh52::signage": {
  "fp": "2a94887d8e4d47b7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::6c9ae0e716405f00": {
  "subjects": [],
  "subject_text": "익산 마을 성당 앞 격투장과 관중석\n성당 앞 마을 중심에 마련된 야외 격투장. 중앙 경기 구역 둘레에 관중석과 횃불이 배치되고 바닥에는 케이지 승강구가 있다.",
  "identity": "canonical",
  "scope_id": "L229",
  "scope_role": "location_exterior",
  "scope_sha": "cd835905025bc38b"
 },
 "S60sh52::bgfirst_bg": {
  "input_fingerprint": "a6abcd17ac433680",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh52__bgfirst_bg.png",
  "asset_id": "aeb46e86-9443-4079-ae58-e593ffa40546",
  "input_asset_ids": [
   "4d52529a-4346-46e6-91fd-c388e0037459",
   "1f832f0c-4b90-4195-8434-227eb2c715f6"
  ]
 },
 "S60sh52": {
  "input_fingerprint": "33b40bec9007e63f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The water tower has been shot through and toppled, releasing a large flood of stored water in the sunset light. Charlie bears accumulated dents and holes from the fighting, while B-200's deployed gun-hands remain intact. 백산: He is soaking wet, with his gold Hahoe mask half broken and his radiation-disfigured face exposed. His royal-style clothing has not yet been stripped away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The water tower has been shot through and toppled, releasing a large flood of stored water in the sunset light. Charlie bears accumulated dents and holes from the fighting, while B-200's deployed gun-hands remain intact. 백산: He is soaking wet, with his gold Hahoe mask half broken and his radiation-disfigured face exposed. His royal-style clothing has not yet been stripped away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The water tower has been shot through and toppled, releasing a large flood of stored water in the sunset light. Charlie bears accumulated dents and holes from the fighting, while B-200's deployed gun-hands remain intact. 백산: He is soaking wet, with his gold Hahoe mask half broken and his radiation-disfigured face exposed. His royal-style clothing has not yet been stripped away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh52__bgfirst_bg.png",
     "asset_id": "aeb46e86-9443-4079-ae58-e593ffa40546",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S60sh52.png",
     "asset_id": "4d52529a-4346-46e6-91fd-c388e0037459",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1083564>",
     "asset_id": "9a307034-865c-47d1-bf2c-6230188e67c7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B03.png",
     "asset_id": "1f832f0c-4b90-4195-8434-227eb2c715f6",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1083564>",
     "asset_id": "9a307034-865c-47d1-bf2c-6230188e67c7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 우측 상단의 허공을 향하고 있습니다.",
    "built_space": "레퍼런스의 성당 건물과 물탱크 파이프가 배경에 배치되었으나, 우측에 프롬프트에 없는 인물의 다리가 자리 잡고 있습니다.",
    "entities": "백산의 화상 얼굴, 깨진 황금 가면, 복장이 나타나지만 프레임 우측에 검은 부츠를 신은 정체불명의 인물이 추가되었습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 지시되지 않은 추가 인물 등장 (우측 검은 부츠)",
     "[gemini-pro] 배경 우측에 아무런 지지대 없이 허공에 떠 있는 의자",
     "[gpt-high] 백산만 등장해야 하는 장면에 검은 옷과 장화를 신은 다른 인물의 하체를 추가했다."
    ],
    "physics": "우측 배경의 나무 의자는 공중에 떠 있으며 어떤 지지대도 없습니다. 백산의 자세는 바닥에 앉은 형태를 취하고 있습니다."
   },
   {
    "label": "B",
    "direction": "시선은 정면 아래쪽을 공허하게 응시하고 있습니다.",
    "built_space": "로케이션 레퍼런스의 무너진 물탱크 철골과 파이프 잔해가 노을빛을 받는 물웅덩이와 함께 배치되었습니다.",
    "entities": "백산의 얼굴 형태, 절반이 깨진 황금 가면, 심하게 훼손된 화상 피부 및 젖은 의상 등 명시된 모든 요소가 정확히 일치합니다.",
    "hard_violations": [],
    "physics": "범람한 물과 진흙 속에 깊이 주저앉아 있으며, 물 표면이 몸을 자연스럽게 감싸 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시되지 않은 인물의 등장과 물리 법칙을 무시한 배경 요소(떠 있는 의자)로 인해 핵심 규칙을 위반했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업 앵글 속 반파된 황금 가면, 화상 입은 피부, 젖은 질감과 배경 잔해까지 프롬프트의 요구사항을 정확하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 상단의 허공을 향하고 있습니다.",
        "built_space": "레퍼런스의 성당 건물과 물탱크 파이프가 배경에 배치되었으나, 우측에 프롬프트에 없는 인물의 다리가 자리 잡고 있습니다.",
        "entities": "백산의 화상 얼굴, 깨진 황금 가면, 복장이 나타나지만 프레임 우측에 검은 부츠를 신은 정체불명의 인물이 추가되었습니다.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 추가 인물 등장 (우측 검은 부츠)",
         "배경 우측에 아무런 지지대 없이 허공에 떠 있는 의자"
        ],
        "physics": "우측 배경의 나무 의자는 공중에 떠 있으며 어떤 지지대도 없습니다. 백산의 자세는 바닥에 앉은 형태를 취하고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선은 정면 아래쪽을 공허하게 응시하고 있습니다.",
        "built_space": "로케이션 레퍼런스의 무너진 물탱크 철골과 파이프 잔해가 노을빛을 받는 물웅덩이와 함께 배치되었습니다.",
        "entities": "백산의 얼굴 형태, 절반이 깨진 황금 가면, 심하게 훼손된 화상 피부 및 젖은 의상 등 명시된 모든 요소가 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "범람한 물과 진흙 속에 깊이 주저앉아 있으며, 물 표면이 몸을 자연스럽게 감싸 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시되지 않은 인물의 등장과 물리 법칙을 무시한 배경 요소(떠 있는 의자)로 인해 핵심 규칙을 위반했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업 앵글 속 반파된 황금 가면, 화상 입은 피부, 젖은 질감과 배경 잔해까지 프롬프트의 요구사항을 정확하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 상단의 허공을 향하고 있습니다.",
        "built_space": "레퍼런스의 성당 건물과 물탱크 파이프가 배경에 배치되었으나, 우측에 프롬프트에 없는 인물의 다리가 자리 잡고 있습니다.",
        "entities": "백산의 화상 얼굴, 깨진 황금 가면, 복장이 나타나지만 프레임 우측에 검은 부츠를 신은 정체불명의 인물이 추가되었습니다.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 추가 인물 등장 (우측 검은 부츠)",
         "배경 우측에 아무런 지지대 없이 허공에 떠 있는 의자"
        ],
        "physics": "우측 배경의 나무 의자는 공중에 떠 있으며 어떤 지지대도 없습니다. 백산의 자세는 바닥에 앉은 형태를 취하고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선은 정면 아래쪽을 공허하게 응시하고 있습니다.",
        "built_space": "로케이션 레퍼런스의 무너진 물탱크 철골과 파이프 잔해가 노을빛을 받는 물웅덩이와 함께 배치되었습니다.",
        "entities": "백산의 얼굴 형태, 절반이 깨진 황금 가면, 심하게 훼손된 화상 피부 및 젖은 의상 등 명시된 모든 요소가 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "범람한 물과 진흙 속에 깊이 주저앉아 있으며, 물 표면이 몸을 자연스럽게 감싸 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "화상 얼굴과 반쯤 깨진 금빛 가면의 클로즈업은 충실하지만, 어깨까지 잠긴 수위는 젖은 바닥에 주저앉은 장면과 어긋나며 가면도 하회탈보다 장식적인 형태다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "장소와 젖은 왕실풍 의상은 잘 맞지만, 오른쪽에 다른 인물의 하체를 추가해 백산만 등장해야 한다는 필수 조건을 위반했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "노출된 눈은 거의 정면에서 약간 아래를 향하며 뚜렷한 응시 대상은 보이지 않는다. 얼굴은 정면에 가깝고 가면의 바깥 면과 파단면이 함께 드러난다. 무기나 이동 중인 인물은 없다.",
        "built_space": "배경에 무너진 금속 지지대와 굵은 배관 여러 개, 콘크리트 잔해가 보인다. 가까운 얼굴 중심의 구도라 참고 장소의 성당과 관람석은 확인할 수 없으며, 이 제외 자체는 문제가 아니다. 다만 참고의 얕게 물이 흐르는 바닥과 달리 백산의 양어깨 아래가 물에 잠겨 있어 깊은 웅덩이처럼 보인다.",
        "entities": "검은 머리의 성인 동아시아계 남성 한 명으로, 백산의 기본 외형과 부합한다. 드러난 이마와 뺨, 턱에 심한 화상 변형이 있고 눈은 자연스러운 사람 눈이다. 얼굴 절반에 남은 금빛 가면은 금속 질감과 깨진 가장자리가 명확하지만 소용돌이 장식이 강해 하회탈 특유의 형태와는 차이가 있다. 보이는 어깨에는 참고와 유사한 자주색·금색 왕실풍 의상이 남아 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되고 가면은 얼굴에 착용된 상태로 보인다. 몸통과 하체는 물에 가려져 엉덩이의 바닥 접촉이나 앉은 자세를 확인할 수 없다. 공중에 떠 있는 신체나 물체는 보이지 않지만, 어깨 높이까지 잠긴 상태는 요청한 바닥 착석을 명확하게 전달하지 못한다."
       },
       {
        "label": "B",
        "direction": "백산은 고개와 두 눈을 화면 오른쪽 위로 돌려, 오른쪽에 서 있는 추가 인물 쪽을 올려다본다. 응시 방향은 그 인물의 위치와 맞지만 그 상대 자체가 요청되지 않았다. 금빛 가면의 외측과 깨진 경계는 얼굴 옆에서 함께 보인다.",
        "built_space": "왼쪽에 파손된 물탱크와 물이 쏟아지는 배관, 뒤쪽에 중앙 아치 출입구 하나와 계단, 출입구 양옆 등 두 개와 세로 배너 두 개, 문 앞 의자 하나가 보여 참고 장소와 잘 맞는다. 오른쪽에는 넘어진 의자가 추가되어 있다. 백산은 젖은 잔해 바닥 가까이에 낮게 위치하지만, 바로 뒤 오른쪽에는 검은 옷과 장화를 착용한 다른 인물이 서 있다.",
        "entities": "백산은 검은 젖은 머리의 성인 동아시아계 남성으로 보이고, 자주색 바탕의 금색 자수 의상도 참고의 왕실풍 복장과 부합한다. 노출된 얼굴 반쪽에는 울퉁불퉁한 화상 흉터가 있으며 자연스러운 눈동자가 보인다. 금빛 반쪽 가면은 깨진 금속으로 읽히지만 하회탈의 특징은 제한적이다. 오른쪽에는 백산과 별개인 인물의 하체가 명백히 추가되어 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "백산만 등장해야 하는 장면에 검은 옷과 장화를 신은 다른 인물의 하체를 추가했다."
        ],
        "physics": "백산의 상체는 낮은 위치에서 기울어 있고 목과 어깨의 연결은 자연스럽다. 엉덩이와 손은 프레임 밖이어서 직접적인 바닥 지지는 확인할 수 없지만, 보이는 자세에 부유의 징후는 없다. 추가 인물의 장화는 바닥을 딛고 있으며, 가면은 얼굴에 밀착되어 있다. 배관의 물은 아래로 쏟아지고 가면과 머리카락의 물방울도 중력 방향으로 매달려 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "화상 얼굴과 반쯤 깨진 금빛 가면의 클로즈업은 충실하지만, 어깨까지 잠긴 수위는 젖은 바닥에 주저앉은 장면과 어긋나며 가면도 하회탈보다 장식적인 형태다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장소와 젖은 왕실풍 의상은 잘 맞지만, 오른쪽에 다른 인물의 하체를 추가해 백산만 등장해야 한다는 필수 조건을 위반했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "노출된 눈은 거의 정면에서 약간 아래를 향하며 뚜렷한 응시 대상은 보이지 않는다. 얼굴은 정면에 가깝고 가면의 바깥 면과 파단면이 함께 드러난다. 무기나 이동 중인 인물은 없다.",
        "built_space": "배경에 무너진 금속 지지대와 굵은 배관 여러 개, 콘크리트 잔해가 보인다. 가까운 얼굴 중심의 구도라 참고 장소의 성당과 관람석은 확인할 수 없으며, 이 제외 자체는 문제가 아니다. 다만 참고의 얕게 물이 흐르는 바닥과 달리 백산의 양어깨 아래가 물에 잠겨 있어 깊은 웅덩이처럼 보인다.",
        "entities": "검은 머리의 성인 동아시아계 남성 한 명으로, 백산의 기본 외형과 부합한다. 드러난 이마와 뺨, 턱에 심한 화상 변형이 있고 눈은 자연스러운 사람 눈이다. 얼굴 절반에 남은 금빛 가면은 금속 질감과 깨진 가장자리가 명확하지만 소용돌이 장식이 강해 하회탈 특유의 형태와는 차이가 있다. 보이는 어깨에는 참고와 유사한 자주색·금색 왕실풍 의상이 남아 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되고 가면은 얼굴에 착용된 상태로 보인다. 몸통과 하체는 물에 가려져 엉덩이의 바닥 접촉이나 앉은 자세를 확인할 수 없다. 공중에 떠 있는 신체나 물체는 보이지 않지만, 어깨 높이까지 잠긴 상태는 요청한 바닥 착석을 명확하게 전달하지 못한다."
       },
       {
        "label": "A",
        "direction": "백산은 고개와 두 눈을 화면 오른쪽 위로 돌려, 오른쪽에 서 있는 추가 인물 쪽을 올려다본다. 응시 방향은 그 인물의 위치와 맞지만 그 상대 자체가 요청되지 않았다. 금빛 가면의 외측과 깨진 경계는 얼굴 옆에서 함께 보인다.",
        "built_space": "왼쪽에 파손된 물탱크와 물이 쏟아지는 배관, 뒤쪽에 중앙 아치 출입구 하나와 계단, 출입구 양옆 등 두 개와 세로 배너 두 개, 문 앞 의자 하나가 보여 참고 장소와 잘 맞는다. 오른쪽에는 넘어진 의자가 추가되어 있다. 백산은 젖은 잔해 바닥 가까이에 낮게 위치하지만, 바로 뒤 오른쪽에는 검은 옷과 장화를 착용한 다른 인물이 서 있다.",
        "entities": "백산은 검은 젖은 머리의 성인 동아시아계 남성으로 보이고, 자주색 바탕의 금색 자수 의상도 참고의 왕실풍 복장과 부합한다. 노출된 얼굴 반쪽에는 울퉁불퉁한 화상 흉터가 있으며 자연스러운 눈동자가 보인다. 금빛 반쪽 가면은 깨진 금속으로 읽히지만 하회탈의 특징은 제한적이다. 오른쪽에는 백산과 별개인 인물의 하체가 명백히 추가되어 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "백산만 등장해야 하는 장면에 검은 옷과 장화를 신은 다른 인물의 하체를 추가했다."
        ],
        "physics": "백산의 상체는 낮은 위치에서 기울어 있고 목과 어깨의 연결은 자연스럽다. 엉덩이와 손은 프레임 밖이어서 직접적인 바닥 지지는 확인할 수 없지만, 보이는 자세에 부유의 징후는 없다. 추가 인물의 장화는 바닥을 딛고 있으며, 가면은 얼굴에 밀착되어 있다. 배관의 물은 아래로 쏟아지고 가면과 머리카락의 물방울도 중력 방향으로 매달려 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.929,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.679,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 지시되지 않은 추가 인물 등장 (우측 검은 부츠)",
     "[gemini-pro] 배경 우측에 아무런 지지대 없이 허공에 떠 있는 의자",
     "[gpt-high] 백산만 등장해야 하는 장면에 검은 옷과 장화를 신은 다른 인물의 하체를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 679,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 679,
    "verdict_ko": "지시되지 않은 인물의 등장과 물리 법칙을 무시한 배경 요소(떠 있는 의자)로 인해 핵심 규칙을 위반했습니다.  ★위반: [gemini-pro] 프롬프트에 지시되지 않은 추가 인물 등장 (우측 검은 부츠) / [gemini-pro] 배경 우측에 아무런 지지대 없이 허공에 떠 있는 의자 / [gpt-high] 백산만 등장해야 하는 장면에 검은 옷과 장화를 신은 다른 인물의 하체를 추가했다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "클로즈업 앵글 속 반파된 황금 가면, 화상 입은 피부, 젖은 질감과 배경 잔해까지 프롬프트의 요구사항을 정확하게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B03.png",
    "asset_id": "1f832f0c-4b90-4195-8434-227eb2c715f6",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1083564>",
    "asset_id": "9a307034-865c-47d1-bf2c-6230188e67c7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-35d5-711d-a82b-cf96a47ad6f1",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh52__bgfirst_bg.png",
   "bg_asset_id": "aeb46e86-9443-4079-ae58-e593ffa40546",
   "bg_record_key": "S60sh52::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S60sh52::cine": {
  "applied": false,
  "attempted_at": "2026-09-19T18:50:20.842137+00:00",
  "fingerprint": "2fb5d30e17712eb4dbf528fa0d8202b45c178e9a345e947fb2f7f1546f9e5f62",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S60sh52_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S60sh52_sel.png",
  "source_sha256": "379fe86ed947aad1b47cd62cd811540512e0dd2bb758f344743f49dded9f4e1e",
  "error": "RuntimeError: moderation blocked: {\"code\":\"imagine:content-moderated\",\"error\":\"Generated image rejected by content moderation.\",\"usage\":{\"cost_in_usd_ticks\":220000000}}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S60sh63::signage": {
  "fp": "5b70c7e5af8ef902",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S60sh63::bgfirst_bg": {
  "input_fingerprint": "1da2083e740ebd82",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh63__bgfirst_bg.png",
  "asset_id": "1a5d48e7-71af-4630-8d5b-960f4796ee43",
  "input_asset_ids": [
   "3f4ba729-6aad-496e-a709-774fcddbcb83",
   "5382b8c1-3827-43d2-bb92-334028ead33d"
  ]
 },
 "S60sh63": {
  "input_fingerprint": "65f5c87a771224a5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Sunset lights the ridge above the devastated village. Charlie remains badly dented and punctured, with malfunctioning sensors, and extends a metal hand. 현우: He sits on the ridge, visibly battered from the fighting, with a brighter expression and one hand extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Sunset lights the ridge above the devastated village. Charlie remains badly dented and punctured, with malfunctioning sensors, and extends a metal hand. 현우: He sits on the ridge, visibly battered from the fighting, with a brighter expression and one hand extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Sunset lights the ridge above the devastated village. Charlie remains badly dented and punctured, with malfunctioning sensors, and extends a metal hand. 현우: He sits on the ridge, visibly battered from the fighting, with a brighter expression and one hand extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh63__bgfirst_bg.png",
     "asset_id": "1a5d48e7-71af-4630-8d5b-960f4796ee43",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S60sh63.png",
     "asset_id": "3f4ba729-6aad-496e-a709-774fcddbcb83",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B02.png",
     "asset_id": "5382b8c1-3827-43d2-bb92-334028ead33d",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우(좌)와 찰리(우)가 손을 뻗어 중앙에서 맞잡고 있음.",
    "built_space": "능선 지면 위. 배경은 노을 진 산맥이며 파괴된 마을은 보이지 않음.",
    "entities": "현우는 회색 소매와 상처 난 팔, 찰리는 모래색 금속 장갑판으로 표현됨.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조 (현우와 찰리의 손가락이 비정상적으로 융합됨)"
    ],
    "physics": "두 팔이 공중에서 서로를 지탱하나, 맞잡은 손의 형태가 해부학적으로 성립하지 않음."
   },
   {
    "label": "B",
    "direction": "찰리(좌)와 현우(우)가 손을 뻗어 단단히 맞잡고 있음.",
    "built_space": "능선 지면 위. 팔 아래로 지면이 보이며 배경에 파괴된 마을이 위치함.",
    "entities": "찰리는 마모된 모래색 장갑판, 현우는 구멍 난 회색 소매와 오염된 피부로 표현됨.",
    "hard_violations": [],
    "physics": "찰리의 왼손과 현우의 오른손이 물리적으로 자연스럽게 교차하여 단단히 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "맞잡은 손의 해부학적 구조가 불가능하게 융합되어 프롬프트 조건을 심각하게 위반함."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "요구된 클로즈업 구도, 자연스러운 손의 결합, 파괴된 마을 배경을 모두 충실히 구현함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우(좌)와 찰리(우)가 손을 뻗어 중앙에서 맞잡고 있음.",
        "built_space": "능선 지면 위. 배경은 노을 진 산맥이며 파괴된 마을은 보이지 않음.",
        "entities": "현우는 회색 소매와 상처 난 팔, 찰리는 모래색 금속 장갑판으로 표현됨.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조 (현우와 찰리의 손가락이 비정상적으로 융합됨)"
        ],
        "physics": "두 팔이 공중에서 서로를 지탱하나, 맞잡은 손의 형태가 해부학적으로 성립하지 않음."
       },
       {
        "label": "B",
        "direction": "찰리(좌)와 현우(우)가 손을 뻗어 단단히 맞잡고 있음.",
        "built_space": "능선 지면 위. 팔 아래로 지면이 보이며 배경에 파괴된 마을이 위치함.",
        "entities": "찰리는 마모된 모래색 장갑판, 현우는 구멍 난 회색 소매와 오염된 피부로 표현됨.",
        "hard_violations": [],
        "physics": "찰리의 왼손과 현우의 오른손이 물리적으로 자연스럽게 교차하여 단단히 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "맞잡은 손의 해부학적 구조가 불가능하게 융합되어 프롬프트 조건을 심각하게 위반함."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "요구된 클로즈업 구도, 자연스러운 손의 결합, 파괴된 마을 배경을 모두 충실히 구현함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우(좌)와 찰리(우)가 손을 뻗어 중앙에서 맞잡고 있음.",
        "built_space": "능선 지면 위. 배경은 노을 진 산맥이며 파괴된 마을은 보이지 않음.",
        "entities": "현우는 회색 소매와 상처 난 팔, 찰리는 모래색 금속 장갑판으로 표현됨.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조 (현우와 찰리의 손가락이 비정상적으로 융합됨)"
        ],
        "physics": "두 팔이 공중에서 서로를 지탱하나, 맞잡은 손의 형태가 해부학적으로 성립하지 않음."
       },
       {
        "label": "B",
        "direction": "찰리(좌)와 현우(우)가 손을 뻗어 단단히 맞잡고 있음.",
        "built_space": "능선 지면 위. 팔 아래로 지면이 보이며 배경에 파괴된 마을이 위치함.",
        "entities": "찰리는 마모된 모래색 장갑판, 현우는 구멍 난 회색 소매와 오염된 피부로 표현됨.",
        "hard_violations": [],
        "physics": "찰리의 왼손과 현우의 오른손이 물리적으로 자연스럽게 교차하여 단단히 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "서로 감아 쥔 손과 현우의 거친 피부, 찰리의 손상된 금속이 핵심 순간을 더 분명히 구현하지만, 하늘과 원경까지 드러내 배경을 제한한 클로즈업 지시는 어긴다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "노을 아래 인간 손과 육중한 금속 손의 접촉은 구현했으나, 찰리가 되잡는 동작이 덜 분명하고 넓은 지면과 마을이 손 중심의 제한된 구도를 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔은 왼쪽에서 오른쪽 중앙으로, 현우의 팔은 오른쪽 위에서 왼쪽 아래로 뻗어 두 손이 만난다. 현우의 엄지와 아래쪽 손가락들이 금속 손을 붙잡지만, 찰리의 긴 손가락은 주로 오른쪽 아래로 나란히 향해 현우의 손을 되감아 쥐는 모습이 덜 명확하다. 얼굴과 시선은 화면 밖이다.",
        "built_space": "자갈과 마른 풀이 있는 능선 지면이 화면 하단에 넓게 보이고, 두 팔 위로 여러 파손 건물과 산이 드러난다. 능선의 재질은 장소 참조와 부합하지만, 지면을 팔 아래 작은 틈으로만 보여 달라는 제한과 다르다. 좌석이나 고정 설비, 반사는 보이지 않으며 골반이 잘려 있어 앉은 자세는 확인할 수 없다.",
        "entities": "현우에 해당하는 사람 손 하나와 회색 셔츠 소매, 찰리에 해당하는 모래색 장갑의 금속 손과 전완 하나가 보인다. 사람 손은 때와 피부 주름이 있으나 전투 상처는 두드러지지 않는다. 손만으로 18세 한국계 미국인이라는 신원은 확정할 수 없으며 얼굴과 머리는 평가 범위 밖이다. 찰리의 각진 장갑, 검은 관절, 긁힘과 작은 파손 구멍은 참조 및 손상 상태와 잘 맞는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손은 각각 화면 밖으로 이어지는 손목과 팔에 연결되어 지지된다. 현우의 엄지가 금속 손 위를 누르고 손가락들이 아래를 받쳐 실제 접촉이 성립한다. 금속 손의 맞쥠은 약하게 읽히지만 관절이 불가능하게 꺾였다고 단정할 근거는 없다. 몸통과 하체는 잘려 있어 착석 지지점은 확인되지 않으며, 이를 공중 부양으로 볼 근거도 없다."
       },
       {
        "label": "B",
        "direction": "현우의 팔이 왼쪽에서 오른쪽 중앙으로, 찰리의 팔이 오른쪽에서 왼쪽 중앙으로 뻗는다. 찰리의 엄지가 현우의 손등을 덮고 아래쪽 금속 손가락들이 사람 손을 감싸며, 현우의 손가락도 반대쪽을 붙잡아 서로 단단히 맞잡는 방향과 목표가 분명하다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "팔 아래에는 돌과 마른 풀이 있는 능선이 보이고, 팔 위에는 골짜기의 작은 마을과 연기, 겹친 산줄기, 넓은 노을 하늘이 드러난다. 지형과 석양은 장소 참조에 부합하지만 배경이 팔 아래 작은 지면 틈에 한정되지 않는다. 왼쪽에 현우의 옷 입은 몸 일부가 보일 뿐 엉덩이와 지면의 접촉은 잘려 있다. 고정 설비나 반사는 없다.",
        "entities": "현우의 회색 셔츠와 올리브색 하의 일부, 상처와 흙먼지가 묻은 손과 전완이 보인다. 얼굴이 없어 정확한 나이와 한국계 미국인 신원은 확인할 수 없지만 드러난 신체가 청년 설정과 명백히 충돌하지는 않는다. 찰리의 육중한 모래색 금속 손과 팔에는 긁힘, 패임, 찢어진 장갑 가장자리가 있어 전투 손상 상태가 드러난다. 얼굴과 센서는 화면 밖이므로 평가하지 않는다. 추가 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "사람 손은 손목과 전완으로, 금속 손은 기계식 손목과 장갑 팔로 연속해서 연결된다. 양쪽 엄지와 굽힌 손가락이 상대 손에 접촉해 서로 잡는 힘의 방향이 자연스럽다. 지지 없이 떠 있는 물체나 분리된 신체는 보이지 않는다. 착석 여부를 확정할 하체 지지점은 프레임 밖이다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "서로 감아 쥔 손과 현우의 거친 피부, 찰리의 손상된 금속이 핵심 순간을 더 분명히 구현하지만, 하늘과 원경까지 드러내 배경을 제한한 클로즈업 지시는 어긴다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "노을 아래 인간 손과 육중한 금속 손의 접촉은 구현했으나, 찰리가 되잡는 동작이 덜 분명하고 넓은 지면과 마을이 손 중심의 제한된 구도를 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 팔은 왼쪽에서 오른쪽 중앙으로, 현우의 팔은 오른쪽 위에서 왼쪽 아래로 뻗어 두 손이 만난다. 현우의 엄지와 아래쪽 손가락들이 금속 손을 붙잡지만, 찰리의 긴 손가락은 주로 오른쪽 아래로 나란히 향해 현우의 손을 되감아 쥐는 모습이 덜 명확하다. 얼굴과 시선은 화면 밖이다.",
        "built_space": "자갈과 마른 풀이 있는 능선 지면이 화면 하단에 넓게 보이고, 두 팔 위로 여러 파손 건물과 산이 드러난다. 능선의 재질은 장소 참조와 부합하지만, 지면을 팔 아래 작은 틈으로만 보여 달라는 제한과 다르다. 좌석이나 고정 설비, 반사는 보이지 않으며 골반이 잘려 있어 앉은 자세는 확인할 수 없다.",
        "entities": "현우에 해당하는 사람 손 하나와 회색 셔츠 소매, 찰리에 해당하는 모래색 장갑의 금속 손과 전완 하나가 보인다. 사람 손은 때와 피부 주름이 있으나 전투 상처는 두드러지지 않는다. 손만으로 18세 한국계 미국인이라는 신원은 확정할 수 없으며 얼굴과 머리는 평가 범위 밖이다. 찰리의 각진 장갑, 검은 관절, 긁힘과 작은 파손 구멍은 참조 및 손상 상태와 잘 맞는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손은 각각 화면 밖으로 이어지는 손목과 팔에 연결되어 지지된다. 현우의 엄지가 금속 손 위를 누르고 손가락들이 아래를 받쳐 실제 접촉이 성립한다. 금속 손의 맞쥠은 약하게 읽히지만 관절이 불가능하게 꺾였다고 단정할 근거는 없다. 몸통과 하체는 잘려 있어 착석 지지점은 확인되지 않으며, 이를 공중 부양으로 볼 근거도 없다."
       },
       {
        "label": "A",
        "direction": "현우의 팔이 왼쪽에서 오른쪽 중앙으로, 찰리의 팔이 오른쪽에서 왼쪽 중앙으로 뻗는다. 찰리의 엄지가 현우의 손등을 덮고 아래쪽 금속 손가락들이 사람 손을 감싸며, 현우의 손가락도 반대쪽을 붙잡아 서로 단단히 맞잡는 방향과 목표가 분명하다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "팔 아래에는 돌과 마른 풀이 있는 능선이 보이고, 팔 위에는 골짜기의 작은 마을과 연기, 겹친 산줄기, 넓은 노을 하늘이 드러난다. 지형과 석양은 장소 참조에 부합하지만 배경이 팔 아래 작은 지면 틈에 한정되지 않는다. 왼쪽에 현우의 옷 입은 몸 일부가 보일 뿐 엉덩이와 지면의 접촉은 잘려 있다. 고정 설비나 반사는 없다.",
        "entities": "현우의 회색 셔츠와 올리브색 하의 일부, 상처와 흙먼지가 묻은 손과 전완이 보인다. 얼굴이 없어 정확한 나이와 한국계 미국인 신원은 확인할 수 없지만 드러난 신체가 청년 설정과 명백히 충돌하지는 않는다. 찰리의 육중한 모래색 금속 손과 팔에는 긁힘, 패임, 찢어진 장갑 가장자리가 있어 전투 손상 상태가 드러난다. 얼굴과 센서는 화면 밖이므로 평가하지 않는다. 추가 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "사람 손은 손목과 전완으로, 금속 손은 기계식 손목과 장갑 팔로 연속해서 연결된다. 양쪽 엄지와 굽힌 손가락이 상대 손에 접촉해 서로 잡는 힘의 방향이 자연스럽다. 지지 없이 떠 있는 물체나 분리된 신체는 보이지 않는다. 착석 여부를 확정할 하체 지지점은 프레임 밖이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.857
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.857
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조 (현우와 찰리의 손가락이 비정상적으로 융합됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1857
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "맞잡은 손의 해부학적 구조가 불가능하게 융합되어 프롬프트 조건을 심각하게 위반함.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학적 구조 (현우와 찰리의 손가락이 비정상적으로 융합됨)"
   },
   {
    "label": "B",
    "score": 1857,
    "verdict_ko": "요구된 클로즈업 구도, 자연스러운 손의 결합, 파괴된 마을 배경을 모두 충실히 구현함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B02.png",
    "asset_id": "5382b8c1-3827-43d2-bb92-334028ead33d",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-393f-7822-a4a4-33b02df1f05b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh63__bgfirst_bg.png",
   "bg_asset_id": "1a5d48e7-71af-4630-8d5b-960f4796ee43",
   "bg_record_key": "S60sh63::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S60sh63::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:55:07.421160+00:00",
  "fingerprint": "7312bb593d2464c87a021fda5b43247aa864e9e1b632e7f078eba9578c9d69be",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S60sh63_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S60sh63_sel.png",
  "source_sha256": "6bd8f4a81e149a0e3afcdfcb9f9db27ad321bb59f58981313fe0c8f5b8cb4092",
  "file": "S60sh63_cine.png",
  "staged_sha256": "05118a78bfbc57fef77e4d0c0ace27857e86c97bc7fa5228bb2af274422668e2",
  "latency_ms": 10723
 },
 "S61sh1::signage": {
  "fp": "3d752f0da3adcbcb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S61sh1": {
  "input_fingerprint": "f95149bc39268e4c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어두운 밤, 늪지대 진흙 속에 박힌 낡은 캠핑카의 외관 전경.\n\nLOCATION (lock): At the marsh's muddy vehicle-stranding point at night, where the abandoned camper remains sunk in the ground. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 낡은 캠핑카 (Old and lodged in the swamp mud) — The near end, adjacent side, and part of the upper body are visible from the elevated diagonal viewpoint; used as Establishes the discovered vehicle at a readable environmental scale; 늪지대 진흙 (Surrounding and holding the vehicle's wheels); used as Occupies the foreground and side margins, explaining why the vehicle remains here.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the scene's dark nighttime ambience with restrained tonal separation sufficient to distinguish the vehicle from the swamp, without adding an unestablished light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stuck in the swamp at night, with no successful extraction established. A rusty radiation-zone sign stands nearby, with crows gathered around it.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어두운 밤, 늪지대 진흙 속에 박힌 낡은 캠핑카의 외관 전경.\n\nLOCATION (lock): At the marsh's muddy vehicle-stranding point at night, where the abandoned camper remains sunk in the ground. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 낡은 캠핑카 (Old and lodged in the swamp mud) — The near end, adjacent side, and part of the upper body are visible from the elevated diagonal viewpoint; used as Establishes the discovered vehicle at a readable environmental scale; 늪지대 진흙 (Surrounding and holding the vehicle's wheels); used as Occupies the foreground and side margins, explaining why the vehicle remains here.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the scene's dark nighttime ambience with restrained tonal separation sufficient to distinguish the vehicle from the swamp, without adding an unestablished light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stuck in the swamp at night, with no successful extraction established. A rusty radiation-zone sign stands nearby, with crows gathered around it.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어두운 밤, 늪지대 진흙 속에 박힌 낡은 캠핑카의 외관 전경.\n\nLOCATION (lock): At the marsh's muddy vehicle-stranding point at night, where the abandoned camper remains sunk in the ground. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 낡은 캠핑카 (Old and lodged in the swamp mud) — The near end, adjacent side, and part of the upper body are visible from the elevated diagonal viewpoint; used as Establishes the discovered vehicle at a readable environmental scale; 늪지대 진흙 (Surrounding and holding the vehicle's wheels); used as Occupies the foreground and side margins, explaining why the vehicle remains here.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the scene's dark nighttime ambience with restrained tonal separation sufficient to distinguish the vehicle from the swamp, without adding an unestablished light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stuck in the swamp at night, with no successful extraction established. A rusty radiation-zone sign stands nearby, with crows gathered around it.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 늪지대와 캠핑카, 표지판을 향해 하향식(사선)으로 향함. 까마귀들은 표지판과 바닥에서 다양한 방향을 바라봄.",
    "built_space": "진흙 늪지대 좌측에 캠핑카가 박혀 있고, 우측에 둥근 표지판과 방사능 표지판이 있음. 레퍼런스와 동일하게 캠핑카의 전면부와 측면이 자연스럽게 배치됨.",
    "entities": "낡은 캠핑카, 녹슨 진입 금지 및 방사능 표지판, 까마귀 무리, 진흙. 인물은 없으며 야간 조명이 적용됨.",
    "hard_violations": [],
    "physics": "캠핑카는 진흙 속에 단단히 박혀 지탱됨. 표지판은 땅에 고정되어 있고 까마귀들은 표지판 위와 바닥에 안정적으로 서 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 늪지대 전경과 거대한 발자국, 캠핑카를 향함. 캠핑카는 왼쪽(후면 방향)을 향하고 있음.",
    "built_space": "늪지대 좌측에 캠핑카, 우측 원경에 표지판이 있음. 전경의 발자국과 배경 산세는 레퍼런스와 동일한데, 캠핑카만 방향이 180도 뒤집혀 후면이 노출됨.",
    "entities": "방향이 뒤집힌 낡은 캠핑카, 표지판, 거대한 발자국, 까마귀 무리, 진흙. 인물은 없으며 밤 시간대임.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스와 동일한 전경(발자국) 및 배경을 유지한 채 카메라가 이동하지 않았음에도 캠핑카의 방향만 뒤집혀(전면->후면) 물리적 공간의 연속성 오류 발생",
     "[gemini-pro] 지시된 'elevated diagonal viewpoint' (높은 사선 시점)을 위반하고 레퍼런스의 낮은 시점을 그대로 답습함",
     "[gpt-high] 장소 고정 참조에 없는 전신주 두 개와 연결 전선을 배경에 새로 추가했습니다."
    ],
    "physics": "캠핑카는 진흙에 박혀 있고 발자국은 진흙에 깊게 파여 형태를 유지함. 표지판과 까마귀들도 안정적으로 위치함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 공간적 일관성을 유지하면서 야간 조명, 인물 제거, 요구된 높은 사선 시점(elevated diagonal)을 모두 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "요구된 높은 시점을 반영하지 않았고, 전경과 배경이 동일함에도 캠핑카의 방향만 임의로 뒤집히는 연속성 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 늪지대와 캠핑카, 표지판을 향해 하향식(사선)으로 향함. 까마귀들은 표지판과 바닥에서 다양한 방향을 바라봄.",
        "built_space": "진흙 늪지대 좌측에 캠핑카가 박혀 있고, 우측에 둥근 표지판과 방사능 표지판이 있음. 레퍼런스와 동일하게 캠핑카의 전면부와 측면이 자연스럽게 배치됨.",
        "entities": "낡은 캠핑카, 녹슨 진입 금지 및 방사능 표지판, 까마귀 무리, 진흙. 인물은 없으며 야간 조명이 적용됨.",
        "hard_violations": [],
        "physics": "캠핑카는 진흙 속에 단단히 박혀 지탱됨. 표지판은 땅에 고정되어 있고 까마귀들은 표지판 위와 바닥에 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 늪지대 전경과 거대한 발자국, 캠핑카를 향함. 캠핑카는 왼쪽(후면 방향)을 향하고 있음.",
        "built_space": "늪지대 좌측에 캠핑카, 우측 원경에 표지판이 있음. 전경의 발자국과 배경 산세는 레퍼런스와 동일한데, 캠핑카만 방향이 180도 뒤집혀 후면이 노출됨.",
        "entities": "방향이 뒤집힌 낡은 캠핑카, 표지판, 거대한 발자국, 까마귀 무리, 진흙. 인물은 없으며 밤 시간대임.",
        "hard_violations": [
         "레퍼런스와 동일한 전경(발자국) 및 배경을 유지한 채 카메라가 이동하지 않았음에도 캠핑카의 방향만 뒤집혀(전면->후면) 물리적 공간의 연속성 오류 발생",
         "지시된 'elevated diagonal viewpoint' (높은 사선 시점)을 위반하고 레퍼런스의 낮은 시점을 그대로 답습함"
        ],
        "physics": "캠핑카는 진흙에 박혀 있고 발자국은 진흙에 깊게 파여 형태를 유지함. 표지판과 까마귀들도 안정적으로 위치함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 공간적 일관성을 유지하면서 야간 조명, 인물 제거, 요구된 높은 사선 시점(elevated diagonal)을 모두 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "요구된 높은 시점을 반영하지 않았고, 전경과 배경이 동일함에도 캠핑카의 방향만 임의로 뒤집히는 연속성 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 늪지대와 캠핑카, 표지판을 향해 하향식(사선)으로 향함. 까마귀들은 표지판과 바닥에서 다양한 방향을 바라봄.",
        "built_space": "진흙 늪지대 좌측에 캠핑카가 박혀 있고, 우측에 둥근 표지판과 방사능 표지판이 있음. 레퍼런스와 동일하게 캠핑카의 전면부와 측면이 자연스럽게 배치됨.",
        "entities": "낡은 캠핑카, 녹슨 진입 금지 및 방사능 표지판, 까마귀 무리, 진흙. 인물은 없으며 야간 조명이 적용됨.",
        "hard_violations": [],
        "physics": "캠핑카는 진흙 속에 단단히 박혀 지탱됨. 표지판은 땅에 고정되어 있고 까마귀들은 표지판 위와 바닥에 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 늪지대 전경과 거대한 발자국, 캠핑카를 향함. 캠핑카는 왼쪽(후면 방향)을 향하고 있음.",
        "built_space": "늪지대 좌측에 캠핑카, 우측 원경에 표지판이 있음. 전경의 발자국과 배경 산세는 레퍼런스와 동일한데, 캠핑카만 방향이 180도 뒤집혀 후면이 노출됨.",
        "entities": "방향이 뒤집힌 낡은 캠핑카, 표지판, 거대한 발자국, 까마귀 무리, 진흙. 인물은 없으며 밤 시간대임.",
        "hard_violations": [
         "레퍼런스와 동일한 전경(발자국) 및 배경을 유지한 채 카메라가 이동하지 않았음에도 캠핑카의 방향만 뒤집혀(전면->후면) 물리적 공간의 연속성 오류 발생",
         "지시된 'elevated diagonal viewpoint' (높은 사선 시점)을 위반하고 레퍼런스의 낮은 시점을 그대로 답습함"
        ],
        "physics": "캠핑카는 진흙에 박혀 있고 발자국은 진흙에 깊게 파여 형태를 유지함. 표지판과 까마귀들도 안정적으로 위치함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "높은 사선 시점과 진흙에 박힌 차량은 구현했지만, 참조에 없는 전신주와 전선을 추가해 장소 고정을 위반했습니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "참조 차량의 전면·측면과 늪에 고착된 상태를 달밤의 전경으로 충실히 재현했지만, 표지판을 전경에 크게 부각한 점은 배경 요소의 크기 지시와 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카의 뒤쪽 끝이 화면 왼쪽 아래를 향하고 앞쪽은 오른쪽 뒤로 향해, 후면·인접 측면·지붕이 보입니다. 오른쪽 두 표지판의 표시 면은 카메라 쪽을 향합니다. 까마귀들은 표지판 위와 주변 지면에서 서로 다른 방향을 보고 있으며, 지정된 응시 대상이나 이동 동작은 없습니다.",
        "built_space": "차량 한 대의 측면에는 문 하나와 큰 창 두 개가 보이고, 후면에는 창 하나와 사다리 하나, 지붕에는 냉방 장치 하나와 난간이 있습니다. 오른쪽에는 원형 표지판 하나와 사각 방사능 표지판 하나가 있어 참조의 표지판 수와 맞습니다. 진흙과 웅덩이가 전경과 양옆을 채우지만, 먼 배경에 참조에서 확인되지 않는 전신주 두 개와 연결 전선이 추가되었습니다.",
        "entities": "사람이나 얼굴은 없습니다. 녹슨 흰 차체와 적갈색 띠를 가진 낡은 캠핑카, 진흙 늪, 마른 갈대, 방사능 표지판, 그 주변 까마귀 무리가 있습니다. 차량의 색과 노후 재질은 참조와 유사하지만, 후면 중심 시점이라 참조의 특징적인 전면은 확인할 수 없습니다. 전경에는 참조의 밑창 자국과 유사한 자국이 남아 있습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "장소 고정 참조에 없는 전신주 두 개와 연결 전선을 배경에 새로 추가했습니다."
        ],
        "physics": "차량의 보이는 바퀴 하부가 진흙에 묻혀 있고 차체는 바퀴와 진흙에 지지됩니다. 표지판 기둥은 땅에 박혀 있으며, 까마귀들은 표지판 윗부분이나 진흙 지면에 발을 대고 있습니다. 지붕 장비와 사다리도 차체에 부착되어 있고, 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "캠핑카 전면은 화면 오른쪽 아래를 향하고 긴 측면은 왼쪽 뒤로 이어져, 가까운 전면·인접 측면·지붕 일부가 함께 보입니다. 두 표지판은 표시 면을 카메라 쪽으로 내놓습니다. 표지판 위 까마귀들은 주로 왼쪽 또는 옆을 보고, 지면의 까마귀들은 여러 방향을 향합니다. 요구되지 않은 이동이나 조준 동작은 없습니다.",
        "built_space": "캠핑카 한 대에 전면 유리, 전조등 한 쌍, 측면 문 하나와 창 두 개, 지붕 냉방 장치 하나가 보여 참조 차량의 주요 구성이 이어집니다. 차량 오른쪽에는 원형 표지판 하나와 사각 방사능 표지판 하나가 각각 기둥에 달려 있습니다. 높은 사선 시점에서 늪과 매몰된 바퀴 주변이 읽히지만, 표지판들이 오른쪽 전경으로 크게 들어와 작은 배경 요소로 유지하라는 지시에는 덜 맞습니다.",
        "entities": "사람이나 얼굴은 없습니다. 흰색과 적갈색 띠의 녹슨 캠핑카는 참조의 전면 형태와 측면 배치를 잘 유지합니다. 진흙, 얕은 물, 마른 갈대, 원형 진입금지 표지판, 녹슨 사각 방사능 표지판과 까마귀 무리가 모두 확인됩니다. 구름 사이 달빛에 의한 밤으로 표현되며 별도의 인공 광원은 보이지 않습니다. 표지판에는 도형만 식별되고 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "앞뒤 바퀴 하부가 진흙에 깊이 묻혀 있어 차량이 빠져나오지 못한 상태가 분명합니다. 두 표지판의 기둥은 진흙에 박혀 있습니다. 표지판 위 세 마리의 까마귀는 발로 테두리를 딛고, 나머지는 지면에 서 있습니다. 지붕 장치는 차체에 고정되어 있으며, 떠 있는 몸이나 지지 없는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "높은 사선 시점과 진흙에 박힌 차량은 구현했지만, 참조에 없는 전신주와 전선을 추가해 장소 고정을 위반했습니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "참조 차량의 전면·측면과 늪에 고착된 상태를 달밤의 전경으로 충실히 재현했지만, 표지판을 전경에 크게 부각한 점은 배경 요소의 크기 지시와 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카의 뒤쪽 끝이 화면 왼쪽 아래를 향하고 앞쪽은 오른쪽 뒤로 향해, 후면·인접 측면·지붕이 보입니다. 오른쪽 두 표지판의 표시 면은 카메라 쪽을 향합니다. 까마귀들은 표지판 위와 주변 지면에서 서로 다른 방향을 보고 있으며, 지정된 응시 대상이나 이동 동작은 없습니다.",
        "built_space": "차량 한 대의 측면에는 문 하나와 큰 창 두 개가 보이고, 후면에는 창 하나와 사다리 하나, 지붕에는 냉방 장치 하나와 난간이 있습니다. 오른쪽에는 원형 표지판 하나와 사각 방사능 표지판 하나가 있어 참조의 표지판 수와 맞습니다. 진흙과 웅덩이가 전경과 양옆을 채우지만, 먼 배경에 참조에서 확인되지 않는 전신주 두 개와 연결 전선이 추가되었습니다.",
        "entities": "사람이나 얼굴은 없습니다. 녹슨 흰 차체와 적갈색 띠를 가진 낡은 캠핑카, 진흙 늪, 마른 갈대, 방사능 표지판, 그 주변 까마귀 무리가 있습니다. 차량의 색과 노후 재질은 참조와 유사하지만, 후면 중심 시점이라 참조의 특징적인 전면은 확인할 수 없습니다. 전경에는 참조의 밑창 자국과 유사한 자국이 남아 있습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "장소 고정 참조에 없는 전신주 두 개와 연결 전선을 배경에 새로 추가했습니다."
        ],
        "physics": "차량의 보이는 바퀴 하부가 진흙에 묻혀 있고 차체는 바퀴와 진흙에 지지됩니다. 표지판 기둥은 땅에 박혀 있으며, 까마귀들은 표지판 윗부분이나 진흙 지면에 발을 대고 있습니다. 지붕 장비와 사다리도 차체에 부착되어 있고, 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "캠핑카 전면은 화면 오른쪽 아래를 향하고 긴 측면은 왼쪽 뒤로 이어져, 가까운 전면·인접 측면·지붕 일부가 함께 보입니다. 두 표지판은 표시 면을 카메라 쪽으로 내놓습니다. 표지판 위 까마귀들은 주로 왼쪽 또는 옆을 보고, 지면의 까마귀들은 여러 방향을 향합니다. 요구되지 않은 이동이나 조준 동작은 없습니다.",
        "built_space": "캠핑카 한 대에 전면 유리, 전조등 한 쌍, 측면 문 하나와 창 두 개, 지붕 냉방 장치 하나가 보여 참조 차량의 주요 구성이 이어집니다. 차량 오른쪽에는 원형 표지판 하나와 사각 방사능 표지판 하나가 각각 기둥에 달려 있습니다. 높은 사선 시점에서 늪과 매몰된 바퀴 주변이 읽히지만, 표지판들이 오른쪽 전경으로 크게 들어와 작은 배경 요소로 유지하라는 지시에는 덜 맞습니다.",
        "entities": "사람이나 얼굴은 없습니다. 흰색과 적갈색 띠의 녹슨 캠핑카는 참조의 전면 형태와 측면 배치를 잘 유지합니다. 진흙, 얕은 물, 마른 갈대, 원형 진입금지 표지판, 녹슨 사각 방사능 표지판과 까마귀 무리가 모두 확인됩니다. 구름 사이 달빛에 의한 밤으로 표현되며 별도의 인공 광원은 보이지 않습니다. 표지판에는 도형만 식별되고 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "앞뒤 바퀴 하부가 진흙에 깊이 묻혀 있어 차량이 빠져나오지 못한 상태가 분명합니다. 두 표지판의 기둥은 진흙에 박혀 있습니다. 표지판 위 세 마리의 까마귀는 발로 테두리를 딛고, 나머지는 지면에 서 있습니다. 지붕 장치는 차체에 고정되어 있으며, 떠 있는 몸이나 지지 없는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 레퍼런스와 동일한 전경(발자국) 및 배경을 유지한 채 카메라가 이동하지 않았음에도 캠핑카의 방향만 뒤집혀(전면->후면) 물리적 공간의 연속성 오류 발생",
     "[gemini-pro] 지시된 'elevated diagonal viewpoint' (높은 사선 시점)을 위반하고 레퍼런스의 낮은 시점을 그대로 답습함",
     "[gpt-high] 장소 고정 참조에 없는 전신주 두 개와 연결 전선을 배경에 새로 추가했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스의 공간적 일관성을 유지하면서 야간 조명, 인물 제거, 요구된 높은 사선 시점(elevated diagonal)을 모두 충실히 구현했습니다."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "요구된 높은 시점을 반영하지 않았고, 전경과 배경이 동일함에도 캠핑카의 방향만 임의로 뒤집히는 연속성 오류가 발생했습니다.  ★위반: [gemini-pro] 레퍼런스와 동일한 전경(발자국) 및 배경을 유지한 채 카메라가 이동하지 않았음에도 캠핑카의 방향만 뒤집혀(전면->후면) 물리적 공간의 연속성 오류 발생 / [gemini-pro] 지시된 'elevated diagonal viewpoint' (높은 사선 시점)을 위반하고 레퍼런스의 낮은 시점을 그대로 답습함 / [gpt-high] 장소 고정 참조에 없는 전신주 두 개와 연결 전선을 배경에 새로 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5_sel.png",
    "asset_id": "d0364154-8f6d-4565-b8fd-94785cb7cc07",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-3ca0-7d88-bef6-1ca95f599f6b",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S56sh5"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S61sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:56:12.039781+00:00",
  "fingerprint": "ab78f6a4144f9e636f5516b30ddad4ba9df0212dd424966902bd464967535e9b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S61sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S61sh1_sel.png",
  "source_sha256": "ed9ff83c8e944b38ed1d0a39dcc6a7c96c7c5151489013bf7622975cc6c14dd6",
  "file": "S61sh1_cine.png",
  "staged_sha256": "ff9d94fb404a829e5384e296555bbef5f00ab17f083abb33fff0e3f704f7b2e7",
  "latency_ms": 9891
 },
 "S61sh5::signage": {
  "fp": "f2414ebd8b88048f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S61sh5": {
  "input_fingerprint": "367cdc504ace206d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 한쪽 입꼬리를 비스듬히 올린 채 거만하게 웃는 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the stranded camper in the dark marsh, near a corroded radiation warning sign. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 자동차 옆면 일부 (Stationary beside the inspection) — A small, softly focused section of the near side remains behind 박철진's shoulder; used as Connects the close-up to the preceding approach without revealing additional people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established nighttime ambience with controlled facial contrast that makes the crooked smile readable without introducing a new source or color cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the camper embedded in swamp mud and the same dark nighttime environment. Exclude the children and any daytime sunlight from the earlier swamp sequence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stranded in swamp mud beside the rusty radiation-zone sign. It is night, and crows remain gathered near the sign before taking flight. 박철진: He remains at the swamp inspection site.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 한쪽 입꼬리를 비스듬히 올린 채 거만하게 웃는 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the stranded camper in the dark marsh, near a corroded radiation warning sign. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 자동차 옆면 일부 (Stationary beside the inspection) — A small, softly focused section of the near side remains behind 박철진's shoulder; used as Connects the close-up to the preceding approach without revealing additional people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established nighttime ambience with controlled facial contrast that makes the crooked smile readable without introducing a new source or color cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the camper embedded in swamp mud and the same dark nighttime environment. Exclude the children and any daytime sunlight from the earlier swamp sequence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stranded in swamp mud beside the rusty radiation-zone sign. It is night, and crows remain gathered near the sign before taking flight. 박철진: He remains at the swamp inspection site.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 한쪽 입꼬리를 비스듬히 올린 채 거만하게 웃는 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the stranded camper in the dark marsh, near a corroded radiation warning sign. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 자동차 옆면 일부 (Stationary beside the inspection) — A small, softly focused section of the near side remains behind 박철진's shoulder; used as Connects the close-up to the preceding approach without revealing additional people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established nighttime ambience with controlled facial contrast that makes the crooked smile readable without introducing a new source or color cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the camper embedded in swamp mud and the same dark nighttime environment. Exclude the children and any daytime sunlight from the earlier swamp sequence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stranded in swamp mud beside the rusty radiation-zone sign. It is night, and crows remain gathered near the sign before taking flight. 박철진: He remains at the swamp inspection site.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물은 정면의 카메라를 응시하고 있음.",
    "built_space": "캐릭터 뒤로 캠핑카의 측면이 화면 전체를 덮고 있으며, 레퍼런스의 좌측면에서는 볼 수 없었던 출입문이 중앙에 배치되어 있음.",
    "entities": "박철진이 요구된 40대 한국인 남성의 외모와 검은색 작업복 차림으로 한쪽 입꼬리를 올리고 있음.",
    "hard_violations": [],
    "physics": "인물의 자세와 표정에 어색함이 없고 자연스러움."
   },
   {
    "label": "B",
    "direction": "인물은 정면의 카메라를 향해 시선을 두고 있음.",
    "built_space": "인물 뒤쪽 좌측에 늪에 빠진 캠핑카의 앞면과 측면이 보이고, 우측에는 까마귀가 앉아 있는 녹슨 방사능 경고 표지판이 레퍼런스의 공간감과 일치하게 배치됨.",
    "entities": "박철진이 레퍼런스와 일치하는 얼굴 및 복장으로 한쪽 입꼬리를 올린 거만한 미소를 짓고 있으며, 방사능 표지판과 까마귀도 적절히 묘사됨.",
    "hard_violations": [],
    "physics": "인물의 자세 및 배경 요소들이 물리적으로 자연스럽게 지면에 안착되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "요구된 인물의 표정(비스듬히 올린 입꼬리)을 완벽하게 재현했으며, 캠핑카와 방사능 경고 표지판의 배치 및 크기를 레퍼런스와 동일하게 훌륭하게 유지했습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물의 얼굴과 표정은 프롬프트에 잘 부합하나, 배경의 캠핑카가 프레임 전체를 차지하며 레퍼런스에 없는 측면 문이 나타나는 등 배경 요소의 스케일과 디테일 묘사가 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 정면의 카메라를 응시하고 있음.",
        "built_space": "캐릭터 뒤로 캠핑카의 측면이 화면 전체를 덮고 있으며, 레퍼런스의 좌측면에서는 볼 수 없었던 출입문이 중앙에 배치되어 있음.",
        "entities": "박철진이 요구된 40대 한국인 남성의 외모와 검은색 작업복 차림으로 한쪽 입꼬리를 올리고 있음.",
        "hard_violations": [],
        "physics": "인물의 자세와 표정에 어색함이 없고 자연스러움."
       },
       {
        "label": "B",
        "direction": "인물은 정면의 카메라를 향해 시선을 두고 있음.",
        "built_space": "인물 뒤쪽 좌측에 늪에 빠진 캠핑카의 앞면과 측면이 보이고, 우측에는 까마귀가 앉아 있는 녹슨 방사능 경고 표지판이 레퍼런스의 공간감과 일치하게 배치됨.",
        "entities": "박철진이 레퍼런스와 일치하는 얼굴 및 복장으로 한쪽 입꼬리를 올린 거만한 미소를 짓고 있으며, 방사능 표지판과 까마귀도 적절히 묘사됨.",
        "hard_violations": [],
        "physics": "인물의 자세 및 배경 요소들이 물리적으로 자연스럽게 지면에 안착되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "요구된 인물의 표정(비스듬히 올린 입꼬리)을 완벽하게 재현했으며, 캠핑카와 방사능 경고 표지판의 배치 및 크기를 레퍼런스와 동일하게 훌륭하게 유지했습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물의 얼굴과 표정은 프롬프트에 잘 부합하나, 배경의 캠핑카가 프레임 전체를 차지하며 레퍼런스에 없는 측면 문이 나타나는 등 배경 요소의 스케일과 디테일 묘사가 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 정면의 카메라를 응시하고 있음.",
        "built_space": "캐릭터 뒤로 캠핑카의 측면이 화면 전체를 덮고 있으며, 레퍼런스의 좌측면에서는 볼 수 없었던 출입문이 중앙에 배치되어 있음.",
        "entities": "박철진이 요구된 40대 한국인 남성의 외모와 검은색 작업복 차림으로 한쪽 입꼬리를 올리고 있음.",
        "hard_violations": [],
        "physics": "인물의 자세와 표정에 어색함이 없고 자연스러움."
       },
       {
        "label": "B",
        "direction": "인물은 정면의 카메라를 향해 시선을 두고 있음.",
        "built_space": "인물 뒤쪽 좌측에 늪에 빠진 캠핑카의 앞면과 측면이 보이고, 우측에는 까마귀가 앉아 있는 녹슨 방사능 경고 표지판이 레퍼런스의 공간감과 일치하게 배치됨.",
        "entities": "박철진이 레퍼런스와 일치하는 얼굴 및 복장으로 한쪽 입꼬리를 올린 거만한 미소를 짓고 있으며, 방사능 표지판과 까마귀도 적절히 묘사됨.",
        "hard_violations": [],
        "physics": "인물의 자세 및 배경 요소들이 물리적으로 자연스럽게 지면에 안착되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "비뚤어진 미소와 얼굴 클로즈업은 충실하지만, 캠핑카의 넓은 부분과 크게 드러난 경고판이 작은 차량 옆면만 배경으로 남기라는 구성을 약화한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "한쪽 입꼬리를 올린 거만한 표정과 얼굴 클로즈업, 흐릿한 캠핑카 옆면으로 이어지는 구성이 더 충실하지만 차량 배경의 비중은 여전히 크다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이며 시선은 렌즈를 향한다. 화면 오른쪽 입꼬리가 더 올라가 비웃는 표정이 읽힌다. 명시된 시선 대상은 없지만 자연스러운 연기 기본 지침의 렌즈 응시 회피에는 덜 부합한다. 무기나 손에 든 지향성 물체는 없다.",
        "built_space": "화면 왼쪽에는 캠핑카 한 대의 옆면 창들과 출입문, 지붕 및 앞쪽 일부까지 보인다. 오른쪽에는 녹슨 방사능 경고판 하나와 그 위에 걸터앉은 까마귀 한 마리가 일부 보인다. 인물은 차량과 표지판 앞에 있으며 습지 수면과 풀이 뒤로 이어진다. 고정물 중복이나 불가능한 반사는 없지만, 작은 차량 옆면 조각만 남기는 배경보다 장소 요소를 많이 노출한다.",
        "entities": "인물은 한 명뿐이며 짧게 넘긴 검은 머리의 한국인 중년 남성으로 묘사되어 박철진의 기본 설정과 맞는다. 참고 얼굴의 주요 윤곽은 닮았으나 수염 자국과 피부 주름이 더 두드러진다. 보이는 어두운 작업복 깃은 참고 복장과 대체로 맞는다. 낡은 캠핑카, 부식된 방사능 표지판, 까마귀와 밤의 습지가 확인된다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 목은 어깨 및 몸통에 정상적으로 연결되어 있다. 하체와 지면 접촉은 클로즈업 밖이므로 확인할 수 없으며 부유를 시사하는 모습은 없다. 캠핑카는 진흙 지면에 놓여 있고 표지판은 기둥으로 지지된다. 까마귀는 표지판 윗부분에 앉아 있어 지지점이 있다."
       },
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이고 눈은 렌즈 부근을 향한다. 화면 오른쪽 입꼬리를 비스듬히 올리고 눈꺼풀을 살짝 좁혀 거만한 미소를 만든다. 특정 응시 대상은 화면에 없으며, 렌즈 응시를 피하라는 기본 지침은 충분히 구현되지 않았다. 무기나 방향을 확인해야 할 휴대 물체는 없다.",
        "built_space": "인물 바로 뒤에 캠핑카 한 대의 옆면이 흐릿하게 보인다. 왼쪽에 큰 창 하나, 오른쪽에 세로 창이 달린 출입문 하나, 맨 오른쪽에 다른 창 일부가 있으며 지붕 가장자리도 보인다. 인물은 차량 옆에 서 있는 관계로 읽힌다. 차량 옆면이 배경 대부분을 차지해 지정된 작은 배경 조각보다는 넓지만, 차량 전체나 별도 표지판을 드러내지 않아 얼굴 중심 구성을 더 잘 유지한다. 불가능한 반사나 중복 출입문은 보이지 않는다.",
        "entities": "짧은 검은 머리의 중년 한국인 남성 한 명만 등장한다. 얼굴 윤곽과 머리 모양은 박철진 참고 이미지와 대체로 일치하며, 수염 자국은 참고보다 진하다. 어두운 작업복의 깃과 어깨 부분이 보인다. 배경 차량은 낡은 밝은색 외판과 어두운 가로 띠를 가진 캠핑카로 장소 참고와 연결된다. 표지판, 까마귀, 진흙에 박힌 바퀴는 클로즈업 밖이라 상태를 판단할 수 없다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리, 목, 어깨의 연결과 미소를 만드는 안면 근육이 자연스럽다. 발과 지면은 프레임 밖이며 공중에 떠 있다는 단서는 없다. 캠핑카의 창과 문은 차체에 부착되어 있고, 별도로 떠 있거나 지지 없이 들린 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "비뚤어진 미소와 얼굴 클로즈업은 충실하지만, 캠핑카의 넓은 부분과 크게 드러난 경고판이 작은 차량 옆면만 배경으로 남기라는 구성을 약화한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "한쪽 입꼬리를 올린 거만한 표정과 얼굴 클로즈업, 흐릿한 캠핑카 옆면으로 이어지는 구성이 더 충실하지만 차량 배경의 비중은 여전히 크다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이며 시선은 렌즈를 향한다. 화면 오른쪽 입꼬리가 더 올라가 비웃는 표정이 읽힌다. 명시된 시선 대상은 없지만 자연스러운 연기 기본 지침의 렌즈 응시 회피에는 덜 부합한다. 무기나 손에 든 지향성 물체는 없다.",
        "built_space": "화면 왼쪽에는 캠핑카 한 대의 옆면 창들과 출입문, 지붕 및 앞쪽 일부까지 보인다. 오른쪽에는 녹슨 방사능 경고판 하나와 그 위에 걸터앉은 까마귀 한 마리가 일부 보인다. 인물은 차량과 표지판 앞에 있으며 습지 수면과 풀이 뒤로 이어진다. 고정물 중복이나 불가능한 반사는 없지만, 작은 차량 옆면 조각만 남기는 배경보다 장소 요소를 많이 노출한다.",
        "entities": "인물은 한 명뿐이며 짧게 넘긴 검은 머리의 한국인 중년 남성으로 묘사되어 박철진의 기본 설정과 맞는다. 참고 얼굴의 주요 윤곽은 닮았으나 수염 자국과 피부 주름이 더 두드러진다. 보이는 어두운 작업복 깃은 참고 복장과 대체로 맞는다. 낡은 캠핑카, 부식된 방사능 표지판, 까마귀와 밤의 습지가 확인된다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 목은 어깨 및 몸통에 정상적으로 연결되어 있다. 하체와 지면 접촉은 클로즈업 밖이므로 확인할 수 없으며 부유를 시사하는 모습은 없다. 캠핑카는 진흙 지면에 놓여 있고 표지판은 기둥으로 지지된다. 까마귀는 표지판 윗부분에 앉아 있어 지지점이 있다."
       },
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이고 눈은 렌즈 부근을 향한다. 화면 오른쪽 입꼬리를 비스듬히 올리고 눈꺼풀을 살짝 좁혀 거만한 미소를 만든다. 특정 응시 대상은 화면에 없으며, 렌즈 응시를 피하라는 기본 지침은 충분히 구현되지 않았다. 무기나 방향을 확인해야 할 휴대 물체는 없다.",
        "built_space": "인물 바로 뒤에 캠핑카 한 대의 옆면이 흐릿하게 보인다. 왼쪽에 큰 창 하나, 오른쪽에 세로 창이 달린 출입문 하나, 맨 오른쪽에 다른 창 일부가 있으며 지붕 가장자리도 보인다. 인물은 차량 옆에 서 있는 관계로 읽힌다. 차량 옆면이 배경 대부분을 차지해 지정된 작은 배경 조각보다는 넓지만, 차량 전체나 별도 표지판을 드러내지 않아 얼굴 중심 구성을 더 잘 유지한다. 불가능한 반사나 중복 출입문은 보이지 않는다.",
        "entities": "짧은 검은 머리의 중년 한국인 남성 한 명만 등장한다. 얼굴 윤곽과 머리 모양은 박철진 참고 이미지와 대체로 일치하며, 수염 자국은 참고보다 진하다. 어두운 작업복의 깃과 어깨 부분이 보인다. 배경 차량은 낡은 밝은색 외판과 어두운 가로 띠를 가진 캠핑카로 장소 참고와 연결된다. 표지판, 까마귀, 진흙에 박힌 바퀴는 클로즈업 밖이라 상태를 판단할 수 없다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리, 목, 어깨의 연결과 미소를 만드는 안면 근육이 자연스럽다. 발과 지면은 프레임 밖이며 공중에 떠 있다는 단서는 없다. 캠핑카의 창과 문은 차체에 부착되어 있고, 별도로 떠 있거나 지지 없이 들린 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.7,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.7,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1700
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "요구된 인물의 표정(비스듬히 올린 입꼬리)을 완벽하게 재현했으며, 캠핑카와 방사능 경고 표지판의 배치 및 크기를 레퍼런스와 동일하게 훌륭하게 유지했습니다."
   },
   {
    "label": "A",
    "score": 1700,
    "verdict_ko": "인물의 얼굴과 표정은 프롬프트에 잘 부합하나, 배경의 캠핑카가 프레임 전체를 차지하며 레퍼런스에 없는 측면 문이 나타나는 등 배경 요소의 스케일과 디테일 묘사가 아쉽습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S61sh1_sel.png",
    "asset_id": "01a84a5e-4855-4dcb-a58c-6c7036065a30",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-3e41-799f-9993-fdfd9b872919",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S61sh1"
  }
 },
 "S61sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:57:11.758733+00:00",
  "fingerprint": "bb9acf6ccb1c57bb17f1d0510f4877dbbd20507a48082ae2b0a2966ee2241714",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S61sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S61sh5_sel.png",
  "source_sha256": "efabffdf8afe179831adc17fcfdfa63552445919995c820d2285578396ca38f1",
  "file": "S61sh5_cine.png",
  "staged_sha256": "532e909edc046e5b92f149698abf47f9840e94c9d69a3c847bacf0b9a1fbb157",
  "latency_ms": 10322
 },
 "S61sh7::signage": {
  "fp": "f4d270f2c24fcfab",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "방사능 구역 팻말"
   }
  ],
  "dropped": []
 },
 "S61sh7": {
  "input_fingerprint": "211e8f2dbdd07509",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어둠 속 진흙 바닥에 꽂힌 낡고 녹슨 방사능 구역 팻말 클로즈업.\n\nLOCATION (lock): At the rusted radiation warning sign planted in the muddy marsh ground at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 방사능 구역 팻말 (Old and rusted, planted in the muddy ground) — The camera sees the front bearing the radiation-area warning at a shallow oblique angle; the rear is not presented; used as Provides the focused narrative evidence while remaining surrounded by visible environmental context; 팻말 아래 진흙 바닥 (The sign support is embedded in the mud); used as Anchors the warning physically in the swamp and prevents an isolated floating-sign composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the dark nighttime ambience while separating the rusty warning face from its surroundings just enough for the warning to read.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rusty radiation-zone sign remains planted beside the swamp, where the old camper is still stranded. Night surrounds the site as the crows take flight and the moon becomes visible.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어둠 속 진흙 바닥에 꽂힌 낡고 녹슨 방사능 구역 팻말 클로즈업.\n\nLOCATION (lock): At the rusted radiation warning sign planted in the muddy marsh ground at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 방사능 구역 팻말 (Old and rusted, planted in the muddy ground) — The camera sees the front bearing the radiation-area warning at a shallow oblique angle; the rear is not presented; used as Provides the focused narrative evidence while remaining surrounded by visible environmental context; 팻말 아래 진흙 바닥 (The sign support is embedded in the mud); used as Anchors the warning physically in the swamp and prevents an isolated floating-sign composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the dark nighttime ambience while separating the rusty warning face from its surroundings just enough for the warning to read.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rusty radiation-zone sign remains planted beside the swamp, where the old camper is still stranded. Night surrounds the site as the crows take flight and the moon becomes visible.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어둠 속 진흙 바닥에 꽂힌 낡고 녹슨 방사능 구역 팻말 클로즈업.\n\nLOCATION (lock): At the rusted radiation warning sign planted in the muddy marsh ground at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 방사능 구역 팻말 (Old and rusted, planted in the muddy ground) — The camera sees the front bearing the radiation-area warning at a shallow oblique angle; the rear is not presented; used as Provides the focused narrative evidence while remaining surrounded by visible environmental context; 팻말 아래 진흙 바닥 (The sign support is embedded in the mud); used as Anchors the warning physically in the swamp and prevents an isolated floating-sign composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the dark nighttime ambience while separating the rusty warning face from its surroundings just enough for the warning to read.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rusty radiation-zone sign remains planted beside the swamp, where the old camper is still stranded. Night surrounds the site as the crows take flight and the moon becomes visible.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 진흙 바닥에 꽂힌 방사능 표지판의 정면을 향하고 있습니다.",
    "built_space": "야외의 늪지대이며, 배경 좌측에 낡은 캠핑카가 위치해 레퍼런스의 공간과 일치합니다.",
    "entities": "녹슨 사각형 방사능 표지판과 진흙 바닥, 배경의 캠핑카 모두 프롬프트 요구사항과 일치합니다.",
    "hard_violations": [],
    "physics": "표지판의 기둥이 진흙 바닥에 깊게 박혀 안정적으로 서 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 진흙에 꽂힌 방사능 표지판을 비스듬한 각도로 바라보고 있습니다.",
    "built_space": "야외 늪지대이며 배경 좌측에 캠핑카, 우측에 원형 표지판이 배치되어 레퍼런스와 일치합니다.",
    "entities": "녹슨 사각형 방사능 표지판, 진흙 바닥, 까마귀 등 요소가 존재하나, 표지판 하단에 한글 텍스트가 있습니다.",
    "hard_violations": [
     "[gemini-pro] 텍스트 유출 (표지판에 읽을 수 있는 한글 텍스트가 포함됨)",
     "[gpt-high] 팻말에 판독 가능한 한글이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "표지판 기둥이 진흙에 박혀 형태를 유지하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "글자가 없어야 한다는 지침을 준수하며, 진흙에 꽂힌 녹슨 방사능 표지판의 클로즈업 샷을 성공적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "표지판 앵글과 배경 배치는 좋으나, 프롬프트에서 엄격히 금지한 읽을 수 있는 텍스트가 포함되어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 진흙 바닥에 꽂힌 방사능 표지판의 정면을 향하고 있습니다.",
        "built_space": "야외의 늪지대이며, 배경 좌측에 낡은 캠핑카가 위치해 레퍼런스의 공간과 일치합니다.",
        "entities": "녹슨 사각형 방사능 표지판과 진흙 바닥, 배경의 캠핑카 모두 프롬프트 요구사항과 일치합니다.",
        "hard_violations": [],
        "physics": "표지판의 기둥이 진흙 바닥에 깊게 박혀 안정적으로 서 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 진흙에 꽂힌 방사능 표지판을 비스듬한 각도로 바라보고 있습니다.",
        "built_space": "야외 늪지대이며 배경 좌측에 캠핑카, 우측에 원형 표지판이 배치되어 레퍼런스와 일치합니다.",
        "entities": "녹슨 사각형 방사능 표지판, 진흙 바닥, 까마귀 등 요소가 존재하나, 표지판 하단에 한글 텍스트가 있습니다.",
        "hard_violations": [
         "텍스트 유출 (표지판에 읽을 수 있는 한글 텍스트가 포함됨)"
        ],
        "physics": "표지판 기둥이 진흙에 박혀 형태를 유지하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "글자가 없어야 한다는 지침을 준수하며, 진흙에 꽂힌 녹슨 방사능 표지판의 클로즈업 샷을 성공적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "표지판 앵글과 배경 배치는 좋으나, 프롬프트에서 엄격히 금지한 읽을 수 있는 텍스트가 포함되어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 진흙 바닥에 꽂힌 방사능 표지판의 정면을 향하고 있습니다.",
        "built_space": "야외의 늪지대이며, 배경 좌측에 낡은 캠핑카가 위치해 레퍼런스의 공간과 일치합니다.",
        "entities": "녹슨 사각형 방사능 표지판과 진흙 바닥, 배경의 캠핑카 모두 프롬프트 요구사항과 일치합니다.",
        "hard_violations": [],
        "physics": "표지판의 기둥이 진흙 바닥에 깊게 박혀 안정적으로 서 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 진흙에 꽂힌 방사능 표지판을 비스듬한 각도로 바라보고 있습니다.",
        "built_space": "야외 늪지대이며 배경 좌측에 캠핑카, 우측에 원형 표지판이 배치되어 레퍼런스와 일치합니다.",
        "entities": "녹슨 사각형 방사능 표지판, 진흙 바닥, 까마귀 등 요소가 존재하나, 표지판 하단에 한글 텍스트가 있습니다.",
        "hard_violations": [
         "텍스트 유출 (표지판에 읽을 수 있는 한글 텍스트가 포함됨)"
        ],
        "physics": "표지판 기둥이 진흙에 박혀 형태를 유지하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "팻말 앞면과 진흙에 박힌 지지대는 보이지만, 판독 가능한 한글이 문자 금지 조건을 위반하고 까마귀도 날아오르지 않고 앉아 있다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "글자 없는 녹슨 방사능 팻말과 진흙 접점을 클로즈업해 핵심 지시를 충족하지만, 지지대의 노출 길이가 참조보다 크게 짧아졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "방사능 표식이 있는 앞면이 카메라를 향해 약간 비스듬히 놓였고 뒷면은 보이지 않는다. 배경 캠핑카의 앞면도 카메라 쪽이다. 까마귀들은 지붕과 원형 표지판 위에 앉아 있으며 날아가는 방향은 나타나지 않는다.",
        "built_space": "전경에 직사각형 방사능 팻말 하나와 지지대 하나, 왼쪽 배경에 캠핑카 한 대, 오른쪽 배경에 원형 표지판 하나가 보인다. 참조의 주요 시설 종류와 개수는 유지되지만 원형 표지판의 화면상 위치는 달라졌다. 팻말 상단은 프레임 밖으로 잘리고 하단에는 지지대가 박힌 진흙과 웅덩이가 보인다.",
        "entities": "사람은 없다. 녹슨 금속 팻말, 방사능 기호, 낡은 캠핑카, 갈대와 진흙 습지가 보인다. 팻말 아래쪽에는 개별 글자를 판독할 수 있는 한글 두 줄이 생겨 참조 및 문자 금지 지시와 어긋난다. 까마귀 여섯 마리가 배경 시설 위에 앉아 있다. 차가운 야간 조명은 유지되지만 달 자체는 보이지 않는다.",
        "hard_violations": [
         "팻말에 판독 가능한 한글이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "금속 팻말은 지지대에 연결되어 있고 지지대의 하단은 진흙에 박혀 있다. 캠핑카는 진흙 지면에 얹혀 있으며 까마귀들은 지붕이나 표지판 가장자리에 발을 대고 있다. 지지 없이 떠 있는 물체는 없지만 까마귀가 날아오르는 순간은 구현하지 않았다."
       },
       {
        "label": "B",
        "direction": "방사능 팻말의 앞면이 카메라를 거의 정면으로 향하되 약한 사선 각도를 보이며, 기호 전체가 드러나고 뒷면은 보이지 않는다. 왼쪽 가장자리의 캠핑카는 전면 일부가 카메라 쪽을 향한다. 사람의 시선이나 이동하는 생물은 없다.",
        "built_space": "중앙에 직사각형 팻말 하나와 짧게 노출된 지지대 하나가 있고, 왼쪽 배경에는 캠핑카 한 대의 전면 일부만 들어온다. 진흙, 얕은 물과 갈대가 팻말 주변을 채운다. 참조의 원형 표지판은 이 좁은 구도에서 보이지 않으며 중복 시설은 없다. 다만 참조에서 길게 드러났던 지지대가 여기서는 대부분 사라져 팻말이 지면에 훨씬 가까워졌다.",
        "entities": "사람이나 얼굴은 없다. 녹슨 금속판 위 방사능 기호, 진흙에 박힌 지지대, 낡은 캠핑카 일부와 습지가 식별된다. 읽을 수 있는 문자나 추가 문구는 없고 기호와 녹이 금속 표면에 붙어 있는 것으로 보인다. 참조의 차가운 달빛 야간 분위기를 유지한다. 까마귀와 달 자체는 프레임에 없다.",
        "hard_violations": [],
        "physics": "팻말 아래 중앙의 지지대가 금속판을 받치고 진흙 속으로 들어가는 접점까지 보인다. 캠핑카는 배경 지면에 놓여 있고 물은 낮은 곳에 고여 빛을 반사한다. 공중에 떠 있거나 지지체가 없는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "팻말 앞면과 진흙에 박힌 지지대는 보이지만, 판독 가능한 한글이 문자 금지 조건을 위반하고 까마귀도 날아오르지 않고 앉아 있다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "글자 없는 녹슨 방사능 팻말과 진흙 접점을 클로즈업해 핵심 지시를 충족하지만, 지지대의 노출 길이가 참조보다 크게 짧아졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "방사능 표식이 있는 앞면이 카메라를 향해 약간 비스듬히 놓였고 뒷면은 보이지 않는다. 배경 캠핑카의 앞면도 카메라 쪽이다. 까마귀들은 지붕과 원형 표지판 위에 앉아 있으며 날아가는 방향은 나타나지 않는다.",
        "built_space": "전경에 직사각형 방사능 팻말 하나와 지지대 하나, 왼쪽 배경에 캠핑카 한 대, 오른쪽 배경에 원형 표지판 하나가 보인다. 참조의 주요 시설 종류와 개수는 유지되지만 원형 표지판의 화면상 위치는 달라졌다. 팻말 상단은 프레임 밖으로 잘리고 하단에는 지지대가 박힌 진흙과 웅덩이가 보인다.",
        "entities": "사람은 없다. 녹슨 금속 팻말, 방사능 기호, 낡은 캠핑카, 갈대와 진흙 습지가 보인다. 팻말 아래쪽에는 개별 글자를 판독할 수 있는 한글 두 줄이 생겨 참조 및 문자 금지 지시와 어긋난다. 까마귀 여섯 마리가 배경 시설 위에 앉아 있다. 차가운 야간 조명은 유지되지만 달 자체는 보이지 않는다.",
        "hard_violations": [
         "팻말에 판독 가능한 한글이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "금속 팻말은 지지대에 연결되어 있고 지지대의 하단은 진흙에 박혀 있다. 캠핑카는 진흙 지면에 얹혀 있으며 까마귀들은 지붕이나 표지판 가장자리에 발을 대고 있다. 지지 없이 떠 있는 물체는 없지만 까마귀가 날아오르는 순간은 구현하지 않았다."
       },
       {
        "label": "A",
        "direction": "방사능 팻말의 앞면이 카메라를 거의 정면으로 향하되 약한 사선 각도를 보이며, 기호 전체가 드러나고 뒷면은 보이지 않는다. 왼쪽 가장자리의 캠핑카는 전면 일부가 카메라 쪽을 향한다. 사람의 시선이나 이동하는 생물은 없다.",
        "built_space": "중앙에 직사각형 팻말 하나와 짧게 노출된 지지대 하나가 있고, 왼쪽 배경에는 캠핑카 한 대의 전면 일부만 들어온다. 진흙, 얕은 물과 갈대가 팻말 주변을 채운다. 참조의 원형 표지판은 이 좁은 구도에서 보이지 않으며 중복 시설은 없다. 다만 참조에서 길게 드러났던 지지대가 여기서는 대부분 사라져 팻말이 지면에 훨씬 가까워졌다.",
        "entities": "사람이나 얼굴은 없다. 녹슨 금속판 위 방사능 기호, 진흙에 박힌 지지대, 낡은 캠핑카 일부와 습지가 식별된다. 읽을 수 있는 문자나 추가 문구는 없고 기호와 녹이 금속 표면에 붙어 있는 것으로 보인다. 참조의 차가운 달빛 야간 분위기를 유지한다. 까마귀와 달 자체는 프레임에 없다.",
        "hard_violations": [],
        "physics": "팻말 아래 중앙의 지지대가 금속판을 받치고 진흙 속으로 들어가는 접점까지 보인다. 캠핑카는 배경 지면에 놓여 있고 물은 낮은 곳에 고여 빛을 반사한다. 공중에 떠 있거나 지지체가 없는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 텍스트 유출 (표지판에 읽을 수 있는 한글 텍스트가 포함됨)",
     "[gpt-high] 팻말에 판독 가능한 한글이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "글자가 없어야 한다는 지침을 준수하며, 진흙에 꽂힌 녹슨 방사능 표지판의 클로즈업 샷을 성공적으로 구현했습니다."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "표지판 앵글과 배경 배치는 좋으나, 프롬프트에서 엄격히 금지한 읽을 수 있는 텍스트가 포함되어 감점되었습니다.  ★위반: [gemini-pro] 텍스트 유출 (표지판에 읽을 수 있는 한글 텍스트가 포함됨) / [gpt-high] 팻말에 판독 가능한 한글이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S61sh1_sel.png",
    "asset_id": "01a84a5e-4855-4dcb-a58c-6c7036065a30",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-3fe7-7781-a877-ca2daf203472",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S61sh1"
  },
  "lane_policy": "share_plan_prev_bgonly"
 },
 "S61sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:58:02.746436+00:00",
  "fingerprint": "6896945ebc3e00444ab0beb60e0a36719e2a4e0c26d438d89a3f58bc0c845d18",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S61sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S61sh7_sel.png",
  "source_sha256": "01c027b8bae61a45341b892885d6c580bb761480485533fd2a986939961f322b",
  "file": "S61sh7_cine.png",
  "staged_sha256": "a68ed6b5ff33c583439cf3cfddaf792b49b6c2653a0941fe9ffa3692c98ff4e0",
  "latency_ms": 8903
 },
 "S62sh3::signage": {
  "fp": "a03021a1549dfde5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S62sh3": {
  "input_fingerprint": "47b4842d7fa356b7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 활짝 열린 철문 너머, 쌀과 감자 포대 몇 개만 덩그러니 놓인 텅 빈 식량창고 내부.\n\nLOCATION (lock): Inside the village food storehouse, with a few rice and potato sacks illuminated by daylight through the open metal door. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 식량창고 철문 (Fully open) — The open door leaf is seen obliquely along the left edge, leaving the inward view unobstructed; used as Creates a threshold frame and establishes that the empty space is the revealed storage interior; 쌀 포대 몇 개와 감자 한 포대 (The few remaining provisions in the otherwise empty warehouse); used as Forms a small, readable supply cluster whose modest scale is measured against the surrounding empty floor; 비어 있는 창고 내부 (Largely empty) — The diagonal view reveals the interior floor extending beyond the remaining sacks; used as Makes absence, rather than the sacks themselves, the main spatial statement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient daylight appropriate to the open storage entrance, preserving enough interior detail for the sparse supplies and surrounding emptiness to register.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The food store's old iron door is open, revealing only a few sacks of rice and one sack of potatoes. Elsewhere in the village square, an old truck is undergoing starting attempts; Charlie's dents, holes and sensor damage remain unrepaired.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 활짝 열린 철문 너머, 쌀과 감자 포대 몇 개만 덩그러니 놓인 텅 빈 식량창고 내부.\n\nLOCATION (lock): Inside the village food storehouse, with a few rice and potato sacks illuminated by daylight through the open metal door. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 식량창고 철문 (Fully open) — The open door leaf is seen obliquely along the left edge, leaving the inward view unobstructed; used as Creates a threshold frame and establishes that the empty space is the revealed storage interior; 쌀 포대 몇 개와 감자 한 포대 (The few remaining provisions in the otherwise empty warehouse); used as Forms a small, readable supply cluster whose modest scale is measured against the surrounding empty floor; 비어 있는 창고 내부 (Largely empty) — The diagonal view reveals the interior floor extending beyond the remaining sacks; used as Makes absence, rather than the sacks themselves, the main spatial statement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient daylight appropriate to the open storage entrance, preserving enough interior detail for the sparse supplies and surrounding emptiness to register.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The food store's old iron door is open, revealing only a few sacks of rice and one sack of potatoes. Elsewhere in the village square, an old truck is undergoing starting attempts; Charlie's dents, holes and sensor damage remain unrepaired.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 활짝 열린 철문 너머, 쌀과 감자 포대 몇 개만 덩그러니 놓인 텅 빈 식량창고 내부.\n\nLOCATION (lock): Inside the village food storehouse, with a few rice and potato sacks illuminated by daylight through the open metal door. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 식량창고 철문 (Fully open) — The open door leaf is seen obliquely along the left edge, leaving the inward view unobstructed; used as Creates a threshold frame and establishes that the empty space is the revealed storage interior; 쌀 포대 몇 개와 감자 한 포대 (The few remaining provisions in the otherwise empty warehouse); used as Forms a small, readable supply cluster whose modest scale is measured against the surrounding empty floor; 비어 있는 창고 내부 (Largely empty) — The diagonal view reveals the interior floor extending beyond the remaining sacks; used as Makes absence, rather than the sacks themselves, the main spatial statement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient daylight appropriate to the open storage entrance, preserving enough interior detail for the sparse supplies and surrounding emptiness to register.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The food store's old iron door is open, revealing only a few sacks of rice and one sack of potatoes. Elsewhere in the village square, an old truck is undergoing starting attempts; Charlie's dents, holes and sensor damage remain unrepaired.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 창고 내부를 비추며 중앙 안쪽의 또 다른 출입구를 향함.",
    "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 정면 안쪽에 또 다른 열린 문이 보이나 바깥 풍경이 레퍼런스와 일치하지 않음.",
    "entities": "바닥에 여러 개의 포대가 모여 있음. 사람 없음. 포대 겉면에 'POTATO' 등의 영어 단어가 명확하게 적혀 있음.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)",
     "[gpt-high] 감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "포대들이 바닥에 자연스럽게 놓여 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 창고 내부를 향해 사선으로 넓게 비추고 있으며, 시선은 바닥에 놓인 포대들과 텅 빈 공간으로 향함.",
    "built_space": "창고 내부. 화면 왼쪽에 활짝 열린 철문이 비스듬히 자리 잡고 있으며, 문 너머로 레퍼런스의 녹슨 급수탑과 주변 건물이 올바른 위치 관계로 보임. 내부는 넓고 비어 있음.",
    "entities": "바닥에 쌀 포대와 감자가 노출된 포대가 놓여 있음. 사람 없음. 포대에 적힌 글씨는 뭉개져 있어 읽을 수 없음.",
    "hard_violations": [
     "[gpt-high] 오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
    ],
    "physics": "포대들이 바닥에 안정적으로 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "금지된 식별 가능한 텍스트가 없으며, 열린 문 너머로 레퍼런스의 급수탑을 배치하여 지정된 로케이션을 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시적으로 금지한 읽을 수 있는 텍스트가 포함되어 심각한 위반에 해당합니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 창고 내부를 향해 사선으로 넓게 비추고 있으며, 시선은 바닥에 놓인 포대들과 텅 빈 공간으로 향함.",
        "built_space": "창고 내부. 화면 왼쪽에 활짝 열린 철문이 비스듬히 자리 잡고 있으며, 문 너머로 레퍼런스의 녹슨 급수탑과 주변 건물이 올바른 위치 관계로 보임. 내부는 넓고 비어 있음.",
        "entities": "바닥에 쌀 포대와 감자가 노출된 포대가 놓여 있음. 사람 없음. 포대에 적힌 글씨는 뭉개져 있어 읽을 수 없음.",
        "hard_violations": [],
        "physics": "포대들이 바닥에 안정적으로 놓여 있음."
       },
       {
        "label": "A",
        "direction": "카메라는 창고 내부를 비추며 중앙 안쪽의 또 다른 출입구를 향함.",
        "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 정면 안쪽에 또 다른 열린 문이 보이나 바깥 풍경이 레퍼런스와 일치하지 않음.",
        "entities": "바닥에 여러 개의 포대가 모여 있음. 사람 없음. 포대 겉면에 'POTATO' 등의 영어 단어가 명확하게 적혀 있음.",
        "hard_violations": [
         "읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)"
        ],
        "physics": "포대들이 바닥에 자연스럽게 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "금지된 식별 가능한 텍스트가 없으며, 열린 문 너머로 레퍼런스의 급수탑을 배치하여 지정된 로케이션을 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시적으로 금지한 읽을 수 있는 텍스트가 포함되어 심각한 위반에 해당합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 창고 내부를 향해 사선으로 넓게 비추고 있으며, 시선은 바닥에 놓인 포대들과 텅 빈 공간으로 향함.",
        "built_space": "창고 내부. 화면 왼쪽에 활짝 열린 철문이 비스듬히 자리 잡고 있으며, 문 너머로 레퍼런스의 녹슨 급수탑과 주변 건물이 올바른 위치 관계로 보임. 내부는 넓고 비어 있음.",
        "entities": "바닥에 쌀 포대와 감자가 노출된 포대가 놓여 있음. 사람 없음. 포대에 적힌 글씨는 뭉개져 있어 읽을 수 없음.",
        "hard_violations": [],
        "physics": "포대들이 바닥에 안정적으로 놓여 있음."
       },
       {
        "label": "A",
        "direction": "카메라는 창고 내부를 비추며 중앙 안쪽의 또 다른 출입구를 향함.",
        "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 정면 안쪽에 또 다른 열린 문이 보이나 바깥 풍경이 레퍼런스와 일치하지 않음.",
        "entities": "바닥에 여러 개의 포대가 모여 있음. 사람 없음. 포대 겉면에 'POTATO' 등의 영어 단어가 명확하게 적혀 있음.",
        "hard_violations": [
         "읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)"
        ],
        "physics": "포대들이 바닥에 자연스럽게 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "작은 식량 더미 너머로 넓게 이어지는 빈 바닥과 사선 구도는 더 충실하지만, 뒤쪽에 추가한 판재가 ‘포대 몇 개만 놓인 창고’ 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "왼쪽 철문과 낮의 빈 창고는 표현했지만, 감자 포대의 선명한 영문 글자가 읽을 수 있는 글자 금지 조건을 명백히 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 출입구를 옆에 두고 창고 깊숙한 오른쪽 뒤편으로 비스듬히 향한다. 포대 뒤로 빈 바닥이 길게 이어지지만, 문턱 밖에서 내부를 바라보기보다는 이미 실내에 들어온 시점에 가깝다.",
        "built_space": "왼쪽 출입구 하나에 열린 철문짝 두 개가 보인다. 벗겨진 미장 벽, 콘크리트 바닥, 노출 지붕 트러스, 오른쪽 높은 창 일부, 뒤쪽 왼편의 좁은 문 형태 개구부가 보인다. 외부 급수탑은 참조 장소와 연결된다. 참조 사진은 내부를 보여주지 않아 내부 트러스나 뒤쪽 개구부의 정확한 일치 여부는 확인할 수 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 아래에 곡물 포대로 보이는 네 포대와 감자가 직접 드러난 열린 감자 포대 하나가 있다. 닫힌 곡물 포대의 내용물이 쌀인지는 시각적으로 확정할 수 없다. 포대 인쇄는 보이지만 명확히 읽히는 단어는 확인하기 어렵다. 오른쪽 뒤 벽에는 식량과 무관한 좁은 판재 두 개가 추가되어 있다.",
        "hard_violations": [
         "오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
        ],
        "physics": "맨 아래 포대와 감자 포대는 바닥에 놓였고, 쌓인 포대는 아래 포대가 받친다. 감자는 포대 안에 담겨 있다. 철문은 경첩으로 지지되며 판재는 바닥과 벽에 기대어 있다. 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 열린 철문 옆에서 창고 안쪽 벽과 중앙 오른쪽 식량 더미를 향한다. 내부를 가리는 물체는 없지만, 뒤 벽을 비교적 정면으로 보아 A보다 바닥의 사선 깊이가 덜 강조된다.",
        "built_space": "왼쪽 출입구에 철문짝 두 개가 열려 있고 가까운 문짝이 화면 왼쪽을 크게 차지한다. 콘크리트 바닥, 낡은 미장 벽, 박공지붕 아래 목재 트러스, 오른쪽 상단 창 하나가 보인다. 문과 벽의 재질·색은 참조 창고와 대체로 부합하지만, 참조에서 보이지 않는 내부 구조까지 동일하다고 확인할 수는 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 오른쪽에 곡물 포대로 보이는 세 포대와 감자용으로 표시된 포대 하나가 모여 있다. 모두 닫혀 있어 쌀과 감자 자체는 보이지 않는다. 앞쪽 오른쪽 포대의 ‘POTATO’ 글자는 분명히 읽힌다. 주변 바닥에는 별도 저장 물품이 없다.",
        "hard_violations": [
         "감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "포대 네 개는 바닥에 닿아 있고 서로 기대어 서 있다. 철문은 문틀의 경첩으로 지지되며 지붕 부재는 벽과 트러스에 연결된다. 떠 있거나 지지 관계가 불가능한 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "작은 식량 더미 너머로 넓게 이어지는 빈 바닥과 사선 구도는 더 충실하지만, 뒤쪽에 추가한 판재가 ‘포대 몇 개만 놓인 창고’ 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "왼쪽 철문과 낮의 빈 창고는 표현했지만, 감자 포대의 선명한 영문 글자가 읽을 수 있는 글자 금지 조건을 명백히 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 출입구를 옆에 두고 창고 깊숙한 오른쪽 뒤편으로 비스듬히 향한다. 포대 뒤로 빈 바닥이 길게 이어지지만, 문턱 밖에서 내부를 바라보기보다는 이미 실내에 들어온 시점에 가깝다.",
        "built_space": "왼쪽 출입구 하나에 열린 철문짝 두 개가 보인다. 벗겨진 미장 벽, 콘크리트 바닥, 노출 지붕 트러스, 오른쪽 높은 창 일부, 뒤쪽 왼편의 좁은 문 형태 개구부가 보인다. 외부 급수탑은 참조 장소와 연결된다. 참조 사진은 내부를 보여주지 않아 내부 트러스나 뒤쪽 개구부의 정확한 일치 여부는 확인할 수 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 아래에 곡물 포대로 보이는 네 포대와 감자가 직접 드러난 열린 감자 포대 하나가 있다. 닫힌 곡물 포대의 내용물이 쌀인지는 시각적으로 확정할 수 없다. 포대 인쇄는 보이지만 명확히 읽히는 단어는 확인하기 어렵다. 오른쪽 뒤 벽에는 식량과 무관한 좁은 판재 두 개가 추가되어 있다.",
        "hard_violations": [
         "오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
        ],
        "physics": "맨 아래 포대와 감자 포대는 바닥에 놓였고, 쌓인 포대는 아래 포대가 받친다. 감자는 포대 안에 담겨 있다. 철문은 경첩으로 지지되며 판재는 바닥과 벽에 기대어 있다. 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 열린 철문 옆에서 창고 안쪽 벽과 중앙 오른쪽 식량 더미를 향한다. 내부를 가리는 물체는 없지만, 뒤 벽을 비교적 정면으로 보아 A보다 바닥의 사선 깊이가 덜 강조된다.",
        "built_space": "왼쪽 출입구에 철문짝 두 개가 열려 있고 가까운 문짝이 화면 왼쪽을 크게 차지한다. 콘크리트 바닥, 낡은 미장 벽, 박공지붕 아래 목재 트러스, 오른쪽 상단 창 하나가 보인다. 문과 벽의 재질·색은 참조 창고와 대체로 부합하지만, 참조에서 보이지 않는 내부 구조까지 동일하다고 확인할 수는 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 오른쪽에 곡물 포대로 보이는 세 포대와 감자용으로 표시된 포대 하나가 모여 있다. 모두 닫혀 있어 쌀과 감자 자체는 보이지 않는다. 앞쪽 오른쪽 포대의 ‘POTATO’ 글자는 분명히 읽힌다. 주변 바닥에는 별도 저장 물품이 없다.",
        "hard_violations": [
         "감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "포대 네 개는 바닥에 닿아 있고 서로 기대어 서 있다. 철문은 문틀의 경첩으로 지지되며 지붕 부재는 벽과 트러스에 연결된다. 떠 있거나 지지 관계가 불가능한 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.029,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.779,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)",
     "[gpt-high] 감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ],
    "B": [
     "[gpt-high] 오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 779
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "금지된 식별 가능한 텍스트가 없으며, 열린 문 너머로 레퍼런스의 급수탑을 배치하여 지정된 로케이션을 충실히 구현했습니다.  ★위반: [gpt-high] 오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
   },
   {
    "label": "A",
    "score": 779,
    "verdict_ko": "프롬프트에서 명시적으로 금지한 읽을 수 있는 텍스트가 포함되어 심각한 위반에 해당합니다.  ★위반: [gemini-pro] 읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀) / [gpt-high] 감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_ruined_village_square_sel.png",
    "asset_id": "e716b949-d46e-4479-90bb-be052601519b",
    "role": "location_seed_bg"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-4199-788d-9b0a-adc1eebd9338",
  "ref_mode": "seed-bg만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_bypass:bg_only"
 },
 "S62sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:48:18.117363+00:00",
  "fingerprint": "d3ce69862ad9cd5c243e8ef886d8b4d3412f4a6754c427b0bc67c438f31ec097",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S62sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S62sh3_sel.png",
  "source_sha256": "34406e5983bcdb55e4cbfc76daf6f35a73c8476c960789fdf2a1edf78b8a7d18",
  "file": "S62sh3_cine.png",
  "staged_sha256": "3774df59c680a458c61476915a20248bc75c674602f456bf2f5e87176f106f3d",
  "latency_ms": 11921
 },
 "S62sh15::signage": {
  "fp": "c230a55faf0cfd98",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S62sh15": {
  "input_fingerprint": "54f104ac62589c2a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 투사된 빛나는 홀로그램 지도 위로 서남쪽 위치의 붉은 점이 밝게 켜져 있는 찰리의 시점 쇼트.\n\nLOCATION (lock): On the ground of the village's open square, beside the old truck being repaired. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: navigation map projected onto the ground in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 바닥에 투사된 홀로그램 지도 (Active, with the southwest destination marked by a bright red point) — The map-bearing ground surface faces upward and is viewed steeply from above, with the southwest marker in the map's lower-left sector; used as Carries 찰리's navigation information while retaining visible ground around the projection; 공터 바닥 (Visible beneath and around the projected map); used as Establishes the projection's physical placement without adding interface borders or device hardware.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain daytime ambient visibility while allowing the projected map and its bright red southwest marker to read as localized emitted light without darkening the entire setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains in the village square during repair attempts, and Charlie retains his battle damage and malfunctioning sensors. Charlie's POV contains a navigation map identifying a department store and supermarket 25.6 km to the southwest, not a floor-projected hologram.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 투사된 빛나는 홀로그램 지도 위로 서남쪽 위치의 붉은 점이 밝게 켜져 있는 찰리의 시점 쇼트.\n\nLOCATION (lock): On the ground of the village's open square, beside the old truck being repaired. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: navigation map projected onto the ground in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 바닥에 투사된 홀로그램 지도 (Active, with the southwest destination marked by a bright red point) — The map-bearing ground surface faces upward and is viewed steeply from above, with the southwest marker in the map's lower-left sector; used as Carries 찰리's navigation information while retaining visible ground around the projection; 공터 바닥 (Visible beneath and around the projected map); used as Establishes the projection's physical placement without adding interface borders or device hardware.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain daytime ambient visibility while allowing the projected map and its bright red southwest marker to read as localized emitted light without darkening the entire setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains in the village square during repair attempts, and Charlie retains his battle damage and malfunctioning sensors. Charlie's POV contains a navigation map identifying a department store and supermarket 25.6 km to the southwest, not a floor-projected hologram.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 투사된 빛나는 홀로그램 지도 위로 서남쪽 위치의 붉은 점이 밝게 켜져 있는 찰리의 시점 쇼트.\n\nLOCATION (lock): On the ground of the village's open square, beside the old truck being repaired. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: navigation map projected onto the ground in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 바닥에 투사된 홀로그램 지도 (Active, with the southwest destination marked by a bright red point) — The map-bearing ground surface faces upward and is viewed steeply from above, with the southwest marker in the map's lower-left sector; used as Carries 찰리's navigation information while retaining visible ground around the projection; 공터 바닥 (Visible beneath and around the projected map); used as Establishes the projection's physical placement without adding interface borders or device hardware.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain daytime ambient visibility while allowing the projected map and its bright red southwest marker to read as localized emitted light without darkening the entire setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains in the village square during repair attempts, and Charlie retains his battle damage and malfunctioning sensors. Charlie's POV contains a navigation map identifying a department store and supermarket 25.6 km to the southwest, not a floor-projected hologram.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 바닥을 비스듬히 내려다보고 있으며, 시선은 중앙의 홀로그램 지도를 향함.",
    "built_space": "흙바닥의 공터이며, 배경에 수리 중인 낡은 트럭과 공구들이 놓여 있음.",
    "entities": "바닥의 질감이 비치는 홀로그램 지도, 지도 좌측 하단(서남쪽)의 밝은 붉은 점, 낡은 트럭 모두 지시사항에 부합함.",
    "hard_violations": [
     "[gpt-high] 장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
    ],
    "physics": "홀로그램이 바닥의 굴곡을 따라 자연스럽게 투사되어 있음."
   },
   {
    "label": "B",
    "direction": "1인칭 시점으로 카메라는 바닥의 홀로그램 지도를 내려다보고 있음.",
    "built_space": "바닥에 갈라진 틈이 있는 공터이며, 우측 상단에 트럭의 바퀴가 보임.",
    "entities": "홀로그램 지도와 붉은 점이 있으나, 지도 위에 프롬프트 텍스트가 그대로 노출됨. 화자의 팔과 장갑이 보임.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨.",
     "[gpt-high] 지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
     "[gpt-high] 명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
    ],
    "physics": "홀로그램이 평평하게 투사되었고, 화자의 팔은 자연스러운 자세를 취하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지시된 시점과 홀로그램 지도의 붉은 점 위치를 정확히 묘사했으며, 금지된 텍스트 노출 없이 프롬프트를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트의 지시문이 그대로 이미지 안에 읽을 수 있는 텍스트로 유출되는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 바닥을 비스듬히 내려다보고 있으며, 시선은 중앙의 홀로그램 지도를 향함.",
        "built_space": "흙바닥의 공터이며, 배경에 수리 중인 낡은 트럭과 공구들이 놓여 있음.",
        "entities": "바닥의 질감이 비치는 홀로그램 지도, 지도 좌측 하단(서남쪽)의 밝은 붉은 점, 낡은 트럭 모두 지시사항에 부합함.",
        "hard_violations": [],
        "physics": "홀로그램이 바닥의 굴곡을 따라 자연스럽게 투사되어 있음."
       },
       {
        "label": "B",
        "direction": "1인칭 시점으로 카메라는 바닥의 홀로그램 지도를 내려다보고 있음.",
        "built_space": "바닥에 갈라진 틈이 있는 공터이며, 우측 상단에 트럭의 바퀴가 보임.",
        "entities": "홀로그램 지도와 붉은 점이 있으나, 지도 위에 프롬프트 텍스트가 그대로 노출됨. 화자의 팔과 장갑이 보임.",
        "hard_violations": [
         "읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨."
        ],
        "physics": "홀로그램이 평평하게 투사되었고, 화자의 팔은 자연스러운 자세를 취하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지시된 시점과 홀로그램 지도의 붉은 점 위치를 정확히 묘사했으며, 금지된 텍스트 노출 없이 프롬프트를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트의 지시문이 그대로 이미지 안에 읽을 수 있는 텍스트로 유출되는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 바닥을 비스듬히 내려다보고 있으며, 시선은 중앙의 홀로그램 지도를 향함.",
        "built_space": "흙바닥의 공터이며, 배경에 수리 중인 낡은 트럭과 공구들이 놓여 있음.",
        "entities": "바닥의 질감이 비치는 홀로그램 지도, 지도 좌측 하단(서남쪽)의 밝은 붉은 점, 낡은 트럭 모두 지시사항에 부합함.",
        "hard_violations": [],
        "physics": "홀로그램이 바닥의 굴곡을 따라 자연스럽게 투사되어 있음."
       },
       {
        "label": "B",
        "direction": "1인칭 시점으로 카메라는 바닥의 홀로그램 지도를 내려다보고 있음.",
        "built_space": "바닥에 갈라진 틈이 있는 공터이며, 우측 상단에 트럭의 바퀴가 보임.",
        "entities": "홀로그램 지도와 붉은 점이 있으나, 지도 위에 프롬프트 텍스트가 그대로 노출됨. 화자의 팔과 장갑이 보임.",
        "hard_violations": [
         "읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨."
        ],
        "physics": "홀로그램이 평평하게 투사되었고, 화자의 팔은 자연스러운 자세를 취하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "서남쪽 붉은 점과 내려다보는 구도는 맞지만, 읽을 수 있는 설명문과 지도 테두리·조작 아이콘이 명시적 금지사항을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지면이 비치는 발광 지도와 좌하단 붉은 점은 더 충실하지만, 불필요한 전경 기계장치와 넓고 낮아진 시점 때문에 요구된 클로즈업을 완전히 실현하지 못한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 바닥의 지도를 비교적 가파르게 내려다본다. 밝은 붉은 점은 지도 좌하단에 있어 지정된 서남쪽 위치와 맞는다. 지도 면은 위를 향하며 관찰자에게 보인다. 인물의 눈이나 무기는 보이지 않는다.",
        "built_space": "갈라진 포장 바닥 중앙에 직사각형 지도 하나가 있고, 오른쪽 위에는 트럭 바퀴 하나와 그 뒤의 받침이 보인다. 바퀴 주변에는 공구와 부품이 여러 개 놓여 있다. 지도에 청록색 외곽 테두리와 왼쪽 조작 아이콘 열이 추가되어, 경계 없는 바닥 투사라는 요구와 다르다. 전경의 팔이 바닥 일부를 가린다.",
        "entities": "발광 지도, 좌하단 붉은 점, 공터 바닥, 낡은 트럭 일부가 보인다. 지도에는 상점 표식과 거리 숫자가 있지만, 설명 문구와 연도까지 선명하게 읽혀 글자 금지를 어긴다. 찢어진 소매와 장갑 낀 팔은 보이나 찰리의 정체나 센서 고장을 확인할 수는 없다. 인물 얼굴은 없고 별도의 참조 이미지도 제공되지 않았다.",
        "hard_violations": [
         "지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
         "명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
        ],
        "physics": "지도는 포장 바닥에 밀착된 빛으로 표현되며, 붉은 점 주변에는 국소적인 발광이 있다. 바퀴와 공구는 지면에 놓여 있다. 전경 팔은 화면 밖 신체로 이어지는 구도이며, 떠 있는 독립 신체로 보이지 않는다. 다만 불투명한 위성사진 면과 선명한 설명문은 실제 지면 투사보다 화면 합성처럼 보인다."
       },
       {
        "label": "B",
        "direction": "카메라는 흙바닥 지도를 비스듬히 내려다보고, 지도 좌하단의 밝은 붉은 점으로 밝은 경로가 이어진다. 목적지 위치는 서남쪽 요구에 맞는다. 다만 먼 트럭과 배경까지 보이는 각도여서 요구된 가파른 하향 시점보다 낮다. 인물의 시선이나 무기는 없다.",
        "built_space": "흙과 자갈, 타이어 자국이 있는 공터 중앙에 지도 하나가 투사되어 있다. 위쪽 배경에는 트럭 한 대와 보이는 바퀴 두 개, 그 앞의 수리 도구들이 있다. 지도 주변 지면이 충분히 보이고 뚜렷한 인터페이스 외곽 틀은 없다. 그러나 아래쪽과 오른쪽 전경에 금속 프레임과 배선이 크게 들어와 장치 하드웨어 없는 구도와 충돌한다.",
        "entities": "투명한 도로 지도, 좌하단 붉은 목적지 점, 공터 바닥, 수리 중인 것으로 보이는 낡은 트럭이 있다. 지도에는 작은 문자형 표식들이 있지만 백화점·슈퍼마켓과 25.6km라는 정보는 확실히 판독되지 않는다. 사람은 보이지 않는다. 전경의 손상된 듯한 기계 부품을 찰리의 센서라고 확정할 근거는 없다.",
        "hard_violations": [
         "장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
        ],
        "physics": "지도 선 아래로 흙의 질감이 비치고 붉은 빛이 주변 지면에 번져 바닥 투사로 읽힌다. 트럭 바퀴와 수리 도구는 지면에 놓여 있다. 전경 배선은 화면 가장자리의 기계 구조에 연결되어 있으며, 지지 없이 떠 있는 물체나 인체는 보이지 않는다. 낮의 주변 밝기도 유지된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "서남쪽 붉은 점과 내려다보는 구도는 맞지만, 읽을 수 있는 설명문과 지도 테두리·조작 아이콘이 명시적 금지사항을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지면이 비치는 발광 지도와 좌하단 붉은 점은 더 충실하지만, 불필요한 전경 기계장치와 넓고 낮아진 시점 때문에 요구된 클로즈업을 완전히 실현하지 못한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 바닥의 지도를 비교적 가파르게 내려다본다. 밝은 붉은 점은 지도 좌하단에 있어 지정된 서남쪽 위치와 맞는다. 지도 면은 위를 향하며 관찰자에게 보인다. 인물의 눈이나 무기는 보이지 않는다.",
        "built_space": "갈라진 포장 바닥 중앙에 직사각형 지도 하나가 있고, 오른쪽 위에는 트럭 바퀴 하나와 그 뒤의 받침이 보인다. 바퀴 주변에는 공구와 부품이 여러 개 놓여 있다. 지도에 청록색 외곽 테두리와 왼쪽 조작 아이콘 열이 추가되어, 경계 없는 바닥 투사라는 요구와 다르다. 전경의 팔이 바닥 일부를 가린다.",
        "entities": "발광 지도, 좌하단 붉은 점, 공터 바닥, 낡은 트럭 일부가 보인다. 지도에는 상점 표식과 거리 숫자가 있지만, 설명 문구와 연도까지 선명하게 읽혀 글자 금지를 어긴다. 찢어진 소매와 장갑 낀 팔은 보이나 찰리의 정체나 센서 고장을 확인할 수는 없다. 인물 얼굴은 없고 별도의 참조 이미지도 제공되지 않았다.",
        "hard_violations": [
         "지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
         "명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
        ],
        "physics": "지도는 포장 바닥에 밀착된 빛으로 표현되며, 붉은 점 주변에는 국소적인 발광이 있다. 바퀴와 공구는 지면에 놓여 있다. 전경 팔은 화면 밖 신체로 이어지는 구도이며, 떠 있는 독립 신체로 보이지 않는다. 다만 불투명한 위성사진 면과 선명한 설명문은 실제 지면 투사보다 화면 합성처럼 보인다."
       },
       {
        "label": "A",
        "direction": "카메라는 흙바닥 지도를 비스듬히 내려다보고, 지도 좌하단의 밝은 붉은 점으로 밝은 경로가 이어진다. 목적지 위치는 서남쪽 요구에 맞는다. 다만 먼 트럭과 배경까지 보이는 각도여서 요구된 가파른 하향 시점보다 낮다. 인물의 시선이나 무기는 없다.",
        "built_space": "흙과 자갈, 타이어 자국이 있는 공터 중앙에 지도 하나가 투사되어 있다. 위쪽 배경에는 트럭 한 대와 보이는 바퀴 두 개, 그 앞의 수리 도구들이 있다. 지도 주변 지면이 충분히 보이고 뚜렷한 인터페이스 외곽 틀은 없다. 그러나 아래쪽과 오른쪽 전경에 금속 프레임과 배선이 크게 들어와 장치 하드웨어 없는 구도와 충돌한다.",
        "entities": "투명한 도로 지도, 좌하단 붉은 목적지 점, 공터 바닥, 수리 중인 것으로 보이는 낡은 트럭이 있다. 지도에는 작은 문자형 표식들이 있지만 백화점·슈퍼마켓과 25.6km라는 정보는 확실히 판독되지 않는다. 사람은 보이지 않는다. 전경의 손상된 듯한 기계 부품을 찰리의 센서라고 확정할 근거는 없다.",
        "hard_violations": [
         "장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
        ],
        "physics": "지도 선 아래로 흙의 질감이 비치고 붉은 빛이 주변 지면에 번져 바닥 투사로 읽힌다. 트럭 바퀴와 수리 도구는 지면에 놓여 있다. 전경 배선은 화면 가장자리의 기계 구조에 연결되어 있으며, 지지 없이 떠 있는 물체나 인체는 보이지 않는다. 낮의 주변 밝기도 유지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.5
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨.",
     "[gpt-high] 지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
     "[gpt-high] 명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
    ],
    "A": [
     "[gpt-high] 장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지시된 시점과 홀로그램 지도의 붉은 점 위치를 정확히 묘사했으며, 금지된 텍스트 노출 없이 프롬프트를 잘 따랐습니다.  ★위반: [gpt-high] 장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
   },
   {
    "label": "B",
    "score": 250,
    "verdict_ko": "프롬프트의 지시문이 그대로 이미지 안에 읽을 수 있는 텍스트로 유출되는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨. / [gpt-high] 지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다. / [gpt-high] 명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-434f-7a58-afba-556530a7f138",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S62sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:49:11.399577+00:00",
  "fingerprint": "894de6c2ff9503a055b819b25203d8b0eed713492aae5a377de71af2dd7e1859",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S62sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S62sh15_sel.png",
  "source_sha256": "c3d2fe42bc3bdb3e666265c1dde589f700b5edadb32f227e59b3d942581ba402",
  "file": "S62sh15_cine.png",
  "staged_sha256": "f5f5e48cc691b37f8c29c63576332017de2600b607b1eced6dfd6e29e31c59e1",
  "latency_ms": 13116
 },
 "S62sh18::signage": {
  "fp": "f3f240bb7116c54f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S62sh18": {
  "input_fingerprint": "caf6be1335bd6c7c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 자신감에 찬 표정으로 한쪽 입꼬리를 올린 채 미소 짓고 있는 현우의 얼굴 클로즈업.\n\nLOCATION (lock): In the village's open square near the stalled old truck, where the food-search plan is discussed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Village clearing (The conversation remains in the clearing beside the truck); used as A narrow, softly resolved background margin preserves the location without competing with the half-smile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the open clearing maintains restrained contrast and natural facial detail without turning the confident smile into a heroic lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains the vehicle being prepared in the village square, with an old football in use nearby. Charlie's punctured, dented body and sensor damage have not been repaired. 현우: He stands in the village square, retaining his battle injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 자신감에 찬 표정으로 한쪽 입꼬리를 올린 채 미소 짓고 있는 현우의 얼굴 클로즈업.\n\nLOCATION (lock): In the village's open square near the stalled old truck, where the food-search plan is discussed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Village clearing (The conversation remains in the clearing beside the truck); used as A narrow, softly resolved background margin preserves the location without competing with the half-smile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the open clearing maintains restrained contrast and natural facial detail without turning the confident smile into a heroic lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains the vehicle being prepared in the village square, with an old football in use nearby. Charlie's punctured, dented body and sensor damage have not been repaired. 현우: He stands in the village square, retaining his battle injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 자신감에 찬 표정으로 한쪽 입꼬리를 올린 채 미소 짓고 있는 현우의 얼굴 클로즈업.\n\nLOCATION (lock): In the village's open square near the stalled old truck, where the food-search plan is discussed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Village clearing (The conversation remains in the clearing beside the truck); used as A narrow, softly resolved background margin preserves the location without competing with the half-smile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the open clearing maintains restrained contrast and natural facial detail without turning the confident smile into a heroic lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains the vehicle being prepared in the village square, with an old football in use nearby. Charlie's punctured, dented body and sensor damage have not been repaired. 현우: He stands in the village square, retaining his battle injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선이 카메라 렌즈를 정면으로 향하고 있음.",
    "built_space": "흙으로 된 마을 공터. 배경 우측에 낡은 트럭과 공구들이 흐릿하게 배치되어 있음.",
    "entities": "현우(참조와 일치하는 앳된 얼굴, 헝클어진 머리, 회색 셔츠, 얼굴의 흙먼지와 상처). 양쪽 입꼬리가 대칭적으로 올라간 옅은 미소.",
    "hard_violations": [],
    "physics": "땅에 발을 딛고 안정적인 자세로 서 있음."
   },
   {
    "label": "B",
    "direction": "시선이 카메라를 벗어나 화면 좌측(캐릭터의 우측)을 향하고 있음.",
    "built_space": "흙으로 된 마을 공터. 배경 좌측에 낡은 트럭이 흐릿하게 보임.",
    "entities": "현우(참조와 일치하는 외모, 복장, 얼굴의 상처와 오염). 캐릭터 기준 오른쪽 입꼬리가 확연히 더 높게 올라간 비대칭 미소(썩소)를 짓고 있음.",
    "hard_violations": [],
    "physics": "땅에 발을 딛고 안정적인 자세로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "자연스러운 시선 처리와 함께 텍스트가 명시한 '한쪽 입꼬리를 올린 채 미소 짓는' 표정을 정확히 비대칭적으로 구현하여 프롬프트 충실도가 매우 높습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "캐릭터와 배경의 일치도는 높으나, 한쪽 입꼬리만 올리라는 구체적인 표정 지시와 달리 대칭적인 미소를 짓고 있으며 카메라 렌즈를 정면으로 응시해 다소 어색합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있음.",
        "built_space": "흙으로 된 마을 공터. 배경 우측에 낡은 트럭과 공구들이 흐릿하게 배치되어 있음.",
        "entities": "현우(참조와 일치하는 앳된 얼굴, 헝클어진 머리, 회색 셔츠, 얼굴의 흙먼지와 상처). 양쪽 입꼬리가 대칭적으로 올라간 옅은 미소.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 안정적인 자세로 서 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라를 벗어나 화면 좌측(캐릭터의 우측)을 향하고 있음.",
        "built_space": "흙으로 된 마을 공터. 배경 좌측에 낡은 트럭이 흐릿하게 보임.",
        "entities": "현우(참조와 일치하는 외모, 복장, 얼굴의 상처와 오염). 캐릭터 기준 오른쪽 입꼬리가 확연히 더 높게 올라간 비대칭 미소(썩소)를 짓고 있음.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 안정적인 자세로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "자연스러운 시선 처리와 함께 텍스트가 명시한 '한쪽 입꼬리를 올린 채 미소 짓는' 표정을 정확히 비대칭적으로 구현하여 프롬프트 충실도가 매우 높습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "캐릭터와 배경의 일치도는 높으나, 한쪽 입꼬리만 올리라는 구체적인 표정 지시와 달리 대칭적인 미소를 짓고 있으며 카메라 렌즈를 정면으로 응시해 다소 어색합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있음.",
        "built_space": "흙으로 된 마을 공터. 배경 우측에 낡은 트럭과 공구들이 흐릿하게 배치되어 있음.",
        "entities": "현우(참조와 일치하는 앳된 얼굴, 헝클어진 머리, 회색 셔츠, 얼굴의 흙먼지와 상처). 양쪽 입꼬리가 대칭적으로 올라간 옅은 미소.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 안정적인 자세로 서 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라를 벗어나 화면 좌측(캐릭터의 우측)을 향하고 있음.",
        "built_space": "흙으로 된 마을 공터. 배경 좌측에 낡은 트럭이 흐릿하게 보임.",
        "entities": "현우(참조와 일치하는 외모, 복장, 얼굴의 상처와 오염). 캐릭터 기준 오른쪽 입꼬리가 확연히 더 높게 올라간 비대칭 미소(썩소)를 짓고 있음.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 안정적인 자세로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴을 더 크게 잡고 한쪽 입꼬리가 올라간 자신감 있는 미소를 살려 우세하나, 왼쪽 트럭이 차지하는 배경 면적은 요구보다 크다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 외모와 부상·의상은 잘 이어지지만, 어깨와 광장까지 넓게 담은 구도와 비교적 대칭적인 미소가 얼굴 클로즈업 및 한쪽 입꼬리 연기를 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이며 눈은 카메라보다 화면 왼쪽의 가까운 화면 밖 지점을 향한다. 바라보는 상대는 보이지 않으며 지시문도 특정 상대를 지정하지 않는다. 화면 왼쪽 입꼬리가 조금 더 올라가 비대칭적인 미소가 읽힌다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "화면 왼쪽 뒤에 오래된 트럭 한 대가 있고, 주변에 낮은 건물 일부와 흙바닥, 흐릿한 정비 도구들이 보인다. 현우는 트럭보다 앞에 있으며 공간 관계는 가능하다. 얼굴과 목이 화면 높이 대부분을 차지하지만 트럭이 왼쪽 배경을 크게 채워 ‘좁은 배경 여백’에는 덜 충실하다. 건물의 세부 고정 설비는 흐림 때문에 대조할 수 없고, 중복 차량이나 불가능한 반사는 없다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 한국계 청년 남성으로 보이는 얼굴, 헝클어진 검은 머리, 회색 셔츠는 인물 참조와 대체로 맞는다. 참조보다 다소 성숙한 인상이지만 눈·코·턱의 형태는 유사하다. 이마와 뺨, 입 주변의 상처와 오염으로 전투 부상이 유지된다. 낡은 트럭도 보인다. 축구공과 찰리는 이 얼굴 중심 구도에서 식별되지 않으며, 하의와 휴대 상태 역시 평가 범위 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 있고 셔츠는 몸 위에 정상적으로 걸쳐져 있다. 발은 프레임 밖이므로 지면 접촉은 확인할 수 없지만 부유를 시사하는 모습은 없다. 트럭은 보이는 바퀴로 지면에 놓여 있으며, 떠 있거나 지지 없이 들린 물체는 없다."
       },
       {
        "label": "B",
        "direction": "얼굴과 시선이 거의 렌즈 정면을 향한다. 무표정은 아니지만 대화 중 화면 밖 상대를 보는 순간보다 카메라를 의식한 미소에 가깝다. 양쪽 입꼬리가 비슷하게 올라가 있고 화면 오른쪽이 약간 높아, 요구한 한쪽 입꼬리의 강조는 약하다. 방향을 확인할 무기나 휴대 소품은 없다.",
        "built_space": "현우 뒤 오른쪽에 트럭 한 대가 있고 적재함은 왼쪽으로 이어진다. 낮은 건물들, 넓은 흙바닥과 트럭 앞 정비 도구가 흐릿하게 보인다. 현우가 트럭 앞 광장에 서 있는 배치는 자연스럽고 차량 중복은 없다. 다만 머리 전체와 어깨·윗가슴을 담아 배경 광장의 면적이 커졌으며, 얼굴에 밀착한 구도와 좁은 배경 여백 요구에서 A보다 멀다. 세부 설비나 반사는 판독할 수 없다.",
        "entities": "현우로 보이는 청년 남성 한 명만 등장한다. 앳된 동아시아계 얼굴, 검은 헝클어진 머리와 낡은 회색 셔츠가 참조에 부합하며, 국적 자체는 외모로 확인할 수 없다. 얼굴의 긁힘과 먼지, 셔츠의 마모가 부상 및 복장 연속성을 뒷받침한다. 배경의 낡은 트럭도 요구한 차량 유형에 맞는다. 축구공과 찰리는 식별되지 않지만 얼굴 클로즈업에서 반드시 보여야 할 대상은 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "목이 머리를 받치고 어깨와 상체가 정상적으로 연결되어 있으며 옷의 주름도 몸을 따른다. 하체는 보이지 않아 발의 지지는 판단할 수 없고, 공중에 떠 있다는 징후도 없다. 트럭은 바퀴로 흙바닥에 지지되어 있다. 비정상적인 관절이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴을 더 크게 잡고 한쪽 입꼬리가 올라간 자신감 있는 미소를 살려 우세하나, 왼쪽 트럭이 차지하는 배경 면적은 요구보다 크다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "현우의 외모와 부상·의상은 잘 이어지지만, 어깨와 광장까지 넓게 담은 구도와 비교적 대칭적인 미소가 얼굴 클로즈업 및 한쪽 입꼬리 연기를 약화한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이며 눈은 카메라보다 화면 왼쪽의 가까운 화면 밖 지점을 향한다. 바라보는 상대는 보이지 않으며 지시문도 특정 상대를 지정하지 않는다. 화면 왼쪽 입꼬리가 조금 더 올라가 비대칭적인 미소가 읽힌다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "화면 왼쪽 뒤에 오래된 트럭 한 대가 있고, 주변에 낮은 건물 일부와 흙바닥, 흐릿한 정비 도구들이 보인다. 현우는 트럭보다 앞에 있으며 공간 관계는 가능하다. 얼굴과 목이 화면 높이 대부분을 차지하지만 트럭이 왼쪽 배경을 크게 채워 ‘좁은 배경 여백’에는 덜 충실하다. 건물의 세부 고정 설비는 흐림 때문에 대조할 수 없고, 중복 차량이나 불가능한 반사는 없다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 한국계 청년 남성으로 보이는 얼굴, 헝클어진 검은 머리, 회색 셔츠는 인물 참조와 대체로 맞는다. 참조보다 다소 성숙한 인상이지만 눈·코·턱의 형태는 유사하다. 이마와 뺨, 입 주변의 상처와 오염으로 전투 부상이 유지된다. 낡은 트럭도 보인다. 축구공과 찰리는 이 얼굴 중심 구도에서 식별되지 않으며, 하의와 휴대 상태 역시 평가 범위 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 있고 셔츠는 몸 위에 정상적으로 걸쳐져 있다. 발은 프레임 밖이므로 지면 접촉은 확인할 수 없지만 부유를 시사하는 모습은 없다. 트럭은 보이는 바퀴로 지면에 놓여 있으며, 떠 있거나 지지 없이 들린 물체는 없다."
       },
       {
        "label": "A",
        "direction": "얼굴과 시선이 거의 렌즈 정면을 향한다. 무표정은 아니지만 대화 중 화면 밖 상대를 보는 순간보다 카메라를 의식한 미소에 가깝다. 양쪽 입꼬리가 비슷하게 올라가 있고 화면 오른쪽이 약간 높아, 요구한 한쪽 입꼬리의 강조는 약하다. 방향을 확인할 무기나 휴대 소품은 없다.",
        "built_space": "현우 뒤 오른쪽에 트럭 한 대가 있고 적재함은 왼쪽으로 이어진다. 낮은 건물들, 넓은 흙바닥과 트럭 앞 정비 도구가 흐릿하게 보인다. 현우가 트럭 앞 광장에 서 있는 배치는 자연스럽고 차량 중복은 없다. 다만 머리 전체와 어깨·윗가슴을 담아 배경 광장의 면적이 커졌으며, 얼굴에 밀착한 구도와 좁은 배경 여백 요구에서 A보다 멀다. 세부 설비나 반사는 판독할 수 없다.",
        "entities": "현우로 보이는 청년 남성 한 명만 등장한다. 앳된 동아시아계 얼굴, 검은 헝클어진 머리와 낡은 회색 셔츠가 참조에 부합하며, 국적 자체는 외모로 확인할 수 없다. 얼굴의 긁힘과 먼지, 셔츠의 마모가 부상 및 복장 연속성을 뒷받침한다. 배경의 낡은 트럭도 요구한 차량 유형에 맞는다. 축구공과 찰리는 식별되지 않지만 얼굴 클로즈업에서 반드시 보여야 할 대상은 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "목이 머리를 받치고 어깨와 상체가 정상적으로 연결되어 있으며 옷의 주름도 몸을 따른다. 하체는 보이지 않아 발의 지지는 판단할 수 없고, 공중에 떠 있다는 징후도 없다. 트럭은 바퀴로 흙바닥에 지지되어 있다. 비정상적인 관절이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.653,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.653,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1653
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "자연스러운 시선 처리와 함께 텍스트가 명시한 '한쪽 입꼬리를 올린 채 미소 짓는' 표정을 정확히 비대칭적으로 구현하여 프롬프트 충실도가 매우 높습니다."
   },
   {
    "label": "A",
    "score": 1653,
    "verdict_ko": "캐릭터와 배경의 일치도는 높으나, 한쪽 입꼬리만 올리라는 구체적인 표정 지시와 달리 대칭적인 미소를 짓고 있으며 카메라 렌즈를 정면으로 응시해 다소 어색합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S62sh15_sel.png",
    "asset_id": "7350440c-3f64-4d57-b9de-c162a9138d92",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-44f6-78f0-87a9-d7bd1c6346a8",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S62sh15"
  }
 },
 "S62sh18::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:58:54.031461+00:00",
  "fingerprint": "b91206caff855c639bc4ddbec16d4436e02633477d22cfd76f3dc55700dd8944",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S62sh18_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S62sh18_sel.png",
  "source_sha256": "d59edd80f0e49643a7d5a57ae51d14923fd458e9fa69f42e39028c3c34322a70",
  "file": "S62sh18_cine.png",
  "staged_sha256": "b8a0dc159aa17a8277c33610b24cedc0125201ba0ab0a61fde4d145c43599703",
  "latency_ms": 10383
 },
 "S63sh1::signage": {
  "fp": "f28921543ee49df5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S63sh1": {
  "input_fingerprint": "3722c1480235373d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 밝은 낮, 황량한 흙먼지 도로 위를 달리며 바퀴 뒤로 흙먼지를 일으키고 있는 낡은 트럭의 외관 전경.\n\nLOCATION (lock): On a dusty rural road outside the village, where the old supply truck travels through the exposed landscape. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old truck (Traveling along the road and raising dust behind its wheels) — The rear and passenger-side exterior are visible from an elevated rear-quarter angle; used as The moving spatial anchor, kept modest in size within the surrounding road; Barren dirt road (Extends ahead of the traveling truck) — Runs from the lower foreground past the truck toward the upper-center distance; used as Establishes forward travel and provides the visual direction to preserve in the subsequent cab cut; Wheel-raised dust (Trailing behind the moving truck); used as Makes the vehicle's movement readable without obscuring its exterior.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright daytime light makes the barren road and wheel-raised dust legible with controlled contrast and an understated tonal range.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck is traveling along the road in daylight.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 밝은 낮, 황량한 흙먼지 도로 위를 달리며 바퀴 뒤로 흙먼지를 일으키고 있는 낡은 트럭의 외관 전경.\n\nLOCATION (lock): On a dusty rural road outside the village, where the old supply truck travels through the exposed landscape. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old truck (Traveling along the road and raising dust behind its wheels) — The rear and passenger-side exterior are visible from an elevated rear-quarter angle; used as The moving spatial anchor, kept modest in size within the surrounding road; Barren dirt road (Extends ahead of the traveling truck) — Runs from the lower foreground past the truck toward the upper-center distance; used as Establishes forward travel and provides the visual direction to preserve in the subsequent cab cut; Wheel-raised dust (Trailing behind the moving truck); used as Makes the vehicle's movement readable without obscuring its exterior.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright daytime light makes the barren road and wheel-raised dust legible with controlled contrast and an understated tonal range.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck is traveling along the road in daylight.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 밝은 낮, 황량한 흙먼지 도로 위를 달리며 바퀴 뒤로 흙먼지를 일으키고 있는 낡은 트럭의 외관 전경.\n\nLOCATION (lock): On a dusty rural road outside the village, where the old supply truck travels through the exposed landscape. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old truck (Traveling along the road and raising dust behind its wheels) — The rear and passenger-side exterior are visible from an elevated rear-quarter angle; used as The moving spatial anchor, kept modest in size within the surrounding road; Barren dirt road (Extends ahead of the traveling truck) — Runs from the lower foreground past the truck toward the upper-center distance; used as Establishes forward travel and provides the visual direction to preserve in the subsequent cab cut; Wheel-raised dust (Trailing behind the moving truck); used as Makes the vehicle's movement readable without obscuring its exterior.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright daytime light makes the barren road and wheel-raised dust legible with controlled contrast and an understated tonal range.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck is traveling along the road in daylight.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "트럭이 화면 중앙 먼 곳을 향해 주행하며 바퀴 뒤로 흙먼지를 일으키고 있습니다.",
    "built_space": "참고 이미지의 좁은 바위 지형이 배경에 배치되었으나, 바위의 하단부가 도로 지면과 이어지지 않고 일직선으로 잘려 있습니다.",
    "entities": "낡은 트럭, 흙먼지가 날리는 황량한 비포장 도로는 프롬프트의 설명과 일치합니다. 인물은 등장하지 않습니다.",
    "hard_violations": [
     "[gemini-pro] 콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
    ],
    "physics": "트럭과 흙먼지는 지면에 닿아 있으나, 배경의 바위는 지지되는 바닥 형태 없이 지평선 위에 부자연스럽게 떠 있습니다."
   },
   {
    "label": "B",
    "direction": "트럭이 화면 전방의 흙길을 따라 주행 중입니다.",
    "built_space": "도로 좌측에 참고 이미지의 바위 지형이 얹혀 있으며 주변 풍경과 어색하게 결합되어 있습니다.",
    "entities": "트럭과 흙길은 존재하나, 트럭의 후면에 '낡은 트럭'이라는 명확한 한글 텍스트가 노출되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
     "[gemini-pro] 콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)",
     "[gpt-high] 적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
    ],
    "physics": "트럭 아래 발생하는 흙먼지가 자연스럽게 흩날리지 않고 직사각형 형태의 평면적인 덩어리로 바닥 위에 떠 있어 물리적으로 불가능합니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트의 차량과 도로 묘사는 따랐으나, 참고 이미지를 평면적인 스티커처럼 배경에 그대로 오려 붙여 실사 요건을 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트가 명시적으로 금지한 텍스트('낡은 트럭')가 포함되었으며, 흙먼지가 직사각형 그래픽으로 조잡하게 합성되어 기각 대상입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭이 화면 중앙 먼 곳을 향해 주행하며 바퀴 뒤로 흙먼지를 일으키고 있습니다.",
        "built_space": "참고 이미지의 좁은 바위 지형이 배경에 배치되었으나, 바위의 하단부가 도로 지면과 이어지지 않고 일직선으로 잘려 있습니다.",
        "entities": "낡은 트럭, 흙먼지가 날리는 황량한 비포장 도로는 프롬프트의 설명과 일치합니다. 인물은 등장하지 않습니다.",
        "hard_violations": [
         "콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
        ],
        "physics": "트럭과 흙먼지는 지면에 닿아 있으나, 배경의 바위는 지지되는 바닥 형태 없이 지평선 위에 부자연스럽게 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "트럭이 화면 전방의 흙길을 따라 주행 중입니다.",
        "built_space": "도로 좌측에 참고 이미지의 바위 지형이 얹혀 있으며 주변 풍경과 어색하게 결합되어 있습니다.",
        "entities": "트럭과 흙길은 존재하나, 트럭의 후면에 '낡은 트럭'이라는 명확한 한글 텍스트가 노출되어 있습니다.",
        "hard_violations": [
         "금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
         "콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)"
        ],
        "physics": "트럭 아래 발생하는 흙먼지가 자연스럽게 흩날리지 않고 직사각형 형태의 평면적인 덩어리로 바닥 위에 떠 있어 물리적으로 불가능합니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트의 차량과 도로 묘사는 따랐으나, 참고 이미지를 평면적인 스티커처럼 배경에 그대로 오려 붙여 실사 요건을 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트가 명시적으로 금지한 텍스트('낡은 트럭')가 포함되었으며, 흙먼지가 직사각형 그래픽으로 조잡하게 합성되어 기각 대상입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "트럭이 화면 중앙 먼 곳을 향해 주행하며 바퀴 뒤로 흙먼지를 일으키고 있습니다.",
        "built_space": "참고 이미지의 좁은 바위 지형이 배경에 배치되었으나, 바위의 하단부가 도로 지면과 이어지지 않고 일직선으로 잘려 있습니다.",
        "entities": "낡은 트럭, 흙먼지가 날리는 황량한 비포장 도로는 프롬프트의 설명과 일치합니다. 인물은 등장하지 않습니다.",
        "hard_violations": [
         "콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
        ],
        "physics": "트럭과 흙먼지는 지면에 닿아 있으나, 배경의 바위는 지지되는 바닥 형태 없이 지평선 위에 부자연스럽게 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "트럭이 화면 전방의 흙길을 따라 주행 중입니다.",
        "built_space": "도로 좌측에 참고 이미지의 바위 지형이 얹혀 있으며 주변 풍경과 어색하게 결합되어 있습니다.",
        "entities": "트럭과 흙길은 존재하나, 트럭의 후면에 '낡은 트럭'이라는 명확한 한글 텍스트가 노출되어 있습니다.",
        "hard_violations": [
         "금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
         "콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)"
        ],
        "physics": "트럭 아래 발생하는 흙먼지가 자연스럽게 흩날리지 않고 직사각형 형태의 평면적인 덩어리로 바닥 위에 떠 있어 물리적으로 불가능합니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "후면·조수석 측면과 주행 먼지는 맞지만, 적재함의 읽을 수 있는 한글이 명시적 금지 사항을 위반하고 참조 바위가 과도하게 큰 전경을 차지한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "문자 없이 밝은 낮의 주행 와이드숏을 구현했지만, 조수석이 아닌 운전석 측면이 보이고 도로가 상단 중앙 대신 왼쪽으로 향하며 참조 바위의 규모와 연결이 부자연스럽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭 앞부분은 화면 오른쪽 위의 마을과 도로 소실점을 향한다. 먼지는 뒷바퀴에서 왼쪽 아래로 이어져 전진 방향과 맞는다. 후면과 오른쪽 조수석 측면이 보인다. 다만 도로의 목적 방향은 요구된 상단 중앙보다 오른쪽에 치우친다. 식별 가능한 인물이나 시선은 없다.",
        "built_space": "비포장도로 하나가 하단 전경에서 트럭을 지나 우측 상단 마을로 이어진다. 왼쪽에는 참조의 큰 좌측 바위, 우측 바위 덩어리, 위를 가로지르는 바위와 안쪽 돌출석으로 이루어진 틈이 재현되어 있다. 그러나 이 바위군이 화면 왼쪽 절반을 압도하며 도로변 지면과의 연결도 사진 조각을 붙인 듯 급격하다. 원경에는 여러 주택과 전신주가 있다. 참조만으로 이 추가 배치나 바위와 도로의 정확한 공간 관계는 확인되지 않는다.",
        "entities": "낡고 도장이 벗겨진 파란 소형 화물트럭 한 대, 흙길, 마른 주변 지형, 바퀴 뒤 흙먼지가 보인다. 적재함에는 자루와 묶인 짐이 있다. 인물은 식별되지 않아 장면 수준 인물 목록을 불필요하게 추가하지 않았다. 적재함 후면에는 읽을 수 있는 한글 표기가 있어 문자 금지 조건에 어긋난다. 참조 바위의 무늬와 틈 형태는 닮았지만 전체 장면 속 규모와 재질 연결은 자연스럽지 않다.",
        "hard_violations": [
         "적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 도로에 닿아 차체를 지지하며, 적재물은 적재함 바닥 위에 놓여 있다. 먼지는 바퀴 주변에서 일어나 차량 뒤로 퍼져 주행으로 발생한 것으로 읽힌다. 먼지가 하부 일부를 가리지만 트럭 외관은 대체로 보인다. 명백히 지지 없이 떠 있는 차량이나 짐은 없다."
       },
       {
        "label": "B",
        "direction": "트럭 앞부분은 화면 왼쪽 위로 이어지는 도로를 향하고 먼지는 오른쪽 아래 후방으로 흐른다. 전진과 먼지 방향은 일치한다. 그러나 보이는 것은 후면과 왼쪽 운전석 측면으로, 지정된 조수석 측면과 반대다. 도로도 상단 중앙이 아니라 왼쪽 원경으로 빠진다. 식별 가능한 사람의 시선은 없다.",
        "built_space": "도로 하나가 하단 오른쪽에서 트럭 아래를 지나 상단 왼쪽으로 굽는다. 주변은 메마른 흙과 돌, 드문 나무로 이루어져 있다. 상단 중앙에는 참조의 좁은 바위 틈과 이를 둘러싼 좌우 바위, 상부 가로 바위 및 안쪽 돌출석이 보인다. 다만 바위군이 먼 배경에 비해 지나치게 크고, 바닥 경계와 주변 지형의 연결이 부자연스럽다. 인공 건축물이나 고정 설비는 보이지 않는다.",
        "entities": "녹슨 파란 화물트럭 한 대와 비포장도로, 바퀴 뒤의 흙먼지가 확인된다. 적재함에는 낮게 놓인 밝은색 물체 일부가 보이지만 종류는 확정할 수 없다. 인물과 읽을 수 있는 문자는 보이지 않는다. 낡은 트럭과 황량한 낮 풍경은 요청에 맞으며, 바위 표면과 틈 형태는 참조를 따르지만 주변 풍경에 자연스럽게 통합되지는 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 앞뒤 타이어가 노면에 닿아 트럭을 지지한다. 적재함 안 물체도 바닥 위에 놓인 것으로 보인다. 흙먼지가 뒷바퀴 부근에서 솟아 뒤로 길게 퍼져 주행 동작을 설명한다. 먼지가 후면 하부를 일부 가리지만 차체 대부분은 드러난다. 지지 없이 공중에 떠 있는 물체나 불가능한 차체 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "후면·조수석 측면과 주행 먼지는 맞지만, 적재함의 읽을 수 있는 한글이 명시적 금지 사항을 위반하고 참조 바위가 과도하게 큰 전경을 차지한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "문자 없이 밝은 낮의 주행 와이드숏을 구현했지만, 조수석이 아닌 운전석 측면이 보이고 도로가 상단 중앙 대신 왼쪽으로 향하며 참조 바위의 규모와 연결이 부자연스럽다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "트럭 앞부분은 화면 오른쪽 위의 마을과 도로 소실점을 향한다. 먼지는 뒷바퀴에서 왼쪽 아래로 이어져 전진 방향과 맞는다. 후면과 오른쪽 조수석 측면이 보인다. 다만 도로의 목적 방향은 요구된 상단 중앙보다 오른쪽에 치우친다. 식별 가능한 인물이나 시선은 없다.",
        "built_space": "비포장도로 하나가 하단 전경에서 트럭을 지나 우측 상단 마을로 이어진다. 왼쪽에는 참조의 큰 좌측 바위, 우측 바위 덩어리, 위를 가로지르는 바위와 안쪽 돌출석으로 이루어진 틈이 재현되어 있다. 그러나 이 바위군이 화면 왼쪽 절반을 압도하며 도로변 지면과의 연결도 사진 조각을 붙인 듯 급격하다. 원경에는 여러 주택과 전신주가 있다. 참조만으로 이 추가 배치나 바위와 도로의 정확한 공간 관계는 확인되지 않는다.",
        "entities": "낡고 도장이 벗겨진 파란 소형 화물트럭 한 대, 흙길, 마른 주변 지형, 바퀴 뒤 흙먼지가 보인다. 적재함에는 자루와 묶인 짐이 있다. 인물은 식별되지 않아 장면 수준 인물 목록을 불필요하게 추가하지 않았다. 적재함 후면에는 읽을 수 있는 한글 표기가 있어 문자 금지 조건에 어긋난다. 참조 바위의 무늬와 틈 형태는 닮았지만 전체 장면 속 규모와 재질 연결은 자연스럽지 않다.",
        "hard_violations": [
         "적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 도로에 닿아 차체를 지지하며, 적재물은 적재함 바닥 위에 놓여 있다. 먼지는 바퀴 주변에서 일어나 차량 뒤로 퍼져 주행으로 발생한 것으로 읽힌다. 먼지가 하부 일부를 가리지만 트럭 외관은 대체로 보인다. 명백히 지지 없이 떠 있는 차량이나 짐은 없다."
       },
       {
        "label": "A",
        "direction": "트럭 앞부분은 화면 왼쪽 위로 이어지는 도로를 향하고 먼지는 오른쪽 아래 후방으로 흐른다. 전진과 먼지 방향은 일치한다. 그러나 보이는 것은 후면과 왼쪽 운전석 측면으로, 지정된 조수석 측면과 반대다. 도로도 상단 중앙이 아니라 왼쪽 원경으로 빠진다. 식별 가능한 사람의 시선은 없다.",
        "built_space": "도로 하나가 하단 오른쪽에서 트럭 아래를 지나 상단 왼쪽으로 굽는다. 주변은 메마른 흙과 돌, 드문 나무로 이루어져 있다. 상단 중앙에는 참조의 좁은 바위 틈과 이를 둘러싼 좌우 바위, 상부 가로 바위 및 안쪽 돌출석이 보인다. 다만 바위군이 먼 배경에 비해 지나치게 크고, 바닥 경계와 주변 지형의 연결이 부자연스럽다. 인공 건축물이나 고정 설비는 보이지 않는다.",
        "entities": "녹슨 파란 화물트럭 한 대와 비포장도로, 바퀴 뒤의 흙먼지가 확인된다. 적재함에는 낮게 놓인 밝은색 물체 일부가 보이지만 종류는 확정할 수 없다. 인물과 읽을 수 있는 문자는 보이지 않는다. 낡은 트럭과 황량한 낮 풍경은 요청에 맞으며, 바위 표면과 틈 형태는 참조를 따르지만 주변 풍경에 자연스럽게 통합되지는 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 앞뒤 타이어가 노면에 닿아 트럭을 지지한다. 적재함 안 물체도 바닥 위에 놓인 것으로 보인다. 흙먼지가 뒷바퀴 부근에서 솟아 뒤로 길게 퍼져 주행 동작을 설명한다. 먼지가 후면 하부를 일부 가리지만 차체 대부분은 드러난다. 지지 없이 공중에 떠 있는 물체나 불가능한 차체 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.067
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.817
   },
   "violations": {
    "A": [
     "[gemini-pro] 콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
    ],
    "B": [
     "[gemini-pro] 금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
     "[gemini-pro] 콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)",
     "[gpt-high] 적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 817
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트의 차량과 도로 묘사는 따랐으나, 참고 이미지를 평면적인 스티커처럼 배경에 그대로 오려 붙여 실사 요건을 심각하게 위반했습니다.  ★위반: [gemini-pro] 콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
   },
   {
    "label": "B",
    "score": 817,
    "verdict_ko": "프롬프트가 명시적으로 금지한 텍스트('낡은 트럭')가 포함되었으며, 흙먼지가 직사각형 그래픽으로 조잡하게 합성되어 기각 대상입니다.  ★위반: [gemini-pro] 금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자) / [gemini-pro] 콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음) / [gpt-high] 적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_open_roof_truck_sel.png",
    "asset_id": "c94c53e3-76a2-4f78-b4ae-95204318d3c2",
    "role": "location_seed_bg"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-46ae-7e63-bbf3-ca42008d7676",
  "ref_mode": "seed-bg만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_bypass:bg_only"
 },
 "S63sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:51:42.117417+00:00",
  "fingerprint": "6f4e62cb2af24c1e47195b254cefa47e712465881adb3b1bc4a8464264a8e2a2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S63sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S63sh1_sel.png",
  "source_sha256": "cb6a30f70a1a2b3aa42444ad36e33c3981d8badcbd4a2bf473bae330dec20701",
  "file": "S63sh1_cine.png",
  "staged_sha256": "3f791cb1c06bee978a66f76876c3dc5dcd43b1bb9c68b73d41ab0eeb80b9e1f6",
  "latency_ms": 13048
 },
 "S63sh16::signage": {
  "fp": "609d812f5f8a669e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S63sh16": {
  "input_fingerprint": "c6a6ee73e8d5aa3d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상) 어두운 방 안, 핏기 없는 창백한 얼굴로 눈을 반쯤 감은 채 누워있는 미연의 얼굴 클로즈업.\n\nLOCATION (lock): On a floating container panel over the inundated refugee settlement at dawn, at the mother's final resting place in the flashback. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Room interior (Only an indistinct portion surrounding the reclining figure is visible); used as Provides minimal spatial context without introducing furnishings absent from the scene.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The remembered room remains dark, with restrained facial exposure preserving 미연's pallor and half-closed eyes without adding a dream glow or a separate color treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the recalled scene, a container panel floats on the open sea at dawn; this is not an indoor deathbed. 미연: She is stretched out on a floating container panel with a bleeding puncture wound in her abdomen and a face swollen from the earlier beating. She is barely conscious, immediately before her death.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상) 어두운 방 안, 핏기 없는 창백한 얼굴로 눈을 반쯤 감은 채 누워있는 미연의 얼굴 클로즈업.\n\nLOCATION (lock): On a floating container panel over the inundated refugee settlement at dawn, at the mother's final resting place in the flashback. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Room interior (Only an indistinct portion surrounding the reclining figure is visible); used as Provides minimal spatial context without introducing furnishings absent from the scene.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The remembered room remains dark, with restrained facial exposure preserving 미연's pallor and half-closed eyes without adding a dream glow or a separate color treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the recalled scene, a container panel floats on the open sea at dawn; this is not an indoor deathbed. 미연: She is stretched out on a floating container panel with a bleeding puncture wound in her abdomen and a face swollen from the earlier beating. She is barely conscious, immediately before her death.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상) 어두운 방 안, 핏기 없는 창백한 얼굴로 눈을 반쯤 감은 채 누워있는 미연의 얼굴 클로즈업.\n\nLOCATION (lock): On a floating container panel over the inundated refugee settlement at dawn, at the mother's final resting place in the flashback. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Room interior (Only an indistinct portion surrounding the reclining figure is visible); used as Provides minimal spatial context without introducing furnishings absent from the scene.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The remembered room remains dark, with restrained facial exposure preserving 미연's pallor and half-closed eyes without adding a dream glow or a separate color treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the recalled scene, a container panel floats on the open sea at dawn; this is not an indoor deathbed. 미연: She is stretched out on a floating container panel with a bleeding puncture wound in her abdomen and a face swollen from the earlier beating. She is barely conscious, immediately before her death.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 위를 향하고 있으며 눈을 반쯤 감고 있음.",
    "built_space": "녹슨 컨테이너 패널 표면이 질감 있게 표현됨.",
    "entities": "미연 (40대 여성, 창백한 피부, 이전 샷과 동일한 모자 착용).",
    "hard_violations": [],
    "physics": "머리와 몸이 컨테이너 바닥에 중력을 받아 자연스럽게 눕혀져 있음."
   },
   {
    "label": "B",
    "direction": "시선은 앵글 측면을 향하며 눈을 반쯤 감고 있음.",
    "built_space": "녹슨 컨테이너 패널 표면이 질감 있게 표현됨.",
    "entities": "미연 (40대 여성, 창백한 피부, 모자 미착용).",
    "hard_violations": [],
    "physics": "머리가 측면으로 컨테이너 바닥에 닿아 안정적으로 지지받고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 복장(모자)을 정확히 유지하고 창백한 얼굴과 반쯤 감은 눈을 명시된 클로즈업 앵글로 충실히 구현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "창백한 표정과 어두운 분위기는 적절하나 레퍼런스에 있는 모자를 누락하여 복장 일관성에서 감점됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 위를 향하고 있으며 눈을 반쯤 감고 있음.",
        "built_space": "녹슨 컨테이너 패널 표면이 질감 있게 표현됨.",
        "entities": "미연 (40대 여성, 창백한 피부, 이전 샷과 동일한 모자 착용).",
        "hard_violations": [],
        "physics": "머리와 몸이 컨테이너 바닥에 중력을 받아 자연스럽게 눕혀져 있음."
       },
       {
        "label": "B",
        "direction": "시선은 앵글 측면을 향하며 눈을 반쯤 감고 있음.",
        "built_space": "녹슨 컨테이너 패널 표면이 질감 있게 표현됨.",
        "entities": "미연 (40대 여성, 창백한 피부, 모자 미착용).",
        "hard_violations": [],
        "physics": "머리가 측면으로 컨테이너 바닥에 닿아 안정적으로 지지받고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 복장(모자)을 정확히 유지하고 창백한 얼굴과 반쯤 감은 눈을 명시된 클로즈업 앵글로 충실히 구현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "창백한 표정과 어두운 분위기는 적절하나 레퍼런스에 있는 모자를 누락하여 복장 일관성에서 감점됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 위를 향하고 있으며 눈을 반쯤 감고 있음.",
        "built_space": "녹슨 컨테이너 패널 표면이 질감 있게 표현됨.",
        "entities": "미연 (40대 여성, 창백한 피부, 이전 샷과 동일한 모자 착용).",
        "hard_violations": [],
        "physics": "머리와 몸이 컨테이너 바닥에 중력을 받아 자연스럽게 눕혀져 있음."
       },
       {
        "label": "B",
        "direction": "시선은 앵글 측면을 향하며 눈을 반쯤 감고 있음.",
        "built_space": "녹슨 컨테이너 패널 표면이 질감 있게 표현됨.",
        "entities": "미연 (40대 여성, 창백한 피부, 모자 미착용).",
        "hard_violations": [],
        "physics": "머리가 측면으로 컨테이너 바닥에 닿아 안정적으로 지지받고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "창백한 얼굴과 반쯤 감긴 눈을 밀착해 담고 패널의 지지도 자연스럽지만, 이전 장면의 모자가 사라지고 머리 모양도 달라 외형 연속성이 떨어진다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴 클로즈업, 반쯤 감긴 눈, 절제된 노출을 충족하면서 모자와 검은 머리, 녹슨 패널까지 유지해 미연의 외형과 장소 연속성이 더 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴이 화면을 향하도록 옆으로 누워 있으며, 반쯤 열린 눈은 카메라 아래쪽 가까운 공간을 멍하게 향한다. 특정 대상을 응시하거나 움직이는 모습은 없으며, 겨누는 물건도 없다.",
        "built_space": "머리 아래에 청회색의 녹슨 금속 패널 하나가 보이고, 왼쪽 뒤에는 구멍이 있는 모서리 결합부 하나가 보인다. 이전 장면의 패널 재질과 부식 상태에 부합한다. 배경은 검게 흐려져 방의 벽이나 실내 구조는 확인되지 않으며, 추가 가구나 반사는 없다.",
        "entities": "중년 동아시아계 여성 한 명만 보이며, 창백한 피부와 검은 머리, 갈색 체크 셔츠가 미연의 설정에 대체로 맞는다. 얼굴 윤곽은 유사하지만 이전 장면의 모자는 없고 머리는 더 매끈하게 뒤로 넘어가 있다. 눈은 정상적인 홍채와 동공을 유지한다. 얼굴의 타박상과 부기는 뚜렷하지 않다. 복부 상처와 나머지 복장은 화면 밖이므로 평가하지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리 옆면과 뺨 아래가 금속 패널에 닿아 있고 머리카락도 표면에 눌려 있어 지지점이 명확하다. 보이는 목과 어깨는 옆으로 누운 자세와 연결되며, 힘을 주어 신체를 들거나 공중에 뜬 부분은 없다."
       },
       {
        "label": "B",
        "direction": "바로 누운 얼굴을 위에서 가까이 내려다본 구도다. 반쯤 열린 눈은 대체로 카메라 쪽 위 공간을 향하지만 초점이 흐려 의도적인 응시보다는 의식이 희미한 상태로 읽힌다. 움직이는 신체나 방향성 있는 소품은 없다.",
        "built_space": "뒤통수 아래에 녹슨 청회색 컨테이너 패널 하나가 있고, 왼쪽 위에 구멍 난 모서리 결합부 하나와 가장자리 보강 부분이 보인다. 재질과 마모가 이전 장면과 이어진다. 얼굴 주변 외에는 거의 보이지 않아 실내 구조는 확인할 수 없으며, 새 가구나 불가능한 반사는 없다.",
        "entities": "미연에 해당하는 중년 동아시아계 여성 한 명만 보인다. 중년의 얼굴 윤곽, 검은 머리, 어두운 챙 모자가 참조와 잘 이어진다. 얼굴은 창백하고 눈은 반쯤 감겨 있으며 홍채와 동공도 자연스럽다. 턱 주변에 약한 변색은 있지만 구타로 인한 부기는 선명하지 않다. 옷과 복부 상처는 대부분 화면 밖이므로 누락으로 평가하지 않는다. 추가 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수가 패널에 놓이고 머리카락이 양옆 표면으로 퍼져 있어 머리의 지지가 자연스럽다. 목은 바로 누운 몸으로 이어지며, 보이는 신체에서 능동적으로 들어 올린 부분이나 지지 없이 떠 있는 부분은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "창백한 얼굴과 반쯤 감긴 눈을 밀착해 담고 패널의 지지도 자연스럽지만, 이전 장면의 모자가 사라지고 머리 모양도 달라 외형 연속성이 떨어진다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴 클로즈업, 반쯤 감긴 눈, 절제된 노출을 충족하면서 모자와 검은 머리, 녹슨 패널까지 유지해 미연의 외형과 장소 연속성이 더 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴이 화면을 향하도록 옆으로 누워 있으며, 반쯤 열린 눈은 카메라 아래쪽 가까운 공간을 멍하게 향한다. 특정 대상을 응시하거나 움직이는 모습은 없으며, 겨누는 물건도 없다.",
        "built_space": "머리 아래에 청회색의 녹슨 금속 패널 하나가 보이고, 왼쪽 뒤에는 구멍이 있는 모서리 결합부 하나가 보인다. 이전 장면의 패널 재질과 부식 상태에 부합한다. 배경은 검게 흐려져 방의 벽이나 실내 구조는 확인되지 않으며, 추가 가구나 반사는 없다.",
        "entities": "중년 동아시아계 여성 한 명만 보이며, 창백한 피부와 검은 머리, 갈색 체크 셔츠가 미연의 설정에 대체로 맞는다. 얼굴 윤곽은 유사하지만 이전 장면의 모자는 없고 머리는 더 매끈하게 뒤로 넘어가 있다. 눈은 정상적인 홍채와 동공을 유지한다. 얼굴의 타박상과 부기는 뚜렷하지 않다. 복부 상처와 나머지 복장은 화면 밖이므로 평가하지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리 옆면과 뺨 아래가 금속 패널에 닿아 있고 머리카락도 표면에 눌려 있어 지지점이 명확하다. 보이는 목과 어깨는 옆으로 누운 자세와 연결되며, 힘을 주어 신체를 들거나 공중에 뜬 부분은 없다."
       },
       {
        "label": "A",
        "direction": "바로 누운 얼굴을 위에서 가까이 내려다본 구도다. 반쯤 열린 눈은 대체로 카메라 쪽 위 공간을 향하지만 초점이 흐려 의도적인 응시보다는 의식이 희미한 상태로 읽힌다. 움직이는 신체나 방향성 있는 소품은 없다.",
        "built_space": "뒤통수 아래에 녹슨 청회색 컨테이너 패널 하나가 있고, 왼쪽 위에 구멍 난 모서리 결합부 하나와 가장자리 보강 부분이 보인다. 재질과 마모가 이전 장면과 이어진다. 얼굴 주변 외에는 거의 보이지 않아 실내 구조는 확인할 수 없으며, 새 가구나 불가능한 반사는 없다.",
        "entities": "미연에 해당하는 중년 동아시아계 여성 한 명만 보인다. 중년의 얼굴 윤곽, 검은 머리, 어두운 챙 모자가 참조와 잘 이어진다. 얼굴은 창백하고 눈은 반쯤 감겨 있으며 홍채와 동공도 자연스럽다. 턱 주변에 약한 변색은 있지만 구타로 인한 부기는 선명하지 않다. 옷과 복부 상처는 대부분 화면 밖이므로 누락으로 평가하지 않는다. 추가 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수가 패널에 놓이고 머리카락이 양옆 표면으로 퍼져 있어 머리의 지지가 자연스럽다. 목은 바로 누운 몸으로 이어지며, 보이는 신체에서 능동적으로 들어 올린 부분이나 지지 없이 떠 있는 부분은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.446
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.446
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1446
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "이전 샷의 복장(모자)을 정확히 유지하고 창백한 얼굴과 반쯤 감은 눈을 명시된 클로즈업 앵글로 충실히 구현함."
   },
   {
    "label": "B",
    "score": 1446,
    "verdict_ko": "창백한 표정과 어두운 분위기는 적절하나 레퍼런스에 있는 모자를 누락하여 복장 일관성에서 감점됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh7_sel.png",
    "asset_id": "2d4ed774-9a0b-4232-9e60-89a67937b101",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:943817>",
    "asset_id": "5b79d22c-bd89-4281-8528-63a9a47cc023",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-485c-78a7-b480-0e4d8d0090ad",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S38sh7"
  }
 },
 "S63sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:59:50.098956+00:00",
  "fingerprint": "b1066e72342c7c678fb41967bd824b5e1a45f57617b3feaa1d5326d3476e9b27",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S63sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S63sh16_sel.png",
  "source_sha256": "667ce28375eaf9f089078f2fee12b483ea764d5755244f403be3777ca96ba3bc",
  "file": "S63sh16_cine.png",
  "staged_sha256": "3931eef21149f85b256260542469393f7285ce03c0e2844083fe0d11e54ee97b",
  "latency_ms": 10863
 },
 "S63sh17::confined_fp_apt": {
  "applies": true,
  "reason_ko": "이 샷은 트럭 운전석이라는 밀폐된 차량 내부 공간에서 진행됩니다. 현우가 운전석에 앉아 있고 빛이 조수석 창문에서 들어오는 구체적인 방향성이 제시되어 있으므로 인물의 착석 위치와 창문의 상대적인 위치를 정확하게 배치하지 않으면 화면의 연속성과 개연성이 깨지는 심각한 오류가 발생할 수 있습니다.",
  "input_fingerprint": "93c5004a51d6e787"
 },
 "S63sh17::signage": {
  "fp": "fb25ace90e92090b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::529f9dcd3f66": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_529f9dcd3f66.png",
  "place_text": "At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window.",
  "input_fingerprint": "ceeaf8d945ef443b"
 },
 "S63sh17::confined_fp": {
  "reads": {
   "controls": "A steering wheel is located at the driver's station on the left.",
   "mirrors": "No mirrors are depicted in the diagram.",
   "camera": "Positioned near the center windshield area on the passenger side, angled leftward and pointing directly at the driver.",
   "occupants": "Hyun-woo occupies the driver seat."
  },
  "mismatches": [],
  "scene_description_en": "The camera, positioned on the passenger side near the windshield, points diagonally left toward the driver's seat in a close-up framing. Hyun-woo occupies the driver's seat on the left side of the screen, facing forward toward the left edge of the frame, presenting his right profile to the lens. The steering wheel sits directly in front of him on the left. Sunlight illuminates the right side of his face, cast from the passenger-side window located off-screen to the right. The blurred interior of the driver's side cabin wall serves as the background behind him on the far left.",
  "fixed": false,
  "input_fingerprint": "ee553b8a16535dd4"
 },
 "era_assess::0d4b3c45090d2aca": {
  "subjects": [],
  "subject_text": "현우와 수빈이 사용하는 낡은 트럭 운전석\n운전석과 조수석이 나란히 놓인 낡은 운전 공간. 앞유리 아래로 운전대와 계기판이 있고 양옆 창문으로 외부가 보인다.",
  "identity": "canonical",
  "scope_id": "L232",
  "scope_role": "location_exterior",
  "scope_sha": "9aec4edb9cd1cee5"
 },
 "S63sh17": {
  "input_fingerprint": "25a32ec057e84073",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 조수석 창문으로 들어온 따스한 햇빛을 받으며 입술을 꽉 다문 현우의 결의에 찬 얼굴 클로즈업.\n\nLOCATION (lock): At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cab interior (현우 remains seated in the driver's position) — A narrow interior portion behind the driver is seen from the passenger side; used as Maintains the restored present-tense location without bringing 수빈 into this close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Warm sunlight entering through the passenger-side window reaches 현우's face, with controlled highlights preserving the tension around his closed lips.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck continues along the road in daylight. 현우: He remains in the driver's seat, with his fighting injuries still present and a resolute expression.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera, positioned on the passenger side near the windshield, points diagonally left toward the driver's seat in a close-up framing. Hyun-woo occupies the driver's seat on the left side of the screen, facing forward toward the left edge of the frame, presenting his right profile to the lens. The steering wheel sits directly in front of him on the left. Sunlight illuminates the right side of his face, cast from the passenger-side window located off-screen to the right. The blurred interior of the driver's side cabin wall serves as the background behind him on the far left.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 조수석 창문으로 들어온 따스한 햇빛을 받으며 입술을 꽉 다문 현우의 결의에 찬 얼굴 클로즈업.\n\nLOCATION (lock): At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Warm sunlight entering through the passenger-side window reaches 현우's face, with controlled highlights preserving the tension around his closed lips.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck continues along the road in daylight. 현우: He remains in the driver's seat, with his fighting injuries still present and a resolute expression.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera, positioned on the passenger side near the windshield, points diagonally left toward the driver's seat in a close-up framing. Hyun-woo occupies the driver's seat on the left side of the screen, facing forward toward the left edge of the frame, presenting his right profile to the lens. The steering wheel sits directly in front of him on the left. Sunlight illuminates the right side of his face, cast from the passenger-side window located off-screen to the right. The blurred interior of the driver's side cabin wall serves as the background behind him on the far left.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 조수석 창문으로 들어온 따스한 햇빛을 받으며 입술을 꽉 다문 현우의 결의에 찬 얼굴 클로즈업.\n\nLOCATION (lock): At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Warm sunlight entering through the passenger-side window reaches 현우's face, with controlled highlights preserving the tension around his closed lips.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck continues along the road in daylight. 현우: He remains in the driver's seat, with his fighting injuries still present and a resolute expression.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S63sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S63sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 왼쪽을 향함.",
    "built_space": "운전석 외부 뷰 또는 우핸들 구조로 조수석 시점 위반.",
    "entities": "인물은 유사하나 검은 티셔츠로 의상 불일치.",
    "hard_violations": [
     "[gemini-pro] 카메라 위치 위반 (조수석 내부 뷰 아님)"
    ],
    "physics": "좌석에 앉아 지지됨."
   },
   {
    "label": "B",
    "direction": "시선은 화면 왼쪽을 향함.",
    "built_space": "전경 사이드미러 존재로 보아 운전석 외부 뷰임. 조수석 시점 위반.",
    "entities": "인물 외형과 회색 셔츠 등 레퍼런스 일치함.",
    "hard_violations": [
     "[gemini-pro] 카메라 위치 위반 (조수석 내부가 아닌 운전석 외부 뷰)",
     "[gpt-high] 외부 측면 거울과 차체 창문 턱을 전경에 둔 차창 밖 시점으로, 지정된 조수석 쪽 운전실 내부 카메라 배치를 벗어난다."
    ],
    "physics": "휠을 쥔 손과 자세가 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "의상과 인물 외형은 잘 일치하나 카메라가 조수석이 아닌 운전석 외부에 위치하여 공간 구조 조건을 위반함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 회색 셔츠가 아닌 검은색 옷을 입었으며 카메라 역시 운전석 밖에 위치하여 시점 조건을 완전히 위반함."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "시선은 화면 왼쪽을 향함.",
        "built_space": "전경 사이드미러 존재로 보아 운전석 외부 뷰임. 조수석 시점 위반.",
        "entities": "인물 외형과 회색 셔츠 등 레퍼런스 일치함.",
        "hard_violations": [
         "카메라 위치 위반 (조수석 내부가 아닌 운전석 외부 뷰)"
        ],
        "physics": "휠을 쥔 손과 자세가 자연스러움."
       },
       {
        "label": "A",
        "direction": "시선은 화면 왼쪽을 향함.",
        "built_space": "운전석 외부 뷰 또는 우핸들 구조로 조수석 시점 위반.",
        "entities": "인물은 유사하나 검은 티셔츠로 의상 불일치.",
        "hard_violations": [
         "카메라 위치 위반 (조수석 내부 뷰 아님)"
        ],
        "physics": "좌석에 앉아 지지됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "의상과 인물 외형은 잘 일치하나 카메라가 조수석이 아닌 운전석 외부에 위치하여 공간 구조 조건을 위반함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 회색 셔츠가 아닌 검은색 옷을 입었으며 카메라 역시 운전석 밖에 위치하여 시점 조건을 완전히 위반함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "시선은 화면 왼쪽을 향함.",
        "built_space": "전경 사이드미러 존재로 보아 운전석 외부 뷰임. 조수석 시점 위반.",
        "entities": "인물 외형과 회색 셔츠 등 레퍼런스 일치함.",
        "hard_violations": [
         "카메라 위치 위반 (조수석 내부가 아닌 운전석 외부 뷰)"
        ],
        "physics": "휠을 쥔 손과 자세가 자연스러움."
       },
       {
        "label": "A",
        "direction": "시선은 화면 왼쪽을 향함.",
        "built_space": "운전석 외부 뷰 또는 우핸들 구조로 조수석 시점 위반.",
        "entities": "인물은 유사하나 검은 티셔츠로 의상 불일치.",
        "hard_violations": [
         "카메라 위치 위반 (조수석 내부 뷰 아님)"
        ],
        "physics": "좌석에 앉아 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "회색 셔츠와 인물 외형은 참조에 가깝지만, 차창 밖에서 팔과 차체까지 담아 조수석 쪽 실내 얼굴 클로즈업이라는 핵심 구도를 어겼다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "조수석 쪽에서 본 운전자의 얼굴 클로즈업, 전방 주시와 다문 입술·싸움 상처를 잘 구현했으나 검은 티셔츠는 참조의 회색 셔츠와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 두 눈은 화면 왼쪽 차량 전방을 향하며 카메라를 보지 않는다. 팔도 왼쪽 아래 운전대 쪽으로 뻗어 있다. 따뜻한 빛은 주로 머리 뒤쪽과 어깨에 강하게 닿고, 얼굴에 들어오는 조수석 쪽 햇빛은 상대적으로 약하다.",
        "built_space": "운전대 일부 하나, 좌석 등받이 일부 하나, 중앙의 측면 창 하나, 오른쪽 뒤 유리 일부와 낡은 천장·내장재가 보인다. 화면 왼쪽에는 커다란 외부 측면 거울 하나가 있고 아래에는 외부 도장과 창문 턱이 전경으로 드러나므로, 카메라는 조수석 실내보다 차창 밖에 놓인 것으로 읽힌다. 얼굴뿐 아니라 상반신과 팔, 차체까지 보여 요구한 좁은 얼굴 클로즈업보다 넓다. 거울에 판별 가능한 반사상은 없다.",
        "entities": "등장인물은 현우로 읽히는 젊은 동아시아계 남성 한 명뿐이다. 앳된 얼굴, 헝클어진 검은 머리와 회색 칼라 셔츠가 참조에 대체로 맞으며, 국적은 외형만으로 확인할 수 없다. 목과 팔에 상처가 남아 있고 입술은 다물려 있다. 낡은 트럭의 재질도 확인된다. 수빈이나 다른 사람은 없으며, 천장 표지는 흐려 읽을 수 없다.",
        "hard_violations": [
         "외부 측면 거울과 차체 창문 턱을 전경에 둔 차창 밖 시점으로, 지정된 조수석 쪽 운전실 내부 카메라 배치를 벗어난다."
        ],
        "physics": "몸통 뒤에 좌석 등받이가 있고 운전석에 앉은 상체의 자세는 자연스럽다. 팔은 운전대 방향으로 뻗으며 손의 접촉 부위는 거울에 가려져 확인되지 않는다. 거울은 수직 지지대에 고정되어 있고 운전대도 조향부에 연결되어 있다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우의 코와 시선은 화면 왼쪽 앞유리 너머 차량 진행 방향을 향한다. 카메라를 응시하지 않으며 입술을 다물고 턱에 힘을 준 모습이다. 얼굴에는 따뜻한 낮빛이 닿지만 광원인 조수석 창 자체는 프레임 밖이므로 유입 경로를 직접 확인할 수는 없다.",
        "built_space": "왼쪽 아래 운전대 하나와 대시보드 일부, 왼쪽 앞유리 일부, 중앙의 운전석 측면 창과 작은 삼각창, 오른쪽 뒤 유리 일부, 오른쪽 좌석 등받이 하나가 보인다. 앞 기둥의 손잡이 하나도 보인다. 조수석 쪽 실내에서 운전자를 옆으로 보는 관계가 성립하며 현우는 운전대 뒤, 등받이 앞에 있다. 얼굴이 크게 잡히고 운전자 뒤 공간은 좁게 남는다. 중복된 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "앳된 동아시아계 남성 한 명으로, 검은 헝클어진 머리와 얼굴 외형이 현우 참조에 대체로 부합한다. 정확한 나이와 한국계 미국인이라는 국적·배경은 이미지로 확정할 수 없다. 코, 뺨, 입술과 목에 싸움의 상처가 보이고 눈은 정상적인 인간의 눈이다. 다만 보이는 옷은 검은 라운드넥 티셔츠여서 참조의 회색 칼라 셔츠와 불일치한다. 낡은 트럭 내부 외에 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "상체 바로 뒤에 좌석 등받이가 있어 앉은 자세의 지지가 성립한다. 골반과 다리, 손은 클로즈업 밖이므로 접촉 상태를 평가할 수 없으며, 이를 부유나 운전 동작 오류로 볼 근거는 없다. 운전대와 손잡이, 창틀은 차량에 고정되어 있고 신체의 비정상적 변형이나 지지 없는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "회색 셔츠와 인물 외형은 참조에 가깝지만, 차창 밖에서 팔과 차체까지 담아 조수석 쪽 실내 얼굴 클로즈업이라는 핵심 구도를 어겼다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "조수석 쪽에서 본 운전자의 얼굴 클로즈업, 전방 주시와 다문 입술·싸움 상처를 잘 구현했으나 검은 티셔츠는 참조의 회색 셔츠와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 두 눈은 화면 왼쪽 차량 전방을 향하며 카메라를 보지 않는다. 팔도 왼쪽 아래 운전대 쪽으로 뻗어 있다. 따뜻한 빛은 주로 머리 뒤쪽과 어깨에 강하게 닿고, 얼굴에 들어오는 조수석 쪽 햇빛은 상대적으로 약하다.",
        "built_space": "운전대 일부 하나, 좌석 등받이 일부 하나, 중앙의 측면 창 하나, 오른쪽 뒤 유리 일부와 낡은 천장·내장재가 보인다. 화면 왼쪽에는 커다란 외부 측면 거울 하나가 있고 아래에는 외부 도장과 창문 턱이 전경으로 드러나므로, 카메라는 조수석 실내보다 차창 밖에 놓인 것으로 읽힌다. 얼굴뿐 아니라 상반신과 팔, 차체까지 보여 요구한 좁은 얼굴 클로즈업보다 넓다. 거울에 판별 가능한 반사상은 없다.",
        "entities": "등장인물은 현우로 읽히는 젊은 동아시아계 남성 한 명뿐이다. 앳된 얼굴, 헝클어진 검은 머리와 회색 칼라 셔츠가 참조에 대체로 맞으며, 국적은 외형만으로 확인할 수 없다. 목과 팔에 상처가 남아 있고 입술은 다물려 있다. 낡은 트럭의 재질도 확인된다. 수빈이나 다른 사람은 없으며, 천장 표지는 흐려 읽을 수 없다.",
        "hard_violations": [
         "외부 측면 거울과 차체 창문 턱을 전경에 둔 차창 밖 시점으로, 지정된 조수석 쪽 운전실 내부 카메라 배치를 벗어난다."
        ],
        "physics": "몸통 뒤에 좌석 등받이가 있고 운전석에 앉은 상체의 자세는 자연스럽다. 팔은 운전대 방향으로 뻗으며 손의 접촉 부위는 거울에 가려져 확인되지 않는다. 거울은 수직 지지대에 고정되어 있고 운전대도 조향부에 연결되어 있다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우의 코와 시선은 화면 왼쪽 앞유리 너머 차량 진행 방향을 향한다. 카메라를 응시하지 않으며 입술을 다물고 턱에 힘을 준 모습이다. 얼굴에는 따뜻한 낮빛이 닿지만 광원인 조수석 창 자체는 프레임 밖이므로 유입 경로를 직접 확인할 수는 없다.",
        "built_space": "왼쪽 아래 운전대 하나와 대시보드 일부, 왼쪽 앞유리 일부, 중앙의 운전석 측면 창과 작은 삼각창, 오른쪽 뒤 유리 일부, 오른쪽 좌석 등받이 하나가 보인다. 앞 기둥의 손잡이 하나도 보인다. 조수석 쪽 실내에서 운전자를 옆으로 보는 관계가 성립하며 현우는 운전대 뒤, 등받이 앞에 있다. 얼굴이 크게 잡히고 운전자 뒤 공간은 좁게 남는다. 중복된 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "앳된 동아시아계 남성 한 명으로, 검은 헝클어진 머리와 얼굴 외형이 현우 참조에 대체로 부합한다. 정확한 나이와 한국계 미국인이라는 국적·배경은 이미지로 확정할 수 없다. 코, 뺨, 입술과 목에 싸움의 상처가 보이고 눈은 정상적인 인간의 눈이다. 다만 보이는 옷은 검은 라운드넥 티셔츠여서 참조의 회색 칼라 셔츠와 불일치한다. 낡은 트럭 내부 외에 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "상체 바로 뒤에 좌석 등받이가 있어 앉은 자세의 지지가 성립한다. 골반과 다리, 손은 클로즈업 밖이므로 접촉 상태를 평가할 수 없으며, 이를 부유나 운전 동작 오류로 볼 근거는 없다. 운전대와 손잡이, 창틀은 차량에 고정되어 있고 신체의 비정상적 변형이나 지지 없는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 카메라 위치 위반 (조수석 내부가 아닌 운전석 외부 뷰)",
     "[gpt-high] 외부 측면 거울과 차체 창문 턱을 전경에 둔 차창 밖 시점으로, 지정된 조수석 쪽 운전실 내부 카메라 배치를 벗어난다."
    ],
    "A": [
     "[gemini-pro] 카메라 위치 위반 (조수석 내부 뷰 아님)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "의상과 인물 외형은 잘 일치하나 카메라가 조수석이 아닌 운전석 외부에 위치하여 공간 구조 조건을 위반함.  ★위반: [gemini-pro] 카메라 위치 위반 (조수석 내부가 아닌 운전석 외부 뷰) / [gpt-high] 외부 측면 거울과 차체 창문 턱을 전경에 둔 차창 밖 시점으로, 지정된 조수석 쪽 운전실 내부 카메라 배치를 벗어난다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "지정된 회색 셔츠가 아닌 검은색 옷을 입었으며 카메라 역시 운전석 밖에 위치하여 시점 조건을 완전히 위반함.  ★위반: [gemini-pro] 카메라 위치 위반 (조수석 내부 뷰 아님)"
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S63sh17_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-4a0a-7137-bb64-a14798b9b242",
  "confined_fp": {
   "base_key": "confinedfp::529f9dcd3f66",
   "apt_reason": "이 샷은 트럭 운전석이라는 밀폐된 차량 내부 공간에서 진행됩니다. 현우가 운전석에 앉아 있고 빛이 조수석 창문에서 들어오는 구체적인 방향성이 제시되어 있으므로 인물의 착석 위치와 창문의 상대적인 위치를 정확하게 배치하지 않으면 화면의 연속성과 개연성이 깨지는 심각한 오류가 발생할 수 있습니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S63sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:02:24.376489+00:00",
  "fingerprint": "c7e4b4c4a12f201d56554102fe794e97b155303af389ffa99e9f47d9538bacda",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S63sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S63sh17_sel.png",
  "source_sha256": "539d64703e31be9fbf22131aa19df8c22a61965f4c77db40d78c687f29c62d50",
  "file": "S63sh17_cine.png",
  "staged_sha256": "5859701acff521e03f236143cd9637e18286c58c3471ef14953e64e5b3cb1798",
  "latency_ms": 10122
 },
 "S64sh7::signage": {
  "fp": "516d630c928dd2fb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S64sh7": {
  "input_fingerprint": "a398e5c6201800cb",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The machinery warehouse has a half-open door and a heap of broken robots; baby birds are already concealed behind a box in B-200's corner. B-200 crouches in the darkness with converted gun-hands raised, while Charlie has brought in a football and retains his damaged body, chest ring and impaired systems.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The machinery warehouse has a half-open door and a heap of broken robots; baby birds are already concealed behind a box in B-200's corner. B-200 crouches in the darkness with converted gun-hands raised, while Charlie has brought in a football and retains his damaged body, chest ring and impaired systems.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The machinery warehouse has a half-open door and a heap of broken robots; baby birds are already concealed behind a box in B-200's corner. B-200 crouches in the darkness with converted gun-hands raised, while Charlie has brought in a football and retains his damaged body, chest ring and impaired systems.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7__bgfirst_bg.png",
     "asset_id": "2c542339-d519-40a3-aee9-40b352e75792",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S64sh7.png",
     "asset_id": "cbafc64b-e6e6-4acf-8c80-6fb23ad1a5d9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L238B01.png",
     "asset_id": "e1a9f48b-d9e5-4c90-99a8-fadcfb66b2d8",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "B-200의 오른팔 총구가 왼쪽 인물을 향해 조준되어 있음.",
    "built_space": "창고 내부 구조와 빛의 방향이 레퍼런스와 유사하나 2D 일러스트 톤으로 렌더링됨.",
    "entities": "B-200은 형태상 유사하나 실사가 아님. 왼쪽 인물은 축구공을 들고 있으나 코트, 모자, 얼굴 형태 등 찰리의 레퍼런스와 전혀 일치하지 않음.",
    "hard_violations": [
     "[gemini-pro] 창작된 인물 (찰리 레퍼런스와 전혀 다른 외형의 캐릭터 등장)"
    ],
    "physics": "두 캐릭터 모두 바닥에 안정적으로 서 있음."
   },
   {
    "label": "B",
    "direction": "B-200의 오른팔 고사포가 왼쪽 찰리의 상체를 정확히 정조준함.",
    "built_space": "위치 레퍼런스와 정확히 일치하는 실사 창고 내부이며, 우측 구석에 부서진 로봇 부품 더미가 배치됨.",
    "entities": "B-200은 실사 레퍼런스와 완벽히 일치함. 찰리는 얼굴 마스크와 장갑판이 일치하고 손상된 링 묘사가 있으나, 코트, 모자, 축구공이 없음.",
    "hard_violations": [],
    "physics": "두 캐릭터 모두 무게 중심을 잡고 바닥에 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "실사 배경 묘사와 B-200의 정조준 자세는 매우 훌륭하게 구현되었으나, 찰리의 지정된 의상과 축구공이 누락됨."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "실사 영화 스틸컷 조건에 위배되는 2D 그래픽 스타일이며, 찰리의 외형이 레퍼런스와 완전히 달라 치명적인 감점 요소임."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 오른팔 총구가 왼쪽 인물을 향해 조준되어 있음.",
        "built_space": "창고 내부 구조와 빛의 방향이 레퍼런스와 유사하나 2D 일러스트 톤으로 렌더링됨.",
        "entities": "B-200은 형태상 유사하나 실사가 아님. 왼쪽 인물은 축구공을 들고 있으나 코트, 모자, 얼굴 형태 등 찰리의 레퍼런스와 전혀 일치하지 않음.",
        "hard_violations": [
         "창작된 인물 (찰리 레퍼런스와 전혀 다른 외형의 캐릭터 등장)"
        ],
        "physics": "두 캐릭터 모두 바닥에 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "B-200의 오른팔 고사포가 왼쪽 찰리의 상체를 정확히 정조준함.",
        "built_space": "위치 레퍼런스와 정확히 일치하는 실사 창고 내부이며, 우측 구석에 부서진 로봇 부품 더미가 배치됨.",
        "entities": "B-200은 실사 레퍼런스와 완벽히 일치함. 찰리는 얼굴 마스크와 장갑판이 일치하고 손상된 링 묘사가 있으나, 코트, 모자, 축구공이 없음.",
        "hard_violations": [],
        "physics": "두 캐릭터 모두 무게 중심을 잡고 바닥에 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "실사 배경 묘사와 B-200의 정조준 자세는 매우 훌륭하게 구현되었으나, 찰리의 지정된 의상과 축구공이 누락됨."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "실사 영화 스틸컷 조건에 위배되는 2D 그래픽 스타일이며, 찰리의 외형이 레퍼런스와 완전히 달라 치명적인 감점 요소임."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 오른팔 총구가 왼쪽 인물을 향해 조준되어 있음.",
        "built_space": "창고 내부 구조와 빛의 방향이 레퍼런스와 유사하나 2D 일러스트 톤으로 렌더링됨.",
        "entities": "B-200은 형태상 유사하나 실사가 아님. 왼쪽 인물은 축구공을 들고 있으나 코트, 모자, 얼굴 형태 등 찰리의 레퍼런스와 전혀 일치하지 않음.",
        "hard_violations": [
         "창작된 인물 (찰리 레퍼런스와 전혀 다른 외형의 캐릭터 등장)"
        ],
        "physics": "두 캐릭터 모두 바닥에 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "B-200의 오른팔 고사포가 왼쪽 찰리의 상체를 정확히 정조준함.",
        "built_space": "위치 레퍼런스와 정확히 일치하는 실사 창고 내부이며, 우측 구석에 부서진 로봇 부품 더미가 배치됨.",
        "entities": "B-200은 실사 레퍼런스와 완벽히 일치함. 찰리는 얼굴 마스크와 장갑판이 일치하고 손상된 링 묘사가 있으나, 코트, 모자, 축구공이 없음.",
        "hard_violations": [],
        "physics": "두 캐릭터 모두 무게 중심을 잡고 바닥에 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "총구가 찰리의 머리 위로 빗나가고 B-200이 거의 서 있으며, 전신 위주의 넓은 구도라 정조준하는 미디엄 숏을 구현하지 못했다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "왼쪽 전경의 찰리를 향한 포신과 B-200의 낮춘 자세는 더 정확하지만, 넓은 프레이밍과 찰리의 체형·복장 불일치, 노출된 새끼 새가 충실도를 떨어뜨린다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 들어 올린 포신은 화면 왼쪽 위를 향한다. 포신 축을 연장하면 왼쪽 찰리의 머리 위로 지나가므로 찰리를 정조준하지 않는다. 총구는 카메라 정면이 아닌 비스듬한 측면으로 보인다. B-200의 머리는 왼쪽으로 돌아가 있으나 찰리의 얼굴은 B-200보다 화면 앞쪽을 향한다. 반대쪽 손은 아래로 내려가 있다.",
        "built_space": "왼쪽의 부분 개방된 대형 출입문 하나와 높은 창 하나, 천장 매달림 등 세 개, 뒤쪽 부품 선반군, 오른쪽 벽의 창 하나와 가장자리 선반이 보여 장소 사진의 주요 구조를 잘 유지한다. 부서진 로봇 더미는 B-200 뒤 오른쪽 벽 아래에 있다. 다만 넓은 바닥과 B-200의 발까지 담아 요구된 미디엄 숏보다 훨씬 넓고, 어두운 구석에 웅크린 대치보다는 창고 중앙의 대치에 가깝다.",
        "entities": "B-200의 짙은 회색 중장갑, 붉은 센서, 두꺼운 관절은 참조와 대체로 맞지만 한쪽 팔만 다연장 포신이고 반대쪽은 일반 기계 손이다. 찰리의 흰 마스크형 얼굴, 베이지 장갑, 가슴 고리는 보이나 참조의 모자와 긴 외투가 없고 체형도 달라졌다. 가슴 주변 초록 불꽃은 손상을 표현하지만 참조에 없는 강조 효과다. 축구공은 보이지 않으며 아래쪽 일부가 잘려 소지 여부는 확정하기 어렵다. 새끼 새는 보이지 않아 은폐 설정과 충돌하지 않는다. 판독 가능한 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "B-200은 벌린 두 발로 콘크리트 바닥을 딛고 있으며 무릎을 약간 굽혔지만 웅크린 상태는 아니다. 포신은 팔꿈치와 전완 기구에 연결되어 지지된다. 찰리의 발은 프레임 밖이지만 몸통과 다리의 연결은 자연스럽고 공중에 떠 있다는 증거는 없다. 뒤의 로봇 잔해는 바닥과 서로의 몸체에 기대어 쌓여 있다."
       },
       {
        "label": "B",
        "direction": "높이 든 포신은 오른쪽 B-200에서 왼쪽 전경 찰리의 상체 쪽으로 향하며, 포신 축을 연장하면 찰리의 어깨 아래 몸통에 닿는 방향이다. 낮은 반대쪽 포신도 왼쪽 찰리 쪽을 향한다. 총열의 옆면이 읽히고 카메라 자체를 겨누는 구도는 아니다. B-200의 얼굴은 찰리를 향하며 찰리도 등을 카메라에 둔 채 B-200 쪽으로 고개를 돌리고 있다.",
        "built_space": "왼쪽 부분 개방 출입문 하나, 높은 창 하나, 천장 등 세 개, 뒤쪽 부품 선반군, 오른쪽 벽 창과 가장자리 선반이 장소 참조와 대응한다. B-200은 오른쪽 벽 가까이, 찰리는 왼쪽 전경에 있어 대치 관계가 명확하다. 오른쪽 큰 상자 위로 로봇 잔해가 높게 쌓여 있으며 상자 옆 틈에는 새끼 새가 노출되어 있다. B-200의 전신과 넓은 바닥을 담아 요구된 미디엄 숏보다 넓다.",
        "entities": "B-200은 짙은 회색 중장비형 장갑과 양팔의 포신을 갖춰 주요 정체성이 맞지만 포신 구성은 참조보다 단순하다. 찰리는 베이지색 기계 몸체지만 긴 다리와 좁은 상체 때문에 고릴라형 비례가 약하고, 참조의 모자와 외투도 없다. 등 쪽 원형 부품은 보이나 가슴은 가려져 가슴 고리 유지 여부는 판정할 수 없다. 오른손 가까이에 축구공이 있다. 부서진 로봇 더미와 상자는 보이지만 새끼 새 세 마리가 드러나 은폐 상태를 지키지 못한다. 찰리의 표면은 금속 장갑보다 선으로 윤곽을 강조한 의상처럼 보여 실사 재질감이 약하다. 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "B-200은 양발을 넓게 바닥에 붙이고 무릎과 고관절을 굽혀 몸을 낮춘다. 완전히 깊게 웅크리지는 않았지만 중량을 지탱하는 자세는 성립한다. 두 포신은 각각 전완 기구에 연결되어 있다. 찰리의 발은 프레임 밖이며 지상에 선 자세로 읽힌다. 축구공은 오른손 손가락이 윗부분과 옆면에 닿아 붙잡은 상태로 보이고 허벅지에도 가까워 무지지 부유로 단정할 근거는 없다. 잔해는 상자와 서로의 부품에 받쳐져 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "총구가 찰리의 머리 위로 빗나가고 B-200이 거의 서 있으며, 전신 위주의 넓은 구도라 정조준하는 미디엄 숏을 구현하지 못했다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "왼쪽 전경의 찰리를 향한 포신과 B-200의 낮춘 자세는 더 정확하지만, 넓은 프레이밍과 찰리의 체형·복장 불일치, 노출된 새끼 새가 충실도를 떨어뜨린다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "B-200의 들어 올린 포신은 화면 왼쪽 위를 향한다. 포신 축을 연장하면 왼쪽 찰리의 머리 위로 지나가므로 찰리를 정조준하지 않는다. 총구는 카메라 정면이 아닌 비스듬한 측면으로 보인다. B-200의 머리는 왼쪽으로 돌아가 있으나 찰리의 얼굴은 B-200보다 화면 앞쪽을 향한다. 반대쪽 손은 아래로 내려가 있다.",
        "built_space": "왼쪽의 부분 개방된 대형 출입문 하나와 높은 창 하나, 천장 매달림 등 세 개, 뒤쪽 부품 선반군, 오른쪽 벽의 창 하나와 가장자리 선반이 보여 장소 사진의 주요 구조를 잘 유지한다. 부서진 로봇 더미는 B-200 뒤 오른쪽 벽 아래에 있다. 다만 넓은 바닥과 B-200의 발까지 담아 요구된 미디엄 숏보다 훨씬 넓고, 어두운 구석에 웅크린 대치보다는 창고 중앙의 대치에 가깝다.",
        "entities": "B-200의 짙은 회색 중장갑, 붉은 센서, 두꺼운 관절은 참조와 대체로 맞지만 한쪽 팔만 다연장 포신이고 반대쪽은 일반 기계 손이다. 찰리의 흰 마스크형 얼굴, 베이지 장갑, 가슴 고리는 보이나 참조의 모자와 긴 외투가 없고 체형도 달라졌다. 가슴 주변 초록 불꽃은 손상을 표현하지만 참조에 없는 강조 효과다. 축구공은 보이지 않으며 아래쪽 일부가 잘려 소지 여부는 확정하기 어렵다. 새끼 새는 보이지 않아 은폐 설정과 충돌하지 않는다. 판독 가능한 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "B-200은 벌린 두 발로 콘크리트 바닥을 딛고 있으며 무릎을 약간 굽혔지만 웅크린 상태는 아니다. 포신은 팔꿈치와 전완 기구에 연결되어 지지된다. 찰리의 발은 프레임 밖이지만 몸통과 다리의 연결은 자연스럽고 공중에 떠 있다는 증거는 없다. 뒤의 로봇 잔해는 바닥과 서로의 몸체에 기대어 쌓여 있다."
       },
       {
        "label": "A",
        "direction": "높이 든 포신은 오른쪽 B-200에서 왼쪽 전경 찰리의 상체 쪽으로 향하며, 포신 축을 연장하면 찰리의 어깨 아래 몸통에 닿는 방향이다. 낮은 반대쪽 포신도 왼쪽 찰리 쪽을 향한다. 총열의 옆면이 읽히고 카메라 자체를 겨누는 구도는 아니다. B-200의 얼굴은 찰리를 향하며 찰리도 등을 카메라에 둔 채 B-200 쪽으로 고개를 돌리고 있다.",
        "built_space": "왼쪽 부분 개방 출입문 하나, 높은 창 하나, 천장 등 세 개, 뒤쪽 부품 선반군, 오른쪽 벽 창과 가장자리 선반이 장소 참조와 대응한다. B-200은 오른쪽 벽 가까이, 찰리는 왼쪽 전경에 있어 대치 관계가 명확하다. 오른쪽 큰 상자 위로 로봇 잔해가 높게 쌓여 있으며 상자 옆 틈에는 새끼 새가 노출되어 있다. B-200의 전신과 넓은 바닥을 담아 요구된 미디엄 숏보다 넓다.",
        "entities": "B-200은 짙은 회색 중장비형 장갑과 양팔의 포신을 갖춰 주요 정체성이 맞지만 포신 구성은 참조보다 단순하다. 찰리는 베이지색 기계 몸체지만 긴 다리와 좁은 상체 때문에 고릴라형 비례가 약하고, 참조의 모자와 외투도 없다. 등 쪽 원형 부품은 보이나 가슴은 가려져 가슴 고리 유지 여부는 판정할 수 없다. 오른손 가까이에 축구공이 있다. 부서진 로봇 더미와 상자는 보이지만 새끼 새 세 마리가 드러나 은폐 상태를 지키지 못한다. 찰리의 표면은 금속 장갑보다 선으로 윤곽을 강조한 의상처럼 보여 실사 재질감이 약하다. 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "B-200은 양발을 넓게 바닥에 붙이고 무릎과 고관절을 굽혀 몸을 낮춘다. 완전히 깊게 웅크리지는 않았지만 중량을 지탱하는 자세는 성립한다. 두 포신은 각각 전완 기구에 연결되어 있다. 찰리의 발은 프레임 밖이며 지상에 선 자세로 읽힌다. 축구공은 오른손 손가락이 윗부분과 옆면에 닿아 붙잡은 상태로 보이고 허벅지에도 가까워 무지지 부유로 단정할 근거는 없다. 잔해는 상자와 서로의 부품에 받쳐져 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.6
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.6
   },
   "violations": {
    "A": [
     "[gemini-pro] 창작된 인물 (찰리 레퍼런스와 전혀 다른 외형의 캐릭터 등장)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1600,
   "A": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1600,
    "verdict_ko": "실사 배경 묘사와 B-200의 정조준 자세는 매우 훌륭하게 구현되었으나, 찰리의 지정된 의상과 축구공이 누락됨."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "실사 영화 스틸컷 조건에 위배되는 2D 그래픽 스타일이며, 찰리의 외형이 레퍼런스와 완전히 달라 치명적인 감점 요소임.  ★위반: [gemini-pro] 창작된 인물 (찰리 레퍼런스와 전혀 다른 외형의 캐릭터 등장)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L238B01.png",
    "asset_id": "e1a9f48b-d9e5-4c90-99a8-fadcfb66b2d8",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-4d3c-7d46-9aa4-82cb053632cc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7__bgfirst_bg.png",
   "bg_asset_id": "2c542339-d519-40a3-aee9-40b352e75792",
   "bg_record_key": "S64sh7::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C06"
  ]
 },
 "S64sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:03:38.053416+00:00",
  "fingerprint": "2f35c43e45c08f99d1df9656f3de51aa26b6583adb2d2fe8184c56aa36f6480c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S64sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S64sh7_sel.png",
  "source_sha256": "9377897e8d48975f3cece1e031f988053493d842cd948ce788cce8894fdb6571",
  "file": "S64sh7_cine.png",
  "staged_sha256": "4739d69d4f05763ff97f335ab43058833d352b16fcca1d049b8abf525f3d8a19",
  "latency_ms": 9963
 },
 "S64sh15::signage": {
  "fp": "1b6306e467d60a7b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S64sh15": {
  "input_fingerprint": "ec96bf46d5fe0ea9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑에 달린 둥근 링을 향해 두꺼운 손가락을 뻗은 B-200의 손 클로즈업.\n\nLOCATION (lock): In the same dim corner of the large robot-parts warehouse, near discarded machines and the half-open daylight doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ring attached to 찰리's chest (Attached to the chest armor and being indicated by B-200) — Its outward-facing shape is readable at a slight oblique angle on 찰리's chest; used as Provides the small shared focal point, kept in focus with the fingertip and visibly attached to the surrounding torso; 찰리's metal chest armor (Visible around the attached ring) — Seen from the open side of the seated pair, with the chest surface receding obliquely; used as Supplies bodily context and scale so the ring does not become an isolated enlarged object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's subdued ambient exposure, allowing restrained surface highlights to separate the pointing hand from 찰리's metal chest armor.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse door remains half open, the broken robots remain piled in a corner, and the baby birds are still concealed behind the box. B-200 retains his converted gun-hands, while Charlie remains seated nearby with a damaged chest ring, dents, holes and an impaired language system.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑에 달린 둥근 링을 향해 두꺼운 손가락을 뻗은 B-200의 손 클로즈업.\n\nLOCATION (lock): In the same dim corner of the large robot-parts warehouse, near discarded machines and the half-open daylight doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ring attached to 찰리's chest (Attached to the chest armor and being indicated by B-200) — Its outward-facing shape is readable at a slight oblique angle on 찰리's chest; used as Provides the small shared focal point, kept in focus with the fingertip and visibly attached to the surrounding torso; 찰리's metal chest armor (Visible around the attached ring) — Seen from the open side of the seated pair, with the chest surface receding obliquely; used as Supplies bodily context and scale so the ring does not become an isolated enlarged object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's subdued ambient exposure, allowing restrained surface highlights to separate the pointing hand from 찰리's metal chest armor.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse door remains half open, the broken robots remain piled in a corner, and the baby birds are still concealed behind the box. B-200 retains his converted gun-hands, while Charlie remains seated nearby with a damaged chest ring, dents, holes and an impaired language system.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑에 달린 둥근 링을 향해 두꺼운 손가락을 뻗은 B-200의 손 클로즈업.\n\nLOCATION (lock): In the same dim corner of the large robot-parts warehouse, near discarded machines and the half-open daylight doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ring attached to 찰리's chest (Attached to the chest armor and being indicated by B-200) — Its outward-facing shape is readable at a slight oblique angle on 찰리's chest; used as Provides the small shared focal point, kept in focus with the fingertip and visibly attached to the surrounding torso; 찰리's metal chest armor (Visible around the attached ring) — Seen from the open side of the seated pair, with the chest surface receding obliquely; used as Supplies bodily context and scale so the ring does not become an isolated enlarged object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's subdued ambient exposure, allowing restrained surface highlights to separate the pointing hand from 찰리's metal chest armor.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse door remains half open, the broken robots remain piled in a corner, and the baby birds are still concealed behind the box. B-200 retains his converted gun-hands, while Charlie remains seated nearby with a damaged chest ring, dents, holes and an impaired language system.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "화면 왼쪽에서 나타난 베이지색 기계 손이 화면 오른쪽을 향해 있는 찰리의 가슴 링을 가리키고 있음.",
    "built_space": "어두운 창고 내부이나, 이전 샷에서 확립된 문의 위치나 배경의 로봇 더미 배치가 일치하지 않음.",
    "entities": "찰리의 가슴 장갑은 베이지색으로 손상이 잘 표현되었으나, 가리키는 손은 B-200이어야 함에도 짙은 회색이 아닌 베이지색으로 묘사되었으며 고사포 포신 특성이 전혀 없음.",
    "hard_violations": [
     "[gemini-pro] B-200의 손이 짙은 회색이 아닌 찰리의 장갑과 같은 베이지색으로 잘못 묘사됨",
     "[gemini-pro] 지시문에 명시된 B-200의 고사포 포신(gun-hands) 형태 누락",
     "[gemini-pro] 이전 샷에서 확립된 캐릭터들의 좌우 위치가 반대로 뒤바뀜 (찰리가 왼쪽, B-200이 오른쪽이어야 함)"
    ],
    "physics": "프레임 밖에서 뻗어 나온 팔이 손을 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "화면 오른쪽에서 뻗어나온 짙은 회색 기계 손의 두꺼운 손가락이 화면 왼쪽 찰리의 가슴에 있는 링을 가리키고 있음.",
    "built_space": "어두운 창고 내부로, 배경에 부서진 기계 더미와 반쯤 열린 문이 이전 샷과 동일한 공간적 위치에 묘사되어 있음.",
    "entities": "B-200의 손은 짙은 회색이며 지시된 대로 팔목 주위에 고사포 포신(gun-hands)이 결합된 채로 묘사됨. 찰리의 샌드 베이지색 가슴 장갑과 찌그러진 링이 텍스트에 맞게 손상된 채로 표현됨.",
    "hard_violations": [],
    "physics": "B-200의 팔이 프레임 밖에서 뻗어 나와 손을 지탱하고 있으며, 찰리는 안정적으로 앉아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "B-200의 고사포 포신 형태와 짙은 회색 질감을 정확히 유지하면서, 이전 샷의 공간 배치와 요구된 클로즈업 프레이밍을 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "B-200의 짙은 회색 장갑과 고사포 포신 특성을 누락하고 손을 베이지색으로 렌더링했으며, 이전 샷에 확립된 캐릭터 간의 좌우 위치를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "화면 오른쪽에서 뻗어나온 짙은 회색 기계 손의 두꺼운 손가락이 화면 왼쪽 찰리의 가슴에 있는 링을 가리키고 있음.",
        "built_space": "어두운 창고 내부로, 배경에 부서진 기계 더미와 반쯤 열린 문이 이전 샷과 동일한 공간적 위치에 묘사되어 있음.",
        "entities": "B-200의 손은 짙은 회색이며 지시된 대로 팔목 주위에 고사포 포신(gun-hands)이 결합된 채로 묘사됨. 찰리의 샌드 베이지색 가슴 장갑과 찌그러진 링이 텍스트에 맞게 손상된 채로 표현됨.",
        "hard_violations": [],
        "physics": "B-200의 팔이 프레임 밖에서 뻗어 나와 손을 지탱하고 있으며, 찰리는 안정적으로 앉아 있음."
       },
       {
        "label": "A",
        "direction": "화면 왼쪽에서 나타난 베이지색 기계 손이 화면 오른쪽을 향해 있는 찰리의 가슴 링을 가리키고 있음.",
        "built_space": "어두운 창고 내부이나, 이전 샷에서 확립된 문의 위치나 배경의 로봇 더미 배치가 일치하지 않음.",
        "entities": "찰리의 가슴 장갑은 베이지색으로 손상이 잘 표현되었으나, 가리키는 손은 B-200이어야 함에도 짙은 회색이 아닌 베이지색으로 묘사되었으며 고사포 포신 특성이 전혀 없음.",
        "hard_violations": [
         "B-200의 손이 짙은 회색이 아닌 찰리의 장갑과 같은 베이지색으로 잘못 묘사됨",
         "지시문에 명시된 B-200의 고사포 포신(gun-hands) 형태 누락",
         "이전 샷에서 확립된 캐릭터들의 좌우 위치가 반대로 뒤바뀜 (찰리가 왼쪽, B-200이 오른쪽이어야 함)"
        ],
        "physics": "프레임 밖에서 뻗어 나온 팔이 손을 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "B-200의 고사포 포신 형태와 짙은 회색 질감을 정확히 유지하면서, 이전 샷의 공간 배치와 요구된 클로즈업 프레이밍을 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "B-200의 짙은 회색 장갑과 고사포 포신 특성을 누락하고 손을 베이지색으로 렌더링했으며, 이전 샷에 확립된 캐릭터 간의 좌우 위치를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 오른쪽에서 뻗어나온 짙은 회색 기계 손의 두꺼운 손가락이 화면 왼쪽 찰리의 가슴에 있는 링을 가리키고 있음.",
        "built_space": "어두운 창고 내부로, 배경에 부서진 기계 더미와 반쯤 열린 문이 이전 샷과 동일한 공간적 위치에 묘사되어 있음.",
        "entities": "B-200의 손은 짙은 회색이며 지시된 대로 팔목 주위에 고사포 포신(gun-hands)이 결합된 채로 묘사됨. 찰리의 샌드 베이지색 가슴 장갑과 찌그러진 링이 텍스트에 맞게 손상된 채로 표현됨.",
        "hard_violations": [],
        "physics": "B-200의 팔이 프레임 밖에서 뻗어 나와 손을 지탱하고 있으며, 찰리는 안정적으로 앉아 있음."
       },
       {
        "label": "A",
        "direction": "화면 왼쪽에서 나타난 베이지색 기계 손이 화면 오른쪽을 향해 있는 찰리의 가슴 링을 가리키고 있음.",
        "built_space": "어두운 창고 내부이나, 이전 샷에서 확립된 문의 위치나 배경의 로봇 더미 배치가 일치하지 않음.",
        "entities": "찰리의 가슴 장갑은 베이지색으로 손상이 잘 표현되었으나, 가리키는 손은 B-200이어야 함에도 짙은 회색이 아닌 베이지색으로 묘사되었으며 고사포 포신 특성이 전혀 없음.",
        "hard_violations": [
         "B-200의 손이 짙은 회색이 아닌 찰리의 장갑과 같은 베이지색으로 잘못 묘사됨",
         "지시문에 명시된 B-200의 고사포 포신(gun-hands) 형태 누락",
         "이전 샷에서 확립된 캐릭터들의 좌우 위치가 반대로 뒤바뀜 (찰리가 왼쪽, B-200이 오른쪽이어야 함)"
        ],
        "physics": "프레임 밖에서 뻗어 나온 팔이 손을 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두꺼운 손가락과 손상된 가슴 링을 밀접한 클로즈업으로 묶고 창고의 공간감을 유지하지만, B-200의 손 장갑이 짙은 회색보다 베이지색에 가깝다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "회색 기계 손과 포신, 링을 향한 지시는 잘 구현했지만, 배경 비중이 크고 찰리의 링 손상이 약해 핵심 순간의 충실도가 B보다 낮다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽에서 나온 기계 손의 굵은 검지가 왼쪽 아래로 뻗어 찰리의 가슴 링 안쪽을 향한다. 손끝과 링 사이에는 간격이 있다. 손목 주변 포신들도 대체로 왼쪽 전방을 향하지만, 발사 동작은 없다. 찰리의 눈은 프레임 밖이어서 시선은 확인할 수 없다.",
        "built_space": "찰리의 상체가 왼쪽, B-200의 손과 무장 일부가 오른쪽을 차지한다. 배경에는 밝은 출입구 하나, 뒤쪽 작업 선반, 오른쪽 높은 수납 선반과 폐기 로봇 더미가 보인다. 금속 지붕과 낡은 벽 재질은 참조와 유사하지만, 출입구가 뒤쪽 중앙에 보여 참조의 왼쪽 출입구와 공간 배치가 덜 일치한다. 좌석과 엉덩이는 보이지 않아 착석 접촉은 검증할 수 없다.",
        "entities": "찰리의 샌드 베이지 장갑, 흰 마스크의 아래 부분, 어두운 내부 구조가 참조와 맞는다. 가슴에는 몸통에 결합된 원형 금속 링 하나가 있으나 테두리의 파손은 뚜렷하지 않다. B-200의 손은 굵은 기계 관절과 회색 금속 장갑으로 표현되고 다연장 포신도 보인다. 찰리의 장갑에는 긁힘과 작은 구멍이 있다. 새와 상자는 식별되지 않으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "검지는 손바닥과 관절로 연결되고 손은 오른쪽 프레임 밖으로 이어지는 무장 손목에 연결되어 있어 지지 관계가 자연스럽다. 링은 가슴 장갑의 원형 장착부에 고정되어 있다. 찰리의 보이는 팔은 몸통 옆에서 아래로 굽혀져 있으며, 폐기 로봇들은 바닥과 서로의 몸체에 받쳐져 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 손의 굵은 검지가 오른쪽으로 뻗어 찰리의 가슴 링 왼쪽 가장자리와 내부를 향한다. 손끝은 링 바로 앞에서 멈추며, 지시 대상이 명확하다. 찰리의 눈과 B-200의 얼굴은 프레임 밖이고, 포구도 보이지 않아 시선과 포신 방향은 평가할 수 없다.",
        "built_space": "손이 왼쪽 전경, 찰리의 비스듬한 가슴이 오른쪽을 차지하며 링 주변 몸통이 충분히 남아 실제 크기를 설명한다. 배경 왼쪽에는 밝은 출입구 하나와 그 옆 높은 창 하나가 있고, 뒤쪽에는 작업 선반과 천장 철골이 보인다. 참조 창고의 출입구·고창·선반 관계가 유지된다. 폐기물 더미와 좌석은 이 클로즈업에서 확인되지 않는다.",
        "entities": "찰리의 베이지색 각진 장갑과 흰 마스크 아래 부분이 참조와 맞는다. 원형 링 하나가 가슴에 부착되어 있고, 링 왼쪽 테두리의 깨짐과 주변 장갑의 여러 관통 구멍이 선명하다. B-200의 손은 두꺼운 기계 손가락과 육중한 회색 손목을 갖지만 손등과 손가락 장갑은 베이지색에 가까워 참조의 짙은 회색 정체성이 약하다. 포신은 보이지 않으며 프레임 밖의 보유 여부는 판단할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뻗은 검지는 관절을 통해 손바닥에 연결되고 나머지 손가락은 자연스럽게 굽혀져 있다. 손목은 왼쪽의 굵은 전완 장치와 이어져 지지가 분명하다. 손상된 링도 장갑의 장착부에 고정되어 있다. 찰리의 하단 팔은 굽혀진 상태지만 좌석과 하체의 지지점은 프레임 밖이므로 착석 자세 전체는 검증할 수 없다. 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두꺼운 손가락과 손상된 가슴 링을 밀접한 클로즈업으로 묶고 창고의 공간감을 유지하지만, B-200의 손 장갑이 짙은 회색보다 베이지색에 가깝다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "회색 기계 손과 포신, 링을 향한 지시는 잘 구현했지만, 배경 비중이 크고 찰리의 링 손상이 약해 핵심 순간의 충실도가 B보다 낮다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽에서 나온 기계 손의 굵은 검지가 왼쪽 아래로 뻗어 찰리의 가슴 링 안쪽을 향한다. 손끝과 링 사이에는 간격이 있다. 손목 주변 포신들도 대체로 왼쪽 전방을 향하지만, 발사 동작은 없다. 찰리의 눈은 프레임 밖이어서 시선은 확인할 수 없다.",
        "built_space": "찰리의 상체가 왼쪽, B-200의 손과 무장 일부가 오른쪽을 차지한다. 배경에는 밝은 출입구 하나, 뒤쪽 작업 선반, 오른쪽 높은 수납 선반과 폐기 로봇 더미가 보인다. 금속 지붕과 낡은 벽 재질은 참조와 유사하지만, 출입구가 뒤쪽 중앙에 보여 참조의 왼쪽 출입구와 공간 배치가 덜 일치한다. 좌석과 엉덩이는 보이지 않아 착석 접촉은 검증할 수 없다.",
        "entities": "찰리의 샌드 베이지 장갑, 흰 마스크의 아래 부분, 어두운 내부 구조가 참조와 맞는다. 가슴에는 몸통에 결합된 원형 금속 링 하나가 있으나 테두리의 파손은 뚜렷하지 않다. B-200의 손은 굵은 기계 관절과 회색 금속 장갑으로 표현되고 다연장 포신도 보인다. 찰리의 장갑에는 긁힘과 작은 구멍이 있다. 새와 상자는 식별되지 않으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "검지는 손바닥과 관절로 연결되고 손은 오른쪽 프레임 밖으로 이어지는 무장 손목에 연결되어 있어 지지 관계가 자연스럽다. 링은 가슴 장갑의 원형 장착부에 고정되어 있다. 찰리의 보이는 팔은 몸통 옆에서 아래로 굽혀져 있으며, 폐기 로봇들은 바닥과 서로의 몸체에 받쳐져 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 손의 굵은 검지가 오른쪽으로 뻗어 찰리의 가슴 링 왼쪽 가장자리와 내부를 향한다. 손끝은 링 바로 앞에서 멈추며, 지시 대상이 명확하다. 찰리의 눈과 B-200의 얼굴은 프레임 밖이고, 포구도 보이지 않아 시선과 포신 방향은 평가할 수 없다.",
        "built_space": "손이 왼쪽 전경, 찰리의 비스듬한 가슴이 오른쪽을 차지하며 링 주변 몸통이 충분히 남아 실제 크기를 설명한다. 배경 왼쪽에는 밝은 출입구 하나와 그 옆 높은 창 하나가 있고, 뒤쪽에는 작업 선반과 천장 철골이 보인다. 참조 창고의 출입구·고창·선반 관계가 유지된다. 폐기물 더미와 좌석은 이 클로즈업에서 확인되지 않는다.",
        "entities": "찰리의 베이지색 각진 장갑과 흰 마스크 아래 부분이 참조와 맞는다. 원형 링 하나가 가슴에 부착되어 있고, 링 왼쪽 테두리의 깨짐과 주변 장갑의 여러 관통 구멍이 선명하다. B-200의 손은 두꺼운 기계 손가락과 육중한 회색 손목을 갖지만 손등과 손가락 장갑은 베이지색에 가까워 참조의 짙은 회색 정체성이 약하다. 포신은 보이지 않으며 프레임 밖의 보유 여부는 판단할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뻗은 검지는 관절을 통해 손바닥에 연결되고 나머지 손가락은 자연스럽게 굽혀져 있다. 손목은 왼쪽의 굵은 전완 장치와 이어져 지지가 분명하다. 손상된 링도 장갑의 장착부에 고정되어 있다. 찰리의 하단 팔은 굽혀진 상태지만 좌석과 하체의 지지점은 프레임 밖이므로 착석 자세 전체는 검증할 수 없다. 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.2,
    "B": 1.875
   },
   "adjusted": {
    "A": 0.95,
    "B": 1.875
   },
   "violations": {
    "A": [
     "[gemini-pro] B-200의 손이 짙은 회색이 아닌 찰리의 장갑과 같은 베이지색으로 잘못 묘사됨",
     "[gemini-pro] 지시문에 명시된 B-200의 고사포 포신(gun-hands) 형태 누락",
     "[gemini-pro] 이전 샷에서 확립된 캐릭터들의 좌우 위치가 반대로 뒤바뀜 (찰리가 왼쪽, B-200이 오른쪽이어야 함)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 950
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "B-200의 고사포 포신 형태와 짙은 회색 질감을 정확히 유지하면서, 이전 샷의 공간 배치와 요구된 클로즈업 프레이밍을 완벽하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 950,
    "verdict_ko": "B-200의 짙은 회색 장갑과 고사포 포신 특성을 누락하고 손을 베이지색으로 렌더링했으며, 이전 샷에 확립된 캐릭터 간의 좌우 위치를 위반했습니다.  ★위반: [gemini-pro] B-200의 손이 짙은 회색이 아닌 찰리의 장갑과 같은 베이지색으로 잘못 묘사됨 / [gemini-pro] 지시문에 명시된 B-200의 고사포 포신(gun-hands) 형태 누락 / [gemini-pro] 이전 샷에서 확립된 캐릭터들의 좌우 위치가 반대로 뒤바뀜 (찰리가 왼쪽, B-200이 오른쪽이어야 함)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of B-200, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7_sel.png",
    "asset_id": "8798972d-850d-4150-a31b-80583ad02a1d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-50b2-7778-a591-39440750340b",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S64sh7"
  }
 },
 "S64sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:04:41.786512+00:00",
  "fingerprint": "f00e6f495fa18dac3540ae22d1b72ed76b9cc76a98a63e74ef5c78ea73708695",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S64sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S64sh15_sel.png",
  "source_sha256": "77e35be4c6cc3dc20d669bc12c39b42c2faa0e9d80c01b5074bb4fa5390ca06f",
  "file": "S64sh15_cine.png",
  "staged_sha256": "a1dbad154c0c963d741b0826b8113075fb596d3e4aa086425803762b6bcbd343",
  "latency_ms": 10472
 },
 "S64sh24::signage": {
  "fp": "5f34007e464cce6a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S64sh24": {
  "input_fingerprint": "647a71e277e358a6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아기 새(B-200이 돌보던 새)들을 내려다보며 안광이 부드럽게 반달 모양으로 휘어진 B-200의 따뜻한 금속 얼굴.\n\nLOCATION (lock): Beside a concealed box of nestlings in the warehouse's shadowed corner, with daylight beyond the half-open entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Warehouse interior (Remains behind B-200 after 찰리 has left); used as A minimally resolved background keeps the face and the birds' small silhouettes readable without introducing new objects.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the warehouse illumination subdued and continuous, letting the softened crescent eye shapes convey warmth without adding an unsupported warm light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shifted box exposes the baby birds in B-200's warehouse corner; the broken-robot heap and half-open door remain unchanged. B-200 retains his converted gun-hands, and Charlie has left the corner with his battle damage still unrepaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아기 새(B-200이 돌보던 새)들을 내려다보며 안광이 부드럽게 반달 모양으로 휘어진 B-200의 따뜻한 금속 얼굴.\n\nLOCATION (lock): Beside a concealed box of nestlings in the warehouse's shadowed corner, with daylight beyond the half-open entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Warehouse interior (Remains behind B-200 after 찰리 has left); used as A minimally resolved background keeps the face and the birds' small silhouettes readable without introducing new objects.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the warehouse illumination subdued and continuous, letting the softened crescent eye shapes convey warmth without adding an unsupported warm light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shifted box exposes the baby birds in B-200's warehouse corner; the broken-robot heap and half-open door remain unchanged. B-200 retains his converted gun-hands, and Charlie has left the corner with his battle damage still unrepaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아기 새(B-200이 돌보던 새)들을 내려다보며 안광이 부드럽게 반달 모양으로 휘어진 B-200의 따뜻한 금속 얼굴.\n\nLOCATION (lock): Beside a concealed box of nestlings in the warehouse's shadowed corner, with daylight beyond the half-open entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Warehouse interior (Remains behind B-200 after 찰리 has left); used as A minimally resolved background keeps the face and the birds' small silhouettes readable without introducing new objects.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the warehouse illumination subdued and continuous, letting the softened crescent eye shapes convey warmth without adding an unsupported warm light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shifted box exposes the baby birds in B-200's warehouse corner; the broken-robot heap and half-open door remain unchanged. B-200 retains his converted gun-hands, and Charlie has left the corner with his battle damage still unrepaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "로봇이 고개를 숙여 하단의 상자와 아기 새들을 바라보고 있음.",
    "built_space": "창고 내부의 반쯤 열린 문은 보이나, 반드시 유지해야 하는 고장난 로봇 더미가 배경에서 완전히 누락됨.",
    "entities": "눈이 아래로 볼록한 U자 형태로 그려져 따뜻한 인상을 주지 못하며, 레퍼런스의 기계 손가락이 없고 총신만 렌더링됨.",
    "hard_violations": [
     "[gemini-pro] 이전 샷에서 유지되어야 할 고장난 로봇 더미 누락",
     "[gemini-pro] 레퍼런스에 묘사된 기계 손가락이 생략되고 무기 부분만 렌더링됨",
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 지침 위반 (뺨과 어깨 부위에 문자 표기됨)",
     "[gpt-high] 얼굴 하단 장갑의 영문 표기가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
    ],
    "physics": "새들은 상자 안에 안정적으로 놓여 있으나, 로봇의 총신 팔은 명확히 쥐거나 지탱하는 부분 없이 허공 및 상자 위에 걸쳐 있음."
   },
   {
    "label": "B",
    "direction": "로봇의 고개와 시선이 하단 상자 안의 아기 새들을 명확히 향하고 있음.",
    "built_space": "창고 내부 구조를 따르며, 열린 문과 우측의 고장난 로봇 더미가 이전 샷과 일치하게 배치됨.",
    "entities": "눈이 위로 볼록한 반달 모양으로 휘어져 따뜻함을 주며, 기계 손과 총신이 모두 올바르게 묘사됨. 아기 새의 모습도 레퍼런스와 일치함.",
    "hard_violations": [
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 지침 위반 (가슴 장갑판에 문자 표기됨)",
     "[gpt-high] 가슴 장갑의 기체 이름이 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
    ],
    "physics": "로봇의 오른손이 상자 가장자리를 자연스럽게 쥐고 지탱하며, 새들은 상자 바닥에 안정적으로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "따뜻한 반달 모양의 눈빛과 이전 샷의 배경(문, 로봇 더미)을 완벽히 유지했으나, 표면에 읽을 수 있는 텍스트가 포함된 점이 감점 요인임."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "눈 모양이 프롬프트의 의도(따뜻한 반달)와 다르고 필수 배경 요소인 로봇 더미를 누락했으며, 기계 손가락이 생략되어 완성도가 떨어짐."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "로봇의 고개와 시선이 하단 상자 안의 아기 새들을 명확히 향하고 있음.",
        "built_space": "창고 내부 구조를 따르며, 열린 문과 우측의 고장난 로봇 더미가 이전 샷과 일치하게 배치됨.",
        "entities": "눈이 위로 볼록한 반달 모양으로 휘어져 따뜻함을 주며, 기계 손과 총신이 모두 올바르게 묘사됨. 아기 새의 모습도 레퍼런스와 일치함.",
        "hard_violations": [
         "화면 내 읽을 수 있는 텍스트 금지 지침 위반 (가슴 장갑판에 문자 표기됨)"
        ],
        "physics": "로봇의 오른손이 상자 가장자리를 자연스럽게 쥐고 지탱하며, 새들은 상자 바닥에 안정적으로 서 있음."
       },
       {
        "label": "A",
        "direction": "로봇이 고개를 숙여 하단의 상자와 아기 새들을 바라보고 있음.",
        "built_space": "창고 내부의 반쯤 열린 문은 보이나, 반드시 유지해야 하는 고장난 로봇 더미가 배경에서 완전히 누락됨.",
        "entities": "눈이 아래로 볼록한 U자 형태로 그려져 따뜻한 인상을 주지 못하며, 레퍼런스의 기계 손가락이 없고 총신만 렌더링됨.",
        "hard_violations": [
         "이전 샷에서 유지되어야 할 고장난 로봇 더미 누락",
         "레퍼런스에 묘사된 기계 손가락이 생략되고 무기 부분만 렌더링됨",
         "화면 내 읽을 수 있는 텍스트 금지 지침 위반 (뺨과 어깨 부위에 문자 표기됨)"
        ],
        "physics": "새들은 상자 안에 안정적으로 놓여 있으나, 로봇의 총신 팔은 명확히 쥐거나 지탱하는 부분 없이 허공 및 상자 위에 걸쳐 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "따뜻한 반달 모양의 눈빛과 이전 샷의 배경(문, 로봇 더미)을 완벽히 유지했으나, 표면에 읽을 수 있는 텍스트가 포함된 점이 감점 요인임."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "눈 모양이 프롬프트의 의도(따뜻한 반달)와 다르고 필수 배경 요소인 로봇 더미를 누락했으며, 기계 손가락이 생략되어 완성도가 떨어짐."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "로봇의 고개와 시선이 하단 상자 안의 아기 새들을 명확히 향하고 있음.",
        "built_space": "창고 내부 구조를 따르며, 열린 문과 우측의 고장난 로봇 더미가 이전 샷과 일치하게 배치됨.",
        "entities": "눈이 위로 볼록한 반달 모양으로 휘어져 따뜻함을 주며, 기계 손과 총신이 모두 올바르게 묘사됨. 아기 새의 모습도 레퍼런스와 일치함.",
        "hard_violations": [
         "화면 내 읽을 수 있는 텍스트 금지 지침 위반 (가슴 장갑판에 문자 표기됨)"
        ],
        "physics": "로봇의 오른손이 상자 가장자리를 자연스럽게 쥐고 지탱하며, 새들은 상자 바닥에 안정적으로 서 있음."
       },
       {
        "label": "A",
        "direction": "로봇이 고개를 숙여 하단의 상자와 아기 새들을 바라보고 있음.",
        "built_space": "창고 내부의 반쯤 열린 문은 보이나, 반드시 유지해야 하는 고장난 로봇 더미가 배경에서 완전히 누락됨.",
        "entities": "눈이 아래로 볼록한 U자 형태로 그려져 따뜻한 인상을 주지 못하며, 레퍼런스의 기계 손가락이 없고 총신만 렌더링됨.",
        "hard_violations": [
         "이전 샷에서 유지되어야 할 고장난 로봇 더미 누락",
         "레퍼런스에 묘사된 기계 손가락이 생략되고 무기 부분만 렌더링됨",
         "화면 내 읽을 수 있는 텍스트 금지 지침 위반 (뺨과 어깨 부위에 문자 표기됨)"
        ],
        "physics": "새들은 상자 안에 안정적으로 놓여 있으나, 로봇의 총신 팔은 명확히 쥐거나 지탱하는 부분 없이 허공 및 상자 위에 걸쳐 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "부드러운 반달 눈과 새를 내려다보는 동작은 맞지만, 얼굴보다 상체와 주변 공간을 넓게 담았고 가슴의 판독 가능한 문자가 무문자 조건을 위반합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 중심의 더 밀착된 구도와 아래쪽 새를 향한 고개, 양팔 포신이 요구에 더 가깝지만, 얼굴 장갑의 판독 가능한 영문 표기로 최종 사용에는 부적합합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇은 고개를 화면 오른쪽 아래로 숙여 전경 상자 속 새들을 내려다본다. 새 네 마리 중 일부는 부리를 위쪽이나 로봇 쪽으로 돌리고 있다. 오른쪽 포신은 새가 아니라 화면 오른쪽 바깥을 향하며, 왼쪽 기계 손가락은 상자 가장자리에 걸쳐 있다.",
        "built_space": "뒤쪽 중앙에 밝은 출입구 하나, 오른쪽 벽에 창틀 하나와 고장 난 로봇 더미 하나가 보인다. 전경에는 새가 든 상자 하나와 오른쪽 목재 구조물 일부가 있다. 철제 벽과 천장, 콘크리트 바닥은 이전 장면의 창고와 대체로 이어지지만, 배경과 상체가 상당히 선명하고 넓게 보여 얼굴 중심의 최소 배경 요구는 약하다.",
        "entities": "회색 금속 장갑, 각진 머리, 붉은 보조 센서와 굵은 관절은 B-200의 정체성에 부합한다. 붉은 안광 두 개는 위로 둥글게 휜 웃는 눈 형태다. 한쪽 포신과 반대쪽 기계 손이 보이며, 이 손 형태는 이전 장면에도 존재한다. 작은 부리와 둥근 몸통을 가진 새 네 마리가 보이지만 어두워 참조의 깃털 색은 확인하기 어렵다. 찰리는 없다. 가슴에는 기체 이름이 읽히는 표기가 있다.",
        "hard_violations": [
         "가슴 장갑의 기체 이름이 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
        ],
        "physics": "머리는 목 관절에, 팔과 포신은 몸체에 연결되어 있다. 기계 손가락은 상자 테두리에 접촉한다. 새들의 하체는 상자에 가려져 있지만 상자 안에 앉아 있는 배치이며 공중에 떠 있다는 증거는 없다. 로봇의 하체 지지는 화면 밖이므로 판단할 수 없다."
       },
       {
        "label": "B",
        "direction": "로봇의 머리와 안면은 오른쪽 아래 전경의 새들을 향한다. 새들은 대체로 왼쪽 로봇 방향 또는 전방을 바라본다. 양팔 포신은 화면 오른쪽 위로 비껴 나가며 새를 직접 겨누지 않는다. 눈의 곡선은 아래로 둥근 형태여서 A의 웃는 눈보다 졸린 표정에 가깝다.",
        "built_space": "중앙 뒤쪽에 밝은 출입구 하나, 오른쪽 벽에 창틀 하나, 왼쪽 뒤에 창 일부가 보인다. 전경의 새 상자 하나 뒤로 양팔 포신이 지나간다. 고장 난 로봇 더미가 있던 오른쪽 아래 영역은 포신과 상자에 가려져 있어 유지 여부를 확인할 수 없다. 창고 재질과 낮의 외광은 이어지며, A보다 얼굴을 크게 담아 클로즈업 요구에 더 가깝다.",
        "entities": "짙은 회색의 마모된 장갑, 각진 머리, 붉은 원형 센서와 양팔의 다연장 포신은 B-200 참조와 부합한다. 안광은 주황빛의 둥근 반달 형태로, 참조의 붉은색에서 달라졌지만 비인간 로봇의 발광 눈이라는 설정에는 맞는다. 상자 속에는 둥근 몸통과 작은 부리를 가진 갈색·황갈색 새 네 마리가 보인다. 찰리나 다른 인물은 없다. 얼굴 하단에는 판독 가능한 영문 표기가 있다.",
        "hard_violations": [
         "얼굴 하단 장갑의 영문 표기가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
        ],
        "physics": "머리와 몸체는 목 기구로 연결되고 두 포신 묶음은 각각 팔 구조에 연결되어 있다. 새들은 상자 내부에 몸을 낮춘 상태이며 하체와 접지점은 테두리에 가려져 있다. 지지 없이 떠 있는 새나 분리된 포신은 보이지 않는다. 로봇의 발은 프레임 밖이므로 접지는 평가할 수 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "부드러운 반달 눈과 새를 내려다보는 동작은 맞지만, 얼굴보다 상체와 주변 공간을 넓게 담았고 가슴의 판독 가능한 문자가 무문자 조건을 위반합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "얼굴 중심의 더 밀착된 구도와 아래쪽 새를 향한 고개, 양팔 포신이 요구에 더 가깝지만, 얼굴 장갑의 판독 가능한 영문 표기로 최종 사용에는 부적합합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "로봇은 고개를 화면 오른쪽 아래로 숙여 전경 상자 속 새들을 내려다본다. 새 네 마리 중 일부는 부리를 위쪽이나 로봇 쪽으로 돌리고 있다. 오른쪽 포신은 새가 아니라 화면 오른쪽 바깥을 향하며, 왼쪽 기계 손가락은 상자 가장자리에 걸쳐 있다.",
        "built_space": "뒤쪽 중앙에 밝은 출입구 하나, 오른쪽 벽에 창틀 하나와 고장 난 로봇 더미 하나가 보인다. 전경에는 새가 든 상자 하나와 오른쪽 목재 구조물 일부가 있다. 철제 벽과 천장, 콘크리트 바닥은 이전 장면의 창고와 대체로 이어지지만, 배경과 상체가 상당히 선명하고 넓게 보여 얼굴 중심의 최소 배경 요구는 약하다.",
        "entities": "회색 금속 장갑, 각진 머리, 붉은 보조 센서와 굵은 관절은 B-200의 정체성에 부합한다. 붉은 안광 두 개는 위로 둥글게 휜 웃는 눈 형태다. 한쪽 포신과 반대쪽 기계 손이 보이며, 이 손 형태는 이전 장면에도 존재한다. 작은 부리와 둥근 몸통을 가진 새 네 마리가 보이지만 어두워 참조의 깃털 색은 확인하기 어렵다. 찰리는 없다. 가슴에는 기체 이름이 읽히는 표기가 있다.",
        "hard_violations": [
         "가슴 장갑의 기체 이름이 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
        ],
        "physics": "머리는 목 관절에, 팔과 포신은 몸체에 연결되어 있다. 기계 손가락은 상자 테두리에 접촉한다. 새들의 하체는 상자에 가려져 있지만 상자 안에 앉아 있는 배치이며 공중에 떠 있다는 증거는 없다. 로봇의 하체 지지는 화면 밖이므로 판단할 수 없다."
       },
       {
        "label": "A",
        "direction": "로봇의 머리와 안면은 오른쪽 아래 전경의 새들을 향한다. 새들은 대체로 왼쪽 로봇 방향 또는 전방을 바라본다. 양팔 포신은 화면 오른쪽 위로 비껴 나가며 새를 직접 겨누지 않는다. 눈의 곡선은 아래로 둥근 형태여서 A의 웃는 눈보다 졸린 표정에 가깝다.",
        "built_space": "중앙 뒤쪽에 밝은 출입구 하나, 오른쪽 벽에 창틀 하나, 왼쪽 뒤에 창 일부가 보인다. 전경의 새 상자 하나 뒤로 양팔 포신이 지나간다. 고장 난 로봇 더미가 있던 오른쪽 아래 영역은 포신과 상자에 가려져 있어 유지 여부를 확인할 수 없다. 창고 재질과 낮의 외광은 이어지며, A보다 얼굴을 크게 담아 클로즈업 요구에 더 가깝다.",
        "entities": "짙은 회색의 마모된 장갑, 각진 머리, 붉은 원형 센서와 양팔의 다연장 포신은 B-200 참조와 부합한다. 안광은 주황빛의 둥근 반달 형태로, 참조의 붉은색에서 달라졌지만 비인간 로봇의 발광 눈이라는 설정에는 맞는다. 상자 속에는 둥근 몸통과 작은 부리를 가진 갈색·황갈색 새 네 마리가 보인다. 찰리나 다른 인물은 없다. 얼굴 하단에는 판독 가능한 영문 표기가 있다.",
        "hard_violations": [
         "얼굴 하단 장갑의 영문 표기가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
        ],
        "physics": "머리와 몸체는 목 기구로 연결되고 두 포신 묶음은 각각 팔 구조에 연결되어 있다. 새들은 상자 내부에 몸을 낮춘 상태이며 하체와 접지점은 테두리에 가려져 있다. 지지 없이 떠 있는 새나 분리된 포신은 보이지 않는다. 로봇의 발은 프레임 밖이므로 접지는 평가할 수 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.5
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 지침 위반 (가슴 장갑판에 문자 표기됨)",
     "[gpt-high] 가슴 장갑의 기체 이름이 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
    ],
    "A": [
     "[gemini-pro] 이전 샷에서 유지되어야 할 고장난 로봇 더미 누락",
     "[gemini-pro] 레퍼런스에 묘사된 기계 손가락이 생략되고 무기 부분만 렌더링됨",
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 지침 위반 (뺨과 어깨 부위에 문자 표기됨)",
     "[gpt-high] 얼굴 하단 장갑의 영문 표기가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1500,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "따뜻한 반달 모양의 눈빛과 이전 샷의 배경(문, 로봇 더미)을 완벽히 유지했으나, 표면에 읽을 수 있는 텍스트가 포함된 점이 감점 요인임.  ★위반: [gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 지침 위반 (가슴 장갑판에 문자 표기됨) / [gpt-high] 가슴 장갑의 기체 이름이 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "눈 모양이 프롬프트의 의도(따뜻한 반달)와 다르고 필수 배경 요소인 로봇 더미를 누락했으며, 기계 손가락이 생략되어 완성도가 떨어짐.  ★위반: [gemini-pro] 이전 샷에서 유지되어야 할 고장난 로봇 더미 누락 / [gemini-pro] 레퍼런스에 묘사된 기계 손가락이 생략되고 무기 부분만 렌더링됨 / [gemini-pro] 화면 내 읽을 수 있는 텍스트 금지 지침 위반 (뺨과 어깨 부위에 문자 표기됨) / [gpt-high] 얼굴 하단 장갑의 영문 표기가 판독 가능하여, 읽을 수 있는 글자를 전혀 허용하지 않는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh15_sel.png",
    "asset_id": "f86baf4e-e930-4a9b-868c-74a03e203679",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:833571>",
    "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-5267-7fa1-a0b0-04e702f5520c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S64sh15"
  },
  "staged_characters_added": [
   "C64"
  ]
 },
 "S64sh24::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:05:59.633822+00:00",
  "fingerprint": "1d75afdae1984144ff9c301406f4058e88181956a5b942589807b7962e35c882",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S64sh24_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S64sh24_sel.png",
  "source_sha256": "8d703c1b300996ecb1cb1cd5ce4eaeef0a41f61bbfddd413129e783845384922",
  "file": "S64sh24_cine.png",
  "staged_sha256": "4618b6900e5f6ba6d1adff27d252b57fede6f18274ce6aea75ae4f17daeddee3",
  "latency_ms": 9902
 },
 "S65sh5::signage": {
  "fp": "3dfdfb28a150f0da",
  "inscriptions": [
   {
    "text_native": "0",
    "source": "scene_text_quoted",
    "reason_ko": "방사능 측정기 눈금이 가리키고 있는 숫자가 클로즈업 화면에 표시되어야 합니다.",
    "source_quote": "0"
   }
  ],
  "cues": [],
  "dropped": []
 },
 "S65sh5": {
  "input_fingerprint": "c314461c4486dc0e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 방사능 측정기 눈금이 숫자 0에 정확히 멈춰 선 상태의 클로즈업.\n\nLOCATION (lock): On the overgrown approach to a partly collapsed department store, among moss, puddles, and uncontrolled vegetation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Radiation meter dial (The needle has settled at the zero marking) — The marked face is directed toward the camera at a readable downward oblique angle, showing the needle and numeral 0; used as The primary focal detail, held within the supporting hands and environmental context rather than enlarged to fill the image; Moss-covered ground (Moss spreads across the approach to the collapsed department store); used as Soft surrounding context connects the apparently anomalous reading to the living environment.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the exterior keeps the dial markings readable and the surrounding moss subdued, without glare obscuring the zero reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The department store is half collapsed and overgrown with moss and trees, with standing pools of water. The radiation meter settles near zero, rather than establishing an exact zero reading.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"0\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 방사능 측정기 눈금이 숫자 0에 정확히 멈춰 선 상태의 클로즈업.\n\nLOCATION (lock): On the overgrown approach to a partly collapsed department store, among moss, puddles, and uncontrolled vegetation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Radiation meter dial (The needle has settled at the zero marking) — The marked face is directed toward the camera at a readable downward oblique angle, showing the needle and numeral 0; used as The primary focal detail, held within the supporting hands and environmental context rather than enlarged to fill the image; Moss-covered ground (Moss spreads across the approach to the collapsed department store); used as Soft surrounding context connects the apparently anomalous reading to the living environment.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the exterior keeps the dial markings readable and the surrounding moss subdued, without glare obscuring the zero reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The department store is half collapsed and overgrown with moss and trees, with standing pools of water. The radiation meter settles near zero, rather than establishing an exact zero reading.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"0\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 방사능 측정기 눈금이 숫자 0에 정확히 멈춰 선 상태의 클로즈업.\n\nLOCATION (lock): On the overgrown approach to a partly collapsed department store, among moss, puddles, and uncontrolled vegetation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Radiation meter dial (The needle has settled at the zero marking) — The marked face is directed toward the camera at a readable downward oblique angle, showing the needle and numeral 0; used as The primary focal detail, held within the supporting hands and environmental context rather than enlarged to fill the image; Moss-covered ground (Moss spreads across the approach to the collapsed department store); used as Soft surrounding context connects the apparently anomalous reading to the living environment.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the exterior keeps the dial markings readable and the surrounding moss subdued, without glare obscuring the zero reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The department store is half collapsed and overgrown with moss and trees, with standing pools of water. The radiation meter settles near zero, rather than establishing an exact zero reading.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"0\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "측정기 계기판이 카메라를 향하며, 바늘은 중앙 상단을 가리킴.",
    "built_space": "이끼 낀 바닥이 아닌 레퍼런스 사진의 원경(건물과 연못)을 그대로 배경으로 사용함.",
    "entities": "측정기를 쥔 양손. 오른손에 엄지손가락이 2개 존재함. 계기판에 '0'이 두 번 표기됨.",
    "hard_violations": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
    ],
    "physics": "양손이 기기를 들고 지탱함."
   },
   {
    "label": "B",
    "direction": "측정기 계기판이 카메라를 비스듬히 향하며, 바늘이 정확히 0을 가리킴.",
    "built_space": "이끼 낀 바닥과 웅덩이가 배경을 채우며, 수면에 무너진 건물이 반사됨.",
    "entities": "측정기와 가죽 스트랩을 쥔 두 손. 계기판에 '0'이 하나만 명확히 표기됨.",
    "hard_violations": [],
    "physics": "양손이 기기를 안정적으로 쥐고 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 앵글, 이끼 낀 바닥 배경, 0을 가리키는 바늘 등 프롬프트의 요구사항을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "해부학적으로 불가능한 손가락 기형이 발생했으며, 이끼 낀 바닥 대신 레퍼런스 이미지의 넓은 구도를 그대로 복사하여 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "측정기 계기판이 카메라를 향하며, 바늘은 중앙 상단을 가리킴.",
        "built_space": "이끼 낀 바닥이 아닌 레퍼런스 사진의 원경(건물과 연못)을 그대로 배경으로 사용함.",
        "entities": "측정기를 쥔 양손. 오른손에 엄지손가락이 2개 존재함. 계기판에 '0'이 두 번 표기됨.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
        ],
        "physics": "양손이 기기를 들고 지탱함."
       },
       {
        "label": "B",
        "direction": "측정기 계기판이 카메라를 비스듬히 향하며, 바늘이 정확히 0을 가리킴.",
        "built_space": "이끼 낀 바닥과 웅덩이가 배경을 채우며, 수면에 무너진 건물이 반사됨.",
        "entities": "측정기와 가죽 스트랩을 쥔 두 손. 계기판에 '0'이 하나만 명확히 표기됨.",
        "hard_violations": [],
        "physics": "양손이 기기를 안정적으로 쥐고 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 앵글, 이끼 낀 바닥 배경, 0을 가리키는 바늘 등 프롬프트의 요구사항을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "해부학적으로 불가능한 손가락 기형이 발생했으며, 이끼 낀 바닥 대신 레퍼런스 이미지의 넓은 구도를 그대로 복사하여 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "측정기 계기판이 카메라를 향하며, 바늘은 중앙 상단을 가리킴.",
        "built_space": "이끼 낀 바닥이 아닌 레퍼런스 사진의 원경(건물과 연못)을 그대로 배경으로 사용함.",
        "entities": "측정기를 쥔 양손. 오른손에 엄지손가락이 2개 존재함. 계기판에 '0'이 두 번 표기됨.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
        ],
        "physics": "양손이 기기를 들고 지탱함."
       },
       {
        "label": "B",
        "direction": "측정기 계기판이 카메라를 비스듬히 향하며, 바늘이 정확히 0을 가리킴.",
        "built_space": "이끼 낀 바닥과 웅덩이가 배경을 채우며, 수면에 무너진 건물이 반사됨.",
        "entities": "측정기와 가죽 스트랩을 쥔 두 손. 계기판에 '0'이 하나만 명확히 표기됨.",
        "hard_violations": [],
        "physics": "양손이 기기를 안정적으로 쥐고 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "0의 시작 눈금에 놓인 바늘을 내려다보는 클로즈업으로, 손의 지지와 이끼 배경을 유지하면서 핵심 판독 상태를 가장 명료하게 구현했다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "바늘의 0 지시와 장소는 충실하지만, A보다 정면에 가까운 계기판 각도와 강조된 건물 배경, 중복된 0 표기가 핵심 클로즈업의 집중도를 낮춘다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "계기판은 카메라와 손을 뻗은 관찰자 쪽으로 기울어져 있으며, 위에서 비스듬히 내려다보는 방향으로 읽힌다. 바늘은 오른쪽 아래 회전축에서 왼쪽의 숫자 0 바로 위 시작 눈금을 향한다. 얼굴이나 시선은 보이지 않는다.",
        "built_space": "전경에는 이끼로 덮인 깨진 콘크리트와 작은 고인 물이 있고, 상단에는 큰 물웅덩이 하나가 보인다. 건물 본체 대신 물에 비친 콘크리트 외벽과 세로 창 구획이 나타난다. 맞은편 건물이 수면에 거꾸로 비치는 배치는 이 하향 시점에서 가능하다. 참고 장소의 재료와 식생은 맞지만 건물의 정확한 붕괴 형상은 직접 확인할 수 없다.",
        "entities": "낡은 금속 외함, 유리 덮개, 부채꼴 눈금과 바늘을 가진 휴대용 아날로그 측정기 한 대가 보인다. 기기 종류를 글자로 확인할 수는 없지만 방사능 측정기 소품으로 타당하다. 판독 가능한 문자는 숫자 0 하나이며 다른 문구는 없다. 흙 묻은 손 두 개와 팔 일부만 나오고 얼굴이나 추가 인물은 없다. 손만으로 민족·성별·정확한 나이를 판단할 수 없다. 이끼, 잡초, 물웅덩이와 잔해가 있으며 낮의 자연광으로 읽힌다.",
        "hard_violations": [],
        "physics": "왼손이 기기의 왼쪽과 아래를 받치고 오른손이 오른쪽 손잡이와 외함을 잡아 무게를 지지한다. 가죽 끈은 기기에 연결되고 손에 잡혀 있다. 바늘은 계기 내부 회전축에 연결되어 있으며, 떠 있는 물체나 불가능한 손 자세는 보이지 않는다. 유리의 반사는 있지만 0과 바늘 끝을 가리지 않는다."
       },
       {
        "label": "B",
        "direction": "계기판 앞면은 카메라 쪽으로 향하며 A보다 정면에 가깝게 보인다. 중앙 아래 회전축에서 뻗은 바늘은 위쪽의 작은 0에 해당하는 눈금을 가리킨다. 계기 중앙에도 큰 0이 있어 같은 숫자가 두 번 보인다. 얼굴이나 시선은 없다.",
        "built_space": "상단 중앙에 백화점 입구 한 곳과 그 앞 계단·접근로가 있고, 양쪽에는 무너진 콘크리트와 창 구획이 이어진다. 오른쪽에는 기둥으로 나뉜 창가 공간이 보이며, 접근로 주변에 식생과 여러 고인 물 구역이 있다. 참고 사진의 중앙 진입로와 붕괴된 입면 관계를 잘 유지한다. 인물은 건물 내부가 아니라 접근로에서 기기를 든 배치로 읽히며, 불가능한 반사나 명백히 중복된 고정 시설은 없다.",
        "entities": "낡은 직사각형 휴대용 아날로그 측정기 한 대, 이를 잡은 손 두 개, 회갈색 소매가 보인다. 방사능 측정기 소품으로 읽힐 수 있는 형태이며, 읽히는 문자는 두 곳의 0뿐이다. 추가 얼굴이나 전신 인물은 없다. 손만으로 민족·성별·정확한 나이를 확인할 수 없다. 이끼, 관목, 물웅덩이, 붕괴된 백화점이 낮의 자연광 아래 보인다.",
        "hard_violations": [],
        "physics": "왼손이 기기 왼쪽과 하단을 감싸 받치고 오른손이 오른쪽 외함을 잡고 있어 지지가 명확하다. 손목과 소매의 연결도 자연스럽다. 계기 바늘은 내부 축에서 뻗어 있고 잔해는 지면에 놓여 있다. 지지 없이 떠 있는 물체나 물리적으로 불가능한 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "0의 시작 눈금에 놓인 바늘을 내려다보는 클로즈업으로, 손의 지지와 이끼 배경을 유지하면서 핵심 판독 상태를 가장 명료하게 구현했다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "바늘의 0 지시와 장소는 충실하지만, A보다 정면에 가까운 계기판 각도와 강조된 건물 배경, 중복된 0 표기가 핵심 클로즈업의 집중도를 낮춘다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "계기판은 카메라와 손을 뻗은 관찰자 쪽으로 기울어져 있으며, 위에서 비스듬히 내려다보는 방향으로 읽힌다. 바늘은 오른쪽 아래 회전축에서 왼쪽의 숫자 0 바로 위 시작 눈금을 향한다. 얼굴이나 시선은 보이지 않는다.",
        "built_space": "전경에는 이끼로 덮인 깨진 콘크리트와 작은 고인 물이 있고, 상단에는 큰 물웅덩이 하나가 보인다. 건물 본체 대신 물에 비친 콘크리트 외벽과 세로 창 구획이 나타난다. 맞은편 건물이 수면에 거꾸로 비치는 배치는 이 하향 시점에서 가능하다. 참고 장소의 재료와 식생은 맞지만 건물의 정확한 붕괴 형상은 직접 확인할 수 없다.",
        "entities": "낡은 금속 외함, 유리 덮개, 부채꼴 눈금과 바늘을 가진 휴대용 아날로그 측정기 한 대가 보인다. 기기 종류를 글자로 확인할 수는 없지만 방사능 측정기 소품으로 타당하다. 판독 가능한 문자는 숫자 0 하나이며 다른 문구는 없다. 흙 묻은 손 두 개와 팔 일부만 나오고 얼굴이나 추가 인물은 없다. 손만으로 민족·성별·정확한 나이를 판단할 수 없다. 이끼, 잡초, 물웅덩이와 잔해가 있으며 낮의 자연광으로 읽힌다.",
        "hard_violations": [],
        "physics": "왼손이 기기의 왼쪽과 아래를 받치고 오른손이 오른쪽 손잡이와 외함을 잡아 무게를 지지한다. 가죽 끈은 기기에 연결되고 손에 잡혀 있다. 바늘은 계기 내부 회전축에 연결되어 있으며, 떠 있는 물체나 불가능한 손 자세는 보이지 않는다. 유리의 반사는 있지만 0과 바늘 끝을 가리지 않는다."
       },
       {
        "label": "A",
        "direction": "계기판 앞면은 카메라 쪽으로 향하며 A보다 정면에 가깝게 보인다. 중앙 아래 회전축에서 뻗은 바늘은 위쪽의 작은 0에 해당하는 눈금을 가리킨다. 계기 중앙에도 큰 0이 있어 같은 숫자가 두 번 보인다. 얼굴이나 시선은 없다.",
        "built_space": "상단 중앙에 백화점 입구 한 곳과 그 앞 계단·접근로가 있고, 양쪽에는 무너진 콘크리트와 창 구획이 이어진다. 오른쪽에는 기둥으로 나뉜 창가 공간이 보이며, 접근로 주변에 식생과 여러 고인 물 구역이 있다. 참고 사진의 중앙 진입로와 붕괴된 입면 관계를 잘 유지한다. 인물은 건물 내부가 아니라 접근로에서 기기를 든 배치로 읽히며, 불가능한 반사나 명백히 중복된 고정 시설은 없다.",
        "entities": "낡은 직사각형 휴대용 아날로그 측정기 한 대, 이를 잡은 손 두 개, 회갈색 소매가 보인다. 방사능 측정기 소품으로 읽힐 수 있는 형태이며, 읽히는 문자는 두 곳의 0뿐이다. 추가 얼굴이나 전신 인물은 없다. 손만으로 민족·성별·정확한 나이를 확인할 수 없다. 이끼, 관목, 물웅덩이, 붕괴된 백화점이 낮의 자연광 아래 보인다.",
        "hard_violations": [],
        "physics": "왼손이 기기 왼쪽과 하단을 감싸 받치고 오른손이 오른쪽 외함을 잡고 있어 지지가 명확하다. 손목과 소매의 연결도 자연스럽다. 계기 바늘은 내부 축에서 뻗어 있고 잔해는 지면에 놓여 있다. 지지 없이 떠 있는 물체나 물리적으로 불가능한 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.317,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.067,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1067
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 앵글, 이끼 낀 바닥 배경, 0을 가리키는 바늘 등 프롬프트의 요구사항을 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1067,
    "verdict_ko": "해부학적으로 불가능한 손가락 기형이 발생했으며, 이끼 낀 바닥 대신 레퍼런스 이미지의 넓은 구도를 그대로 복사하여 감점되었습니다.  ★위반: [gemini-pro] 해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L239B01.png",
    "asset_id": "a44a0231-5865-442f-9e00-425f802c380d",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-5415-7a8d-8abf-2cee1b42af3f",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S65sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:00:05.873647+00:00",
  "fingerprint": "56311d2cdcddb5b0a15701247a0986dc765d95afde4a701e87ac9c0fbef3b0b1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S65sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S65sh5_sel.png",
  "source_sha256": "a3678d2fa6db32d80064cefaf8e81ad98b919c32f9429d9286ae157c0d098045",
  "file": "S65sh5_cine.png",
  "staged_sha256": "9df9f5122f408eb2893788d36c3becefedbde972415b9e95c32f12e4656f1bd1",
  "latency_ms": 10042
 },
 "S65sh10::signage": {
  "fp": "b6225f47af2d9930",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S65sh10": {
  "input_fingerprint": "6256b113154c1829",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 맨얼굴로 두 눈을 감고 깊게 숨을 들이마시는 현우의 평온한 얼굴.\n\nLOCATION (lock): Among moss-covered rubble outside the ruined department store, on the approach before entering the building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Collapsed department store (Partly collapsed by an earthquake) — A partial exterior section is visible obliquely behind 현우; used as A soft background edge maintains the dangerous setting against his newfound calm; Uncontrolled tree growth (Growing irregularly around the ruined department store); used as Provides a restrained background indication of returning life without crowding the facial close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight preserves natural skin detail and the supported color of the surrounding vegetation, allowing the peaceful expression to register without an added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed department store remains surrounded by moss, trees, flowers and pooled water, with birds and a water deer present. The radiation meter continues to indicate a value near zero. 현우: He has removed his gas mask and breathes with his face uncovered. His previous fighting injuries remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 맨얼굴로 두 눈을 감고 깊게 숨을 들이마시는 현우의 평온한 얼굴.\n\nLOCATION (lock): Among moss-covered rubble outside the ruined department store, on the approach before entering the building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Collapsed department store (Partly collapsed by an earthquake) — A partial exterior section is visible obliquely behind 현우; used as A soft background edge maintains the dangerous setting against his newfound calm; Uncontrolled tree growth (Growing irregularly around the ruined department store); used as Provides a restrained background indication of returning life without crowding the facial close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight preserves natural skin detail and the supported color of the surrounding vegetation, allowing the peaceful expression to register without an added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed department store remains surrounded by moss, trees, flowers and pooled water, with birds and a water deer present. The radiation meter continues to indicate a value near zero. 현우: He has removed his gas mask and breathes with his face uncovered. His previous fighting injuries remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 맨얼굴로 두 눈을 감고 깊게 숨을 들이마시는 현우의 평온한 얼굴.\n\nLOCATION (lock): Among moss-covered rubble outside the ruined department store, on the approach before entering the building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Collapsed department store (Partly collapsed by an earthquake) — A partial exterior section is visible obliquely behind 현우; used as A soft background edge maintains the dangerous setting against his newfound calm; Uncontrolled tree growth (Growing irregularly around the ruined department store); used as Provides a restrained background indication of returning life without crowding the facial close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight preserves natural skin detail and the supported color of the surrounding vegetation, allowing the peaceful expression to register without an added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed department store remains surrounded by moss, trees, flowers and pooled water, with birds and a water deer present. The radiation meter continues to indicate a value near zero. 현우: He has removed his gas mask and breathes with his face uncovered. His previous fighting injuries remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "정면을 향해 두 눈을 지그시 감고 있는 얼굴.",
    "built_space": "이끼 낀 돌과 흐릿한 건축물 잔해가 배경에 자리하지만 백화점의 형태가 명확하지 않음.",
    "entities": "현우. 하지만 레퍼런스에서 확인되는 셔츠가 보이지 않아 맨몸인 것으로 보이며, 얼굴에 남아 있어야 할 전투 부상의 흔적이 전혀 없이 깨끗함.",
    "hard_violations": [],
    "physics": "가만히 서 있는 머리와 어깨의 자연스러운 형태이며, 지지 오류는 없음."
   },
   {
    "label": "B",
    "direction": "고개를 약간 위로 든 상태로 두 눈을 감고 있으며, 깊게 숨을 들이마시는 평온한 얼굴이 카메라를 향함.",
    "built_space": "이끼 낀 콘크리트 잔해들 너머로 사선 방향 뒤쪽에 뼈대가 드러난 무너진 백화점 건물이 뚜렷하게 보이며, 주변으로 식물들이 자라나 있음.",
    "entities": "현우. 캐릭터 레퍼런스와 일치하는 짙은 회색 셔츠를 입고 있으며, 이마의 상처와 얼굴의 흙먼지 등 프롬프트가 요구한 전투 부상 흔적이 잘 묘사됨.",
    "hard_violations": [],
    "physics": "자연스럽게 서서 숨을 들이마시는 자세로, 물리적인 지지나 중력에 어긋나는 부분이 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "눈을 감고 깊게 숨을 들이마시는 표정을 훌륭히 묘사했으며, 지정된 의상과 얼굴의 상처, 등 뒤의 무너진 백화점 배경 등 프롬프트의 요구사항을 모두 완벽하게 충족했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캐릭터 레퍼런스의 지정된 의상을 입지 않고 상의를 탈의한 것처럼 묘사되었으며, 유지되어야 할 기존의 전투 부상 흔적도 완전히 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "고개를 약간 위로 든 상태로 두 눈을 감고 있으며, 깊게 숨을 들이마시는 평온한 얼굴이 카메라를 향함.",
        "built_space": "이끼 낀 콘크리트 잔해들 너머로 사선 방향 뒤쪽에 뼈대가 드러난 무너진 백화점 건물이 뚜렷하게 보이며, 주변으로 식물들이 자라나 있음.",
        "entities": "현우. 캐릭터 레퍼런스와 일치하는 짙은 회색 셔츠를 입고 있으며, 이마의 상처와 얼굴의 흙먼지 등 프롬프트가 요구한 전투 부상 흔적이 잘 묘사됨.",
        "hard_violations": [],
        "physics": "자연스럽게 서서 숨을 들이마시는 자세로, 물리적인 지지나 중력에 어긋나는 부분이 없음."
       },
       {
        "label": "A",
        "direction": "정면을 향해 두 눈을 지그시 감고 있는 얼굴.",
        "built_space": "이끼 낀 돌과 흐릿한 건축물 잔해가 배경에 자리하지만 백화점의 형태가 명확하지 않음.",
        "entities": "현우. 하지만 레퍼런스에서 확인되는 셔츠가 보이지 않아 맨몸인 것으로 보이며, 얼굴에 남아 있어야 할 전투 부상의 흔적이 전혀 없이 깨끗함.",
        "hard_violations": [],
        "physics": "가만히 서 있는 머리와 어깨의 자연스러운 형태이며, 지지 오류는 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "눈을 감고 깊게 숨을 들이마시는 표정을 훌륭히 묘사했으며, 지정된 의상과 얼굴의 상처, 등 뒤의 무너진 백화점 배경 등 프롬프트의 요구사항을 모두 완벽하게 충족했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캐릭터 레퍼런스의 지정된 의상을 입지 않고 상의를 탈의한 것처럼 묘사되었으며, 유지되어야 할 기존의 전투 부상 흔적도 완전히 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "고개를 약간 위로 든 상태로 두 눈을 감고 있으며, 깊게 숨을 들이마시는 평온한 얼굴이 카메라를 향함.",
        "built_space": "이끼 낀 콘크리트 잔해들 너머로 사선 방향 뒤쪽에 뼈대가 드러난 무너진 백화점 건물이 뚜렷하게 보이며, 주변으로 식물들이 자라나 있음.",
        "entities": "현우. 캐릭터 레퍼런스와 일치하는 짙은 회색 셔츠를 입고 있으며, 이마의 상처와 얼굴의 흙먼지 등 프롬프트가 요구한 전투 부상 흔적이 잘 묘사됨.",
        "hard_violations": [],
        "physics": "자연스럽게 서서 숨을 들이마시는 자세로, 물리적인 지지나 중력에 어긋나는 부분이 없음."
       },
       {
        "label": "A",
        "direction": "정면을 향해 두 눈을 지그시 감고 있는 얼굴.",
        "built_space": "이끼 낀 돌과 흐릿한 건축물 잔해가 배경에 자리하지만 백화점의 형태가 명확하지 않음.",
        "entities": "현우. 하지만 레퍼런스에서 확인되는 셔츠가 보이지 않아 맨몸인 것으로 보이며, 얼굴에 남아 있어야 할 전투 부상의 흔적이 전혀 없이 깨끗함.",
        "hard_violations": [],
        "physics": "가만히 서 있는 머리와 어깨의 자연스러운 형태이며, 지지 오류는 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "눈을 감고 숨을 들이마시는 연기와 회색 셔츠·상처는 충실하지만, 얼굴보다 넓고 선명하게 드러난 폐허가 얼굴 중심의 절제된 클로즈업 지시에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "맨얼굴과 감긴 두 눈을 중심으로 한 밀착 클로즈업과 흐릿한 폐허 배경이 우선 지시에 더 충실하지만, 드러난 어깨에 참조의 셔츠가 없고 전투 상처 표현도 약합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 화면 왼쪽 위를 향해 조금 들려 있고 두 눈은 완전히 감겨 있어 바라보는 대상은 없습니다. 턱을 들고 입 주변을 이완한 모습이 평온하게 숨을 들이마시는 순간으로 읽힙니다. 겨누거나 사용하는 물체는 없습니다.",
        "built_space": "인물은 건물 밖 잔해 앞에 있습니다. 왼쪽 배경에는 여러 층의 콘크리트 외벽 한 구간과 파손된 보·기둥, 노출 철근이 보이고, 아래에는 이끼 덮인 잔해, 오른쪽에는 나무가 있습니다. 참조의 콘크리트·이끼 환경과 부합하지만 건물의 노출 면적과 선명도가 커서 부드러운 배경 가장자리라는 지시에는 덜 맞습니다. 중복된 고정 설비나 반사는 없습니다.",
        "entities": "젊은 동아시아계 남성 한 명이며, 앳된 얼굴과 헝클어진 검은 머리, 회색 셔츠가 현우의 참조와 대체로 일치합니다. 국적은 외형만으로 확인할 수 없습니다. 방독면 없이 얼굴이 드러나 있고 눈썹 부근과 볼에 상처가 있습니다. 나무와 이끼, 붕괴된 건물은 보입니다. 계측기·새·고라니·고인 물은 이 얼굴 중심 화면에서 확인되지 않으며, 이를 보여주기 위해 구도를 넓힐 필요는 없습니다. 읽을 수 있는 글자나 추가 인물은 없습니다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 지지되고, 턱을 약간 드는 자세도 정상적인 호흡 동작으로 가능합니다. 발과 지면 접촉은 화면 밖이지만 공중에 뜬 몸이라는 징후는 없습니다. 콘크리트 잔해는 아래 잔해에 놓여 있고 철근은 파손된 구조물에 연결되어 있습니다."
       },
       {
        "label": "B",
        "direction": "얼굴은 거의 정면을 향하지만 두 눈이 완전히 감겨 있어 카메라를 응시하지 않습니다. 입술과 눈 주변은 편안하며 호흡을 가다듬는 순간으로 읽히지만, 깊게 들이마시는 동작의 단서는 A보다 약합니다. 방향을 확인할 휴대 물체나 무기는 없습니다.",
        "built_space": "얼굴 양옆의 좁은 배경에 이끼 덮인 잔해가 있고, 왼쪽 위에는 세로 창이 있는 밝은 외벽 한 구간, 그 아래에는 비스듬히 무너진 콘크리트 판이 보입니다. 오른쪽에는 나무와 불규칙한 식생이 흐리게 남습니다. 건물 밖 잔해 앞이라는 위치가 성립하며, 참조의 재료와 식생을 얼굴 뒤의 부드러운 배경으로 유지합니다. 중복 설비나 불가능한 반사는 없습니다.",
        "entities": "젊은 동아시아계 남성 한 명의 얼굴이 화면을 크게 채웁니다. 검은 머리와 얼굴 윤곽은 현우 참조에 대체로 부합하며, 맨얼굴과 감긴 두 눈이 명확합니다. 다만 화면 아래 목과 어깨가 맨살로 보여 참조의 회색 셔츠 연속성이 맞지 않습니다. 코와 피부에 작은 붉은 흔적은 있지만 전투 상처는 뚜렷하지 않습니다. 폐허·이끼·나무는 보이고 계측기·동물·고인 물은 프레임 밖입니다. 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "머리는 곧게 선 목과 어깨에 자연스럽게 연결되어 있습니다. 감긴 눈과 이완된 얼굴 근육에 해부학적 이상은 없습니다. 하체와 발은 클로즈업 밖이라 접지 상태를 직접 확인할 수 없지만 부유를 시사하는 모습도 없습니다. 배경의 기울어진 콘크리트 판은 아래 잔해에 걸쳐 있고, 지지 없이 떠 있는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "눈을 감고 숨을 들이마시는 연기와 회색 셔츠·상처는 충실하지만, 얼굴보다 넓고 선명하게 드러난 폐허가 얼굴 중심의 절제된 클로즈업 지시에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "맨얼굴과 감긴 두 눈을 중심으로 한 밀착 클로즈업과 흐릿한 폐허 배경이 우선 지시에 더 충실하지만, 드러난 어깨에 참조의 셔츠가 없고 전투 상처 표현도 약합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 화면 왼쪽 위를 향해 조금 들려 있고 두 눈은 완전히 감겨 있어 바라보는 대상은 없습니다. 턱을 들고 입 주변을 이완한 모습이 평온하게 숨을 들이마시는 순간으로 읽힙니다. 겨누거나 사용하는 물체는 없습니다.",
        "built_space": "인물은 건물 밖 잔해 앞에 있습니다. 왼쪽 배경에는 여러 층의 콘크리트 외벽 한 구간과 파손된 보·기둥, 노출 철근이 보이고, 아래에는 이끼 덮인 잔해, 오른쪽에는 나무가 있습니다. 참조의 콘크리트·이끼 환경과 부합하지만 건물의 노출 면적과 선명도가 커서 부드러운 배경 가장자리라는 지시에는 덜 맞습니다. 중복된 고정 설비나 반사는 없습니다.",
        "entities": "젊은 동아시아계 남성 한 명이며, 앳된 얼굴과 헝클어진 검은 머리, 회색 셔츠가 현우의 참조와 대체로 일치합니다. 국적은 외형만으로 확인할 수 없습니다. 방독면 없이 얼굴이 드러나 있고 눈썹 부근과 볼에 상처가 있습니다. 나무와 이끼, 붕괴된 건물은 보입니다. 계측기·새·고라니·고인 물은 이 얼굴 중심 화면에서 확인되지 않으며, 이를 보여주기 위해 구도를 넓힐 필요는 없습니다. 읽을 수 있는 글자나 추가 인물은 없습니다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 지지되고, 턱을 약간 드는 자세도 정상적인 호흡 동작으로 가능합니다. 발과 지면 접촉은 화면 밖이지만 공중에 뜬 몸이라는 징후는 없습니다. 콘크리트 잔해는 아래 잔해에 놓여 있고 철근은 파손된 구조물에 연결되어 있습니다."
       },
       {
        "label": "A",
        "direction": "얼굴은 거의 정면을 향하지만 두 눈이 완전히 감겨 있어 카메라를 응시하지 않습니다. 입술과 눈 주변은 편안하며 호흡을 가다듬는 순간으로 읽히지만, 깊게 들이마시는 동작의 단서는 A보다 약합니다. 방향을 확인할 휴대 물체나 무기는 없습니다.",
        "built_space": "얼굴 양옆의 좁은 배경에 이끼 덮인 잔해가 있고, 왼쪽 위에는 세로 창이 있는 밝은 외벽 한 구간, 그 아래에는 비스듬히 무너진 콘크리트 판이 보입니다. 오른쪽에는 나무와 불규칙한 식생이 흐리게 남습니다. 건물 밖 잔해 앞이라는 위치가 성립하며, 참조의 재료와 식생을 얼굴 뒤의 부드러운 배경으로 유지합니다. 중복 설비나 불가능한 반사는 없습니다.",
        "entities": "젊은 동아시아계 남성 한 명의 얼굴이 화면을 크게 채웁니다. 검은 머리와 얼굴 윤곽은 현우 참조에 대체로 부합하며, 맨얼굴과 감긴 두 눈이 명확합니다. 다만 화면 아래 목과 어깨가 맨살로 보여 참조의 회색 셔츠 연속성이 맞지 않습니다. 코와 피부에 작은 붉은 흔적은 있지만 전투 상처는 뚜렷하지 않습니다. 폐허·이끼·나무는 보이고 계측기·동물·고인 물은 프레임 밖입니다. 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "머리는 곧게 선 목과 어깨에 자연스럽게 연결되어 있습니다. 감긴 눈과 이완된 얼굴 근육에 해부학적 이상은 없습니다. 하체와 발은 클로즈업 밖이라 접지 상태를 직접 확인할 수 없지만 부유를 시사하는 모습도 없습니다. 배경의 기울어진 콘크리트 판은 아래 잔해에 걸쳐 있고, 지지 없이 떠 있는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.4,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.4,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1400
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "눈을 감고 깊게 숨을 들이마시는 표정을 훌륭히 묘사했으며, 지정된 의상과 얼굴의 상처, 등 뒤의 무너진 백화점 배경 등 프롬프트의 요구사항을 모두 완벽하게 충족했습니다."
   },
   {
    "label": "A",
    "score": 1400,
    "verdict_ko": "캐릭터 레퍼런스의 지정된 의상을 입지 않고 상의를 탈의한 것처럼 묘사되었으며, 유지되어야 할 기존의 전투 부상 흔적도 완전히 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S65sh5_sel.png",
    "asset_id": "e3e1a817-0d43-4ace-9d8f-51484a5940e9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-55c2-7b30-9cd2-561ce08a8225",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S65sh5"
  }
 },
 "S65sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:06:56.760511+00:00",
  "fingerprint": "af28e168d90c288f70ef06b2ee18c471dc08a5f7e78fe505767ac9db6a729d59",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S65sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S65sh10_sel.png",
  "source_sha256": "918c33b406bd1d958e8dda9a9f5080849b2fdf37b97b56ac21469538a0338c5a",
  "file": "S65sh10_cine.png",
  "staged_sha256": "1704ffd05193ef7ec1c7fbfe4b378ace18b036ebaed126d7574b9cdf9e7b1bf7",
  "latency_ms": 9677
 },
 "S65sh20::signage": {
  "fp": "51e81110944353f6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S65sh20": {
  "input_fingerprint": "d22e02c2280b2a5e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 높은 콘크리트 층 위에서 아래의 현우와 수빈을 몰래 내려다보는 낯선 인물의 어두운 뒷모습 전경.\n\nLOCATION (lock): On an elevated exposed concrete remnant of the collapsed department store, overlooking the overgrown approach below. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: High concrete floor edge beside the watcher in the lower-left of the frame, foreground; Lower opening being entered by the pair in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: High concrete floor (An elevated level within the partly collapsed building, occupied by the stranger) — Its near edge crosses the lower-left foreground, with the lower approach visible beyond it; used as Establishes the watcher's elevation and separates the foreground observer from the pair below; Opening into the collapsed department store (Open and being entered by 현우 and 수빈) — Seen steeply from above in the lower-right portion of the composition; used as The shared spatial anchor that makes the stranger's surveillance and the pair's inward route readable; Moss and irregular tree growth (Spread through the ruined approach around the building); used as Connects the distant lower space to the living environment established before the watcher reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight separates the readable lower approach from the stranger's darker foreground back, preserving anonymity without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed, moss-covered department store remains overgrown with trees and surrounded by pools of water and wildlife. The radiation reading established outside is near zero. 현우: His gas mask is off as he enters through the collapsed structure, retaining his earlier injuries. 수빈: Her gas mask is now off, leaving her scarred face uncovered as she enters the building. Her radiation-damaged torso remains unchanged. 한수: He watches from above; he has short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 높은 콘크리트 층 위에서 아래의 현우와 수빈을 몰래 내려다보는 낯선 인물의 어두운 뒷모습 전경.\n\nLOCATION (lock): On an elevated exposed concrete remnant of the collapsed department store, overlooking the overgrown approach below. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: High concrete floor edge beside the watcher in the lower-left of the frame, foreground; Lower opening being entered by the pair in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: High concrete floor (An elevated level within the partly collapsed building, occupied by the stranger) — Its near edge crosses the lower-left foreground, with the lower approach visible beyond it; used as Establishes the watcher's elevation and separates the foreground observer from the pair below; Opening into the collapsed department store (Open and being entered by 현우 and 수빈) — Seen steeply from above in the lower-right portion of the composition; used as The shared spatial anchor that makes the stranger's surveillance and the pair's inward route readable; Moss and irregular tree growth (Spread through the ruined approach around the building); used as Connects the distant lower space to the living environment established before the watcher reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight separates the readable lower approach from the stranger's darker foreground back, preserving anonymity without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed, moss-covered department store remains overgrown with trees and surrounded by pools of water and wildlife. The radiation reading established outside is near zero. 현우: His gas mask is off as he enters through the collapsed structure, retaining his earlier injuries. 수빈: Her gas mask is now off, leaving her scarred face uncovered as she enters the building. Her radiation-damaged torso remains unchanged. 한수: He watches from above; he has short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 높은 콘크리트 층 위에서 아래의 현우와 수빈을 몰래 내려다보는 낯선 인물의 어두운 뒷모습 전경.\n\nLOCATION (lock): On an elevated exposed concrete remnant of the collapsed department store, overlooking the overgrown approach below. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: High concrete floor edge beside the watcher in the lower-left of the frame, foreground; Lower opening being entered by the pair in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: High concrete floor (An elevated level within the partly collapsed building, occupied by the stranger) — Its near edge crosses the lower-left foreground, with the lower approach visible beyond it; used as Establishes the watcher's elevation and separates the foreground observer from the pair below; Opening into the collapsed department store (Open and being entered by 현우 and 수빈) — Seen steeply from above in the lower-right portion of the composition; used as The shared spatial anchor that makes the stranger's surveillance and the pair's inward route readable; Moss and irregular tree growth (Spread through the ruined approach around the building); used as Connects the distant lower space to the living environment established before the watcher reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight separates the readable lower approach from the stranger's darker foreground back, preserving anonymity without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed, moss-covered department store remains overgrown with trees and surrounded by pools of water and wildlife. The radiation reading established outside is near zero. 현우: His gas mask is off as he enters through the collapsed structure, retaining his earlier injuries. 수빈: Her gas mask is now off, leaving her scarred face uncovered as she enters the building. Her radiation-damaged torso remains unchanged. 한수: He watches from above; he has short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "전경의 관찰자가 아래를 내려다보나, 아래의 두 인물이 관찰자를 올려다보고 있어 '몰래' 상황이 실패함.",
    "built_space": "좌측 전경에 높은 콘크리트 바닥, 우측 배경에 터널 형태의 입구가 배치되어 레이아웃 요건을 충족함.",
    "entities": "관찰자는 짧은 검은 머리와 파카를 착용함. 아래 여성은 레퍼런스(남색 스웨터)와 다른 갈색 상의를 착용함.",
    "hard_violations": [],
    "physics": "모든 인물이 바닥에 서 있음. 물리적으로 불가능하거나 지지되지 않는 객체 없음."
   },
   {
    "label": "B",
    "direction": "관찰자가 아래를 내려다보고, 두 인물은 등지고 숲을 향해 걸어가고 있어 '몰래' 지켜보는 상황이 유지됨.",
    "built_space": "좌측 전경에 높은 콘크리트 바닥이 있으나, 우측 배경에 텍스트가 요구한 '건물 입구' 구조물이 존재하지 않음.",
    "entities": "관찰자는 비니를 착용해 프롬프트의 '짧은 머리'가 가려짐. 아래 두 인물의 의상은 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "모든 인물이 지면을 딛고 서 있음. 허공에 떠 있거나 지지되지 않는 객체 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물이 등지고 걸어가는 모습을 통해 '몰래' 지켜보는 핵심 상황과 의상을 잘 구현했으나, 우측 하단의 명시된 건물 입구 구조물이 누락되었습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "우측 하단에 건물 입구는 묘사되었으나, 아래의 두 인물이 관찰자를 대놓고 올려다보고 있어 텍스트의 핵심 행동을 완전히 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "전경의 관찰자가 아래를 내려다보나, 아래의 두 인물이 관찰자를 올려다보고 있어 '몰래' 상황이 실패함.",
        "built_space": "좌측 전경에 높은 콘크리트 바닥, 우측 배경에 터널 형태의 입구가 배치되어 레이아웃 요건을 충족함.",
        "entities": "관찰자는 짧은 검은 머리와 파카를 착용함. 아래 여성은 레퍼런스(남색 스웨터)와 다른 갈색 상의를 착용함.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 서 있음. 물리적으로 불가능하거나 지지되지 않는 객체 없음."
       },
       {
        "label": "B",
        "direction": "관찰자가 아래를 내려다보고, 두 인물은 등지고 숲을 향해 걸어가고 있어 '몰래' 지켜보는 상황이 유지됨.",
        "built_space": "좌측 전경에 높은 콘크리트 바닥이 있으나, 우측 배경에 텍스트가 요구한 '건물 입구' 구조물이 존재하지 않음.",
        "entities": "관찰자는 비니를 착용해 프롬프트의 '짧은 머리'가 가려짐. 아래 두 인물의 의상은 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "모든 인물이 지면을 딛고 서 있음. 허공에 떠 있거나 지지되지 않는 객체 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물이 등지고 걸어가는 모습을 통해 '몰래' 지켜보는 핵심 상황과 의상을 잘 구현했으나, 우측 하단의 명시된 건물 입구 구조물이 누락되었습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "우측 하단에 건물 입구는 묘사되었으나, 아래의 두 인물이 관찰자를 대놓고 올려다보고 있어 텍스트의 핵심 행동을 완전히 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "전경의 관찰자가 아래를 내려다보나, 아래의 두 인물이 관찰자를 올려다보고 있어 '몰래' 상황이 실패함.",
        "built_space": "좌측 전경에 높은 콘크리트 바닥, 우측 배경에 터널 형태의 입구가 배치되어 레이아웃 요건을 충족함.",
        "entities": "관찰자는 짧은 검은 머리와 파카를 착용함. 아래 여성은 레퍼런스(남색 스웨터)와 다른 갈색 상의를 착용함.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 서 있음. 물리적으로 불가능하거나 지지되지 않는 객체 없음."
       },
       {
        "label": "B",
        "direction": "관찰자가 아래를 내려다보고, 두 인물은 등지고 숲을 향해 걸어가고 있어 '몰래' 지켜보는 상황이 유지됨.",
        "built_space": "좌측 전경에 높은 콘크리트 바닥이 있으나, 우측 배경에 텍스트가 요구한 '건물 입구' 구조물이 존재하지 않음.",
        "entities": "관찰자는 비니를 착용해 프롬프트의 '짧은 머리'가 가려짐. 아래 두 인물의 의상은 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "모든 인물이 지면을 딛고 서 있음. 허공에 떠 있거나 지지되지 않는 객체 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "어두운 감시자의 뒷모습과 좌하단 고층 바닥·우하단의 두 사람 배치는 더 충실하지만, 두 사람이 건물 입구로 들어가지 않고 외부 길을 향한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "높낮이와 열린 입구는 명확하지만, 입구가 우하단보다 위에 있고 두 사람이 감시자 쪽을 올려다보며 서 있어 몰래 내려다보는 진입 순간과 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "한수는 등을 보인 채 머리를 오른쪽 아래 두 사람이 있는 방향으로 돌리고 있다. 현우와 수빈은 카메라에 등을 보이며 화면 위쪽의 나무 사이 외부 통로로 향한다. 왼쪽 폐건물의 어두운 개구부로 들어가는 방향은 아니며, 두 사람의 진행 방향에 진입구가 보이지 않는다. 무기나 방향을 확인할 휴대 도구는 없다.",
        "built_space": "좌하단 전경에 한수가 선 높은 콘크리트 바닥과 깨진 가장자리 하나가 있고, 그 너머 낮은 접근로의 우하단에 두 사람이 있다. 왼쪽에는 상층의 뚜렷한 사각 기둥 세 개와 파손된 층 슬래브, 아래층의 어두운 개구부들이 보인다. 노출 철근·회색 콘크리트·이끼·수목은 이전 사진과 연결되지만, 지정된 우하단의 진입구는 구현되지 않았다. 뚜렷한 물웅덩이나 반사는 확인되지 않는다.",
        "entities": "사람은 성인 남성 감시자 한 명과 아래쪽의 젊은 남녀 두 명으로 총 세 명이다. 한수의 녹색 외투·갈색 바지·니트 모자는 참고 이미지와 대체로 맞지만 외투의 털 장식과 배낭은 보이지 않고, 짧은 검은 머리는 모자로 가려져 있다. 현우의 검은 머리와 회갈색 셔츠는 이전 사진에 가깝다. 수빈은 검은 단발이지만 상의가 참고 이미지의 남색 니트와 다른 회색 계열이다. 두 사람 모두 뒷모습이라 얼굴의 정확한 나이·정체성·상처와 방독면 착용 여부는 판정하기 어렵고, 수빈의 몸통 손상도 옷에 가려진다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "한수의 발은 화면 밖이지만 다리가 전경 콘크리트 바닥 위로 자연스럽게 이어져 서 있는 자세로 읽힌다. 아래 두 사람은 이끼 낀 잔해 통로에 발을 딛고 있으며 현우의 벌어진 다리는 보행 순간으로 가능하다. 잔해는 지면이나 남은 구조체에 놓여 있고, 떠 있는 인체나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "한수의 머리와 몸은 오른쪽 아래의 두 사람을 향한다. 그러나 현우와 수빈은 입구 바깥에서 정면을 드러낸 채 머리를 위쪽 감시자 방향으로 들고 있다. 건물 안으로 향하는 뒷모습이나 진입 동작이 아니라 감시자를 발견하고 마주 보는 순간에 가깝다. 무기나 손에 든 도구는 없다.",
        "built_space": "전경의 높은 콘크리트 바닥 가장자리 하나에 한수가 서 있고, 왼쪽에는 큰 사각 기둥과 뒤쪽 기둥 두 개, 파손된 상층 슬래브와 노출 철근이 보인다. 오른쪽에는 콘크리트 테두리의 큰 직사각형 입구 하나가 있으며 두 사람은 그 문턱 바깥 낮은 바닥에 서 있다. 입구와 두 사람은 지정된 우하단보다 높은 우측 중상단에 놓인다. 이끼·수목·파손 콘크리트는 참고 장소의 재질과 이어지지만 입구의 정확한 형태는 이전 사진에서 확인되지 않는다. 아래 물웅덩이의 하늘과 식생 반사는 이 시점에서 가능하다.",
        "entities": "총 세 사람만 있다. 한수는 짧은 검은 머리와 녹색 털 장식 외투·갈색 바지를 갖춰 인물 설정에 가깝고, 얼굴은 뒷모습으로 숨겨져 있다. 현우는 검은 머리의 젊은 동아시아계 남성으로 회갈색 셔츠를 입었고, 수빈은 검은 단발의 젊은 동아시아계 여성으로 보인다. 두 사람의 얼굴은 방독면 없이 드러나지만 거리상 정확한 얼굴 일치와 흉터는 검증하기 어렵다. 수빈은 참고 이미지의 남색 니트 대신 몸통 일부를 드러내는 손상된 어두운 옷을 입었다. 물웅덩이는 보이지만 야생동물은 식별되지 않으며, 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "한수의 두 부츠는 높은 콘크리트 바닥에 닿아 있고 무게를 지탱한다. 아래 두 사람도 입구 앞 바닥에 발을 딛고 서 있어 신체 지지는 정상이다. 다만 그 자세는 들어가는 보행보다 멈춰 올려다보는 동작에 가깝다. 큰 파편들은 지면이나 구조물에 기대어 있으며 공중에 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "어두운 감시자의 뒷모습과 좌하단 고층 바닥·우하단의 두 사람 배치는 더 충실하지만, 두 사람이 건물 입구로 들어가지 않고 외부 길을 향한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "높낮이와 열린 입구는 명확하지만, 입구가 우하단보다 위에 있고 두 사람이 감시자 쪽을 올려다보며 서 있어 몰래 내려다보는 진입 순간과 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "한수는 등을 보인 채 머리를 오른쪽 아래 두 사람이 있는 방향으로 돌리고 있다. 현우와 수빈은 카메라에 등을 보이며 화면 위쪽의 나무 사이 외부 통로로 향한다. 왼쪽 폐건물의 어두운 개구부로 들어가는 방향은 아니며, 두 사람의 진행 방향에 진입구가 보이지 않는다. 무기나 방향을 확인할 휴대 도구는 없다.",
        "built_space": "좌하단 전경에 한수가 선 높은 콘크리트 바닥과 깨진 가장자리 하나가 있고, 그 너머 낮은 접근로의 우하단에 두 사람이 있다. 왼쪽에는 상층의 뚜렷한 사각 기둥 세 개와 파손된 층 슬래브, 아래층의 어두운 개구부들이 보인다. 노출 철근·회색 콘크리트·이끼·수목은 이전 사진과 연결되지만, 지정된 우하단의 진입구는 구현되지 않았다. 뚜렷한 물웅덩이나 반사는 확인되지 않는다.",
        "entities": "사람은 성인 남성 감시자 한 명과 아래쪽의 젊은 남녀 두 명으로 총 세 명이다. 한수의 녹색 외투·갈색 바지·니트 모자는 참고 이미지와 대체로 맞지만 외투의 털 장식과 배낭은 보이지 않고, 짧은 검은 머리는 모자로 가려져 있다. 현우의 검은 머리와 회갈색 셔츠는 이전 사진에 가깝다. 수빈은 검은 단발이지만 상의가 참고 이미지의 남색 니트와 다른 회색 계열이다. 두 사람 모두 뒷모습이라 얼굴의 정확한 나이·정체성·상처와 방독면 착용 여부는 판정하기 어렵고, 수빈의 몸통 손상도 옷에 가려진다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "한수의 발은 화면 밖이지만 다리가 전경 콘크리트 바닥 위로 자연스럽게 이어져 서 있는 자세로 읽힌다. 아래 두 사람은 이끼 낀 잔해 통로에 발을 딛고 있으며 현우의 벌어진 다리는 보행 순간으로 가능하다. 잔해는 지면이나 남은 구조체에 놓여 있고, 떠 있는 인체나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "한수의 머리와 몸은 오른쪽 아래의 두 사람을 향한다. 그러나 현우와 수빈은 입구 바깥에서 정면을 드러낸 채 머리를 위쪽 감시자 방향으로 들고 있다. 건물 안으로 향하는 뒷모습이나 진입 동작이 아니라 감시자를 발견하고 마주 보는 순간에 가깝다. 무기나 손에 든 도구는 없다.",
        "built_space": "전경의 높은 콘크리트 바닥 가장자리 하나에 한수가 서 있고, 왼쪽에는 큰 사각 기둥과 뒤쪽 기둥 두 개, 파손된 상층 슬래브와 노출 철근이 보인다. 오른쪽에는 콘크리트 테두리의 큰 직사각형 입구 하나가 있으며 두 사람은 그 문턱 바깥 낮은 바닥에 서 있다. 입구와 두 사람은 지정된 우하단보다 높은 우측 중상단에 놓인다. 이끼·수목·파손 콘크리트는 참고 장소의 재질과 이어지지만 입구의 정확한 형태는 이전 사진에서 확인되지 않는다. 아래 물웅덩이의 하늘과 식생 반사는 이 시점에서 가능하다.",
        "entities": "총 세 사람만 있다. 한수는 짧은 검은 머리와 녹색 털 장식 외투·갈색 바지를 갖춰 인물 설정에 가깝고, 얼굴은 뒷모습으로 숨겨져 있다. 현우는 검은 머리의 젊은 동아시아계 남성으로 회갈색 셔츠를 입었고, 수빈은 검은 단발의 젊은 동아시아계 여성으로 보인다. 두 사람의 얼굴은 방독면 없이 드러나지만 거리상 정확한 얼굴 일치와 흉터는 검증하기 어렵다. 수빈은 참고 이미지의 남색 니트 대신 몸통 일부를 드러내는 손상된 어두운 옷을 입었다. 물웅덩이는 보이지만 야생동물은 식별되지 않으며, 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "한수의 두 부츠는 높은 콘크리트 바닥에 닿아 있고 무게를 지탱한다. 아래 두 사람도 입구 앞 바닥에 발을 딛고 서 있어 신체 지지는 정상이다. 다만 그 자세는 들어가는 보행보다 멈춰 올려다보는 동작에 가깝다. 큰 파편들은 지면이나 구조물에 기대어 있으며 공중에 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.229,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.229,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1229
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "두 인물이 등지고 걸어가는 모습을 통해 '몰래' 지켜보는 핵심 상황과 의상을 잘 구현했으나, 우측 하단의 명시된 건물 입구 구조물이 누락되었습니다."
   },
   {
    "label": "A",
    "score": 1229,
    "verdict_ko": "우측 하단에 건물 입구는 묘사되었으나, 아래의 두 인물이 관찰자를 대놓고 올려다보고 있어 텍스트의 핵심 행동을 완전히 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S65sh10_sel.png",
    "asset_id": "ef41b208-c789-4220-811d-8cb034e97e93",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:907936>",
    "asset_id": "99c75fca-3287-4d9d-9682-424368fc63be",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1175078>",
    "asset_id": "62c2787c-8e39-42b7-a857-baae4c7b8abe",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-5767-70e2-875f-4414bae44f9f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S65sh10"
  },
  "lane_policy": "ab_select_bypass:prev",
  "staged_characters_added": [
   "C01",
   "C28"
  ]
 },
 "S65sh20::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:08:40.285001+00:00",
  "fingerprint": "60c8d058a0812f7cf2ddb00be221b55dcddcd7f730c17c4174fdf4ceca56a166",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S65sh20_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S65sh20_sel.png",
  "source_sha256": "8b787c1991d9aa8d52f28f052eb3e09380d7f052d5db55e3ac972a5dfa9b5cf7",
  "file": "S65sh20_cine.png",
  "staged_sha256": "5055e987a1737a992e0d367c8069e5c87bef1e4aace113982bf961f2475e7e66",
  "latency_ms": 11745
 },
 "S66sh14::signage": {
  "fp": "dea466e63a23a3f9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S66sh14": {
  "input_fingerprint": "5ddf22fc3b9b7515",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 어둠 속 손전등 불빛 아래로 드러난 먼지 쌓인 진열대와 수많은 통조림 캔들이 널린 지하 식품 코너 전경.\n\nLOCATION (lock): Inside the collapsed department store's underground food section, where flashlights reveal dusty shelves and scattered canned goods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Food-section displays (Dust-covered and still retaining their recognizable layout) — Shelf fronts and oblique ends remain visible along the lateral viewing direction; used as Establish depth and the unexpectedly extensive supply of food; Numerous canned foods (Accumulated throughout the dusty food section); used as Provide repeated small-scale discoveries across the composition without isolating a single oversized object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The flashlight reveals localized patches of dusty stock within the dark department store, with controlled contrast preserving the edges of the surrounding displays.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A newly collapsed floor leaves a high opening above the underground food department. Flashlight beams reveal dusty wine, canned food, snacks and surviving food displays in the dark interior.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 어둠 속 손전등 불빛 아래로 드러난 먼지 쌓인 진열대와 수많은 통조림 캔들이 널린 지하 식품 코너 전경.\n\nLOCATION (lock): Inside the collapsed department store's underground food section, where flashlights reveal dusty shelves and scattered canned goods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Food-section displays (Dust-covered and still retaining their recognizable layout) — Shelf fronts and oblique ends remain visible along the lateral viewing direction; used as Establish depth and the unexpectedly extensive supply of food; Numerous canned foods (Accumulated throughout the dusty food section); used as Provide repeated small-scale discoveries across the composition without isolating a single oversized object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The flashlight reveals localized patches of dusty stock within the dark department store, with controlled contrast preserving the edges of the surrounding displays.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A newly collapsed floor leaves a high opening above the underground food department. Flashlight beams reveal dusty wine, canned food, snacks and surviving food displays in the dark interior.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 어둠 속 손전등 불빛 아래로 드러난 먼지 쌓인 진열대와 수많은 통조림 캔들이 널린 지하 식품 코너 전경.\n\nLOCATION (lock): Inside the collapsed department store's underground food section, where flashlights reveal dusty shelves and scattered canned goods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Food-section displays (Dust-covered and still retaining their recognizable layout) — Shelf fronts and oblique ends remain visible along the lateral viewing direction; used as Establish depth and the unexpectedly extensive supply of food; Numerous canned foods (Accumulated throughout the dusty food section); used as Provide repeated small-scale discoveries across the composition without isolating a single oversized object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The flashlight reveals localized patches of dusty stock within the dark department store, with controlled contrast preserving the edges of the surrounding displays.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A newly collapsed floor leaves a high opening above the underground food department. Flashlight beams reveal dusty wine, canned food, snacks and surviving food displays in the dark interior.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "우측 하단의 손전등 불빛은 진열대를, 좌측 상단 허공의 손전등은 중앙을 향함.",
    "built_space": "레퍼런스 사진의 구조(에스컬레이터, 기둥)가 전혀 반영되지 않은 임의의 마트 통로임.",
    "entities": "통조림 캔, 선반, 손전등을 쥔 손.",
    "hard_violations": [
     "[gemini-pro] 지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
    ],
    "physics": "우측 손전등은 손에 쥐어져 있으나, 좌측 상단의 손전등은 아무 지지대 없이 허공에 떠 있음."
   },
   {
    "label": "B",
    "direction": "화면 밖 시점에서 비추는 손전등 불빛이 중앙의 곡선형 진열대에 정확히 닿음.",
    "built_space": "에스컬레이터와 기둥의 배치 등 레퍼런스 사진의 공간 구조를 완벽하게 일치시킴.",
    "entities": "통조림 캔, 진열대, 잔해. 기둥에 '가공식품', '신선식품' 등 읽을 수 있는 글자가 남아있음.",
    "hard_violations": [
     "[gpt-high] 오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
    ],
    "physics": "모든 캔과 파편들이 바닥과 선반 위에 안정적으로 위치해 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "구도를 복사하지 말라는 지시와 텍스트를 읽을 수 없게 처리하라는 지시를 어겼으나, 지정된 장소의 구조와 조명 분위기를 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스 장소를 무시한 임의의 공간을 생성했으며, 허공에 떠 있는 손전등이 있어 물리적 오류(Hard Violation)로 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "우측 하단의 손전등 불빛은 진열대를, 좌측 상단 허공의 손전등은 중앙을 향함.",
        "built_space": "레퍼런스 사진의 구조(에스컬레이터, 기둥)가 전혀 반영되지 않은 임의의 마트 통로임.",
        "entities": "통조림 캔, 선반, 손전등을 쥔 손.",
        "hard_violations": [
         "지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
        ],
        "physics": "우측 손전등은 손에 쥐어져 있으나, 좌측 상단의 손전등은 아무 지지대 없이 허공에 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 밖 시점에서 비추는 손전등 불빛이 중앙의 곡선형 진열대에 정확히 닿음.",
        "built_space": "에스컬레이터와 기둥의 배치 등 레퍼런스 사진의 공간 구조를 완벽하게 일치시킴.",
        "entities": "통조림 캔, 진열대, 잔해. 기둥에 '가공식품', '신선식품' 등 읽을 수 있는 글자가 남아있음.",
        "hard_violations": [],
        "physics": "모든 캔과 파편들이 바닥과 선반 위에 안정적으로 위치해 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "구도를 복사하지 말라는 지시와 텍스트를 읽을 수 없게 처리하라는 지시를 어겼으나, 지정된 장소의 구조와 조명 분위기를 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스 장소를 무시한 임의의 공간을 생성했으며, 허공에 떠 있는 손전등이 있어 물리적 오류(Hard Violation)로 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "우측 하단의 손전등 불빛은 진열대를, 좌측 상단 허공의 손전등은 중앙을 향함.",
        "built_space": "레퍼런스 사진의 구조(에스컬레이터, 기둥)가 전혀 반영되지 않은 임의의 마트 통로임.",
        "entities": "통조림 캔, 선반, 손전등을 쥔 손.",
        "hard_violations": [
         "지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
        ],
        "physics": "우측 손전등은 손에 쥐어져 있으나, 좌측 상단의 손전등은 아무 지지대 없이 허공에 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 밖 시점에서 비추는 손전등 불빛이 중앙의 곡선형 진열대에 정확히 닿음.",
        "built_space": "에스컬레이터와 기둥의 배치 등 레퍼런스 사진의 공간 구조를 완벽하게 일치시킴.",
        "entities": "통조림 캔, 진열대, 잔해. 기둥에 '가공식품', '신선식품' 등 읽을 수 있는 글자가 남아있음.",
        "hard_violations": [],
        "physics": "모든 캔과 파편들이 바닥과 선반 위에 안정적으로 위치해 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장소의 진열대 배치는 잘 보존했지만, 읽을 수 있는 안내 글자가 금지 조건을 위반하고 참조 사진의 위층 내려다보기 구도를 거의 답습합니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지하 내부의 주관적 와이드 시점과 통조림을 향하는 손전등 빛을 충실히 구현했으나, 높은 선반 중심의 배치는 참조 장소의 낮은 독립 진열대 구성과 차이가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 붕괴 가장자리에서 지하 매장을 비스듬히 내려다봅니다. 밝은 조명 영역은 앞쪽 통조림 진열대와 그 뒤 중앙 진열대에 닿지만, 광원이나 손전등의 조사 방향은 직접 보이지 않습니다. 사람이나 시선은 없습니다.",
        "built_space": "왼쪽 에스컬레이터 한 대, 앞쪽과 중앙의 독립 진열대 두 개, 뒤쪽 낮은 판매대 한 개, 오른쪽 가장자리에 일부 잘린 진열대 한 개와 후면 벽 선반들이 보입니다. 위쪽에는 부서진 슬래브와 늘어진 배선이 있고, 오른쪽에는 안내판이 붙은 기둥이 있습니다. 참조 장소의 주요 배치를 상당히 유지하지만 카메라 위치도 참조의 위층 가장자리 시점과 매우 유사합니다. 반사는 없습니다.",
        "entities": "먼지가 쌓인 진열대, 다수의 통조림, 바닥에 흩어진 캔과 콘크리트 잔해가 보입니다. 사람이나 얼굴은 없습니다. 손전등 자체는 보이지 않으며, 와인과 과자는 명확히 식별하기 어렵습니다. 오른쪽 기둥의 ‘신선식품’과 ‘FRESH FOOD’는 읽을 수 있어 무문자 조건에 어긋납니다.",
        "hard_violations": [
         "오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
        ],
        "physics": "진열대는 바닥에 놓여 있고 캔들은 선반이나 잔해가 있는 바닥에 지지됩니다. 늘어진 배선은 상부 구조에 연결되어 있습니다. 떠 있는 물체나 지지 없는 인체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 지하 매장 통로를 따라 수평에 가깝게 바라봅니다. 오른손에 쥔 손전등은 앞쪽 오른편의 쏟아진 통조림 더미와 통로 안쪽을 향하고, 왼쪽 위에서 들어오는 손전등 빛은 왼쪽 선반 상부를 비춥니다. 빛이 실제 식품 재고에 닿아 발견의 순간을 만듭니다.",
        "built_space": "왼쪽 가장자리 선반 한 줄, 왼쪽 중앙의 긴 독립 선반 한 줄, 오른쪽 전경 선반 한 줄, 오른쪽 중경 선반과 후면 선반이 보입니다. 오른쪽 중경에는 큰 콘크리트 기둥이 있고, 머리 위에는 낮빛이 들어오는 붕괴 개구부 한 곳이 있습니다. 선반 정면과 비스듬한 끝면이 함께 보여 깊이가 형성됩니다. 다만 참조의 낮은 독립 진열대 중심 배치보다 높은 통로형 선반이 지배적입니다. 반사는 없습니다.",
        "entities": "다수의 통조림, 봉지 과자류, 병 제품, 먼지 낀 금속 선반과 붕괴 잔해가 보입니다. 병의 내용물이 와인인지는 확정하기 어렵습니다. 손전등을 조작하는 오른손만 나타나며 얼굴이나 전신 인물은 없습니다. 왼쪽 위에는 두 번째 손전등의 앞부분이 걸쳐 있습니다. 확실히 읽을 수 있는 문구나 자막은 보이지 않습니다.",
        "hard_violations": [],
        "physics": "오른쪽 손전등은 손가락과 엄지가 몸통을 감싸 지지합니다. 왼쪽 위 손전등은 몸통이 화면 밖으로 이어져 파지 부위가 잘린 것으로 보이며 공중에 독립적으로 떠 있지는 않습니다. 캔 더미는 바닥과 서로 맞닿은 캔들에 지지되고, 기울어진 선반 판은 바닥과 주변 잔해에 걸쳐 있습니다. 매달린 철근과 배선은 부서진 천장에 연결되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "장소의 진열대 배치는 잘 보존했지만, 읽을 수 있는 안내 글자가 금지 조건을 위반하고 참조 사진의 위층 내려다보기 구도를 거의 답습합니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지하 내부의 주관적 와이드 시점과 통조림을 향하는 손전등 빛을 충실히 구현했으나, 높은 선반 중심의 배치는 참조 장소의 낮은 독립 진열대 구성과 차이가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 붕괴 가장자리에서 지하 매장을 비스듬히 내려다봅니다. 밝은 조명 영역은 앞쪽 통조림 진열대와 그 뒤 중앙 진열대에 닿지만, 광원이나 손전등의 조사 방향은 직접 보이지 않습니다. 사람이나 시선은 없습니다.",
        "built_space": "왼쪽 에스컬레이터 한 대, 앞쪽과 중앙의 독립 진열대 두 개, 뒤쪽 낮은 판매대 한 개, 오른쪽 가장자리에 일부 잘린 진열대 한 개와 후면 벽 선반들이 보입니다. 위쪽에는 부서진 슬래브와 늘어진 배선이 있고, 오른쪽에는 안내판이 붙은 기둥이 있습니다. 참조 장소의 주요 배치를 상당히 유지하지만 카메라 위치도 참조의 위층 가장자리 시점과 매우 유사합니다. 반사는 없습니다.",
        "entities": "먼지가 쌓인 진열대, 다수의 통조림, 바닥에 흩어진 캔과 콘크리트 잔해가 보입니다. 사람이나 얼굴은 없습니다. 손전등 자체는 보이지 않으며, 와인과 과자는 명확히 식별하기 어렵습니다. 오른쪽 기둥의 ‘신선식품’과 ‘FRESH FOOD’는 읽을 수 있어 무문자 조건에 어긋납니다.",
        "hard_violations": [
         "오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
        ],
        "physics": "진열대는 바닥에 놓여 있고 캔들은 선반이나 잔해가 있는 바닥에 지지됩니다. 늘어진 배선은 상부 구조에 연결되어 있습니다. 떠 있는 물체나 지지 없는 인체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "카메라는 지하 매장 통로를 따라 수평에 가깝게 바라봅니다. 오른손에 쥔 손전등은 앞쪽 오른편의 쏟아진 통조림 더미와 통로 안쪽을 향하고, 왼쪽 위에서 들어오는 손전등 빛은 왼쪽 선반 상부를 비춥니다. 빛이 실제 식품 재고에 닿아 발견의 순간을 만듭니다.",
        "built_space": "왼쪽 가장자리 선반 한 줄, 왼쪽 중앙의 긴 독립 선반 한 줄, 오른쪽 전경 선반 한 줄, 오른쪽 중경 선반과 후면 선반이 보입니다. 오른쪽 중경에는 큰 콘크리트 기둥이 있고, 머리 위에는 낮빛이 들어오는 붕괴 개구부 한 곳이 있습니다. 선반 정면과 비스듬한 끝면이 함께 보여 깊이가 형성됩니다. 다만 참조의 낮은 독립 진열대 중심 배치보다 높은 통로형 선반이 지배적입니다. 반사는 없습니다.",
        "entities": "다수의 통조림, 봉지 과자류, 병 제품, 먼지 낀 금속 선반과 붕괴 잔해가 보입니다. 병의 내용물이 와인인지는 확정하기 어렵습니다. 손전등을 조작하는 오른손만 나타나며 얼굴이나 전신 인물은 없습니다. 왼쪽 위에는 두 번째 손전등의 앞부분이 걸쳐 있습니다. 확실히 읽을 수 있는 문구나 자막은 보이지 않습니다.",
        "hard_violations": [],
        "physics": "오른쪽 손전등은 손가락과 엄지가 몸통을 감싸 지지합니다. 왼쪽 위 손전등은 몸통이 화면 밖으로 이어져 파지 부위가 잘린 것으로 보이며 공중에 독립적으로 떠 있지는 않습니다. 캔 더미는 바닥과 서로 맞닿은 캔들에 지지되고, 기울어진 선반 판은 바닥과 주변 잔해에 걸쳐 있습니다. 매달린 철근과 배선은 부서진 천장에 연결되어 있습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.125
   },
   "violations": {
    "A": [
     "[gemini-pro] 지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
    ],
    "B": [
     "[gpt-high] 오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "구도를 복사하지 말라는 지시와 텍스트를 읽을 수 없게 처리하라는 지시를 어겼으나, 지정된 장소의 구조와 조명 분위기를 충실히 구현했습니다.  ★위반: [gpt-high] 오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "레퍼런스 장소를 무시한 임의의 공간을 생성했으며, 허공에 떠 있는 손전등이 있어 물리적 오류(Hard Violation)로 실격입니다.  ★위반: [gemini-pro] 지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L240B01.png",
    "asset_id": "75e19812-b609-48b1-a66c-c19637b5c909",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-591e-743e-a44f-36f5ccb9a8c3",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S66sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:03:23.677285+00:00",
  "fingerprint": "5295fc4c9948dcf1348f2dcebff707529e3642c60d6729b2dd79a057918a7e55",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S66sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S66sh14_sel.png",
  "source_sha256": "cfd90d10256a5961bf57364a31a095ca60a7d66b589c4a7d4ee202aa6510d56a",
  "file": "S66sh14_cine.png",
  "staged_sha256": "a557ac0eb78611863af0fb71b9fa2b29ef6e67c765bbf717a26c188b6d288d93",
  "latency_ms": 9984
 },
 "S66sh26::signage": {
  "fp": "c85fb77f81b4efbf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9d63210811f5f1fe": {
  "subjects": [],
  "subject_text": "무너진 백화점 지상층과 붕괴 구멍 주변\n무너진 벽과 바닥 틈이 이어지는 어두운 상업시설 지상층. 먼지 쌓인 층별 안내판과 아래층으로 향하는 에스컬레이터가 보인다.",
  "identity": "canonical",
  "scope_id": "L240",
  "scope_role": "location_interior",
  "scope_sha": "cf1e843ecc3e18c5"
 },
 "S66sh26::bgfirst_bg": {
  "input_fingerprint": "a9ea5ff6bc1cbb2b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26__bgfirst_bg.png",
  "asset_id": "46131388-6cc7-4fff-87df-cd5907e8b98a",
  "input_asset_ids": [
   "a6150871-4cda-4118-b344-d3f6a8f9efc5",
   "75e19812-b609-48b1-a66c-c19637b5c909"
  ]
 },
 "S66sh26": {
  "input_fingerprint": "fa452bcd383a72bf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The collapse opening remains far above the food department, and the other exits are blocked. Dusty food stock remains in place, with canned food opened during the wait. 한수: He stands near the upper edge of the collapse opening, with short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The collapse opening remains far above the food department, and the other exits are blocked. Dusty food stock remains in place, with canned food opened during the wait. 한수: He stands near the upper edge of the collapse opening, with short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The collapse opening remains far above the food department, and the other exits are blocked. Dusty food stock remains in place, with canned food opened during the wait. 한수: He stands near the upper edge of the collapse opening, with short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26__bgfirst_bg.png",
     "asset_id": "46131388-6cc7-4fff-87df-cd5907e8b98a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S66sh26.png",
     "asset_id": "a6150871-4cda-4118-b344-d3f6a8f9efc5",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:907936>",
     "asset_id": "99c75fca-3287-4d9d-9682-424368fc63be",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L240B01.png",
     "asset_id": "75e19812-b609-48b1-a66c-c19637b5c909",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:907936>",
     "asset_id": "99c75fca-3287-4d9d-9682-424368fc63be",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "남자의 시선과 몸 방향이 아래층 붕괴 구멍을 향하고 있음.",
    "built_space": "폐허가 된 백화점 상층부 에스컬레이터와 붕괴된 바닥이 보이나, 레퍼런스 좌측 전경의 텍스트 기둥이 공간에서 누락됨.",
    "entities": "녹색 재킷, 배낭, 비니를 착용한 남성으로 레퍼런스의 한수 의상과 일치함.",
    "hard_violations": [],
    "physics": "두 발이 바닥에 안정적으로 지지된 채 서 있음."
   },
   {
    "label": "B",
    "direction": "남자의 시선과 몸 방향이 붕괴된 바닥 구멍을 향하고 있음.",
    "built_space": "레퍼런스의 기둥과 에스컬레이터 등 백화점 내부 구조가 동일하게 구현됨.",
    "entities": "짧은 머리와 녹색 재킷을 착용한 남성의 짙은 실루엣.",
    "hard_violations": [
     "[gemini-pro] invented objects (오른손에 쥔 손전등)",
     "[gemini-pro] leaked text (기둥 및 벽면의 선명하고 읽을 수 있는 한국어 텍스트)",
     "[gpt-high] 안내판과 벽면에 읽을 수 있는 글자가 있어 이미지 전체의 판독 가능한 문자 금지 조건을 위반한다.",
     "[gpt-high] 오른손에 지시와 참조에 없는 손전등 형태의 소품이 추가되어 있다."
    ],
    "physics": "바닥에 두 발로 서 있으며, 오른손은 손전등을 쥐어 지지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "프롬프트가 요구한 짙은 실루엣 연출이 다소 부족하고 일부 구조물(기둥)이 누락되었으나, 읽을 수 있는 텍스트나 발명된 사물 같은 치명적인 위반 사항이 없어 B보다 낫습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "실루엣과 배경 구조 묘사는 우수하나, 프롬프트에서 엄격히 금지한 '읽을 수 있는 텍스트'가 명확히 노출되었고 지시되지 않은 사물(손전등)을 들고 있어 치명적인 감점을 받았습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자의 시선과 몸 방향이 아래층 붕괴 구멍을 향하고 있음.",
        "built_space": "폐허가 된 백화점 상층부 에스컬레이터와 붕괴된 바닥이 보이나, 레퍼런스 좌측 전경의 텍스트 기둥이 공간에서 누락됨.",
        "entities": "녹색 재킷, 배낭, 비니를 착용한 남성으로 레퍼런스의 한수 의상과 일치함.",
        "hard_violations": [],
        "physics": "두 발이 바닥에 안정적으로 지지된 채 서 있음."
       },
       {
        "label": "B",
        "direction": "남자의 시선과 몸 방향이 붕괴된 바닥 구멍을 향하고 있음.",
        "built_space": "레퍼런스의 기둥과 에스컬레이터 등 백화점 내부 구조가 동일하게 구현됨.",
        "entities": "짧은 머리와 녹색 재킷을 착용한 남성의 짙은 실루엣.",
        "hard_violations": [
         "invented objects (오른손에 쥔 손전등)",
         "leaked text (기둥 및 벽면의 선명하고 읽을 수 있는 한국어 텍스트)"
        ],
        "physics": "바닥에 두 발로 서 있으며, 오른손은 손전등을 쥐어 지지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "프롬프트가 요구한 짙은 실루엣 연출이 다소 부족하고 일부 구조물(기둥)이 누락되었으나, 읽을 수 있는 텍스트나 발명된 사물 같은 치명적인 위반 사항이 없어 B보다 낫습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "실루엣과 배경 구조 묘사는 우수하나, 프롬프트에서 엄격히 금지한 '읽을 수 있는 텍스트'가 명확히 노출되었고 지시되지 않은 사물(손전등)을 들고 있어 치명적인 감점을 받았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "남자의 시선과 몸 방향이 아래층 붕괴 구멍을 향하고 있음.",
        "built_space": "폐허가 된 백화점 상층부 에스컬레이터와 붕괴된 바닥이 보이나, 레퍼런스 좌측 전경의 텍스트 기둥이 공간에서 누락됨.",
        "entities": "녹색 재킷, 배낭, 비니를 착용한 남성으로 레퍼런스의 한수 의상과 일치함.",
        "hard_violations": [],
        "physics": "두 발이 바닥에 안정적으로 지지된 채 서 있음."
       },
       {
        "label": "B",
        "direction": "남자의 시선과 몸 방향이 붕괴된 바닥 구멍을 향하고 있음.",
        "built_space": "레퍼런스의 기둥과 에스컬레이터 등 백화점 내부 구조가 동일하게 구현됨.",
        "entities": "짧은 머리와 녹색 재킷을 착용한 남성의 짙은 실루엣.",
        "hard_violations": [
         "invented objects (오른손에 쥔 손전등)",
         "leaked text (기둥 및 벽면의 선명하고 읽을 수 있는 한국어 텍스트)"
        ],
        "physics": "바닥에 두 발로 서 있으며, 오른손은 손전등을 쥐어 지지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "붕괴 구멍을 향한 어두운 뒷모습과 장소는 맞지만, 읽을 수 있는 안내문·벽면 문구와 지시되지 않은 손전등이 명시적 제약을 위반한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "상층 가장자리에서 아래를 살피는 한수의 어두운 후면을 와이드숏으로 담고, 참조 공간·복장과 아래층 인물 비노출 조건을 충실히 지킨다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자의 등과 머리는 카메라 반대편의 붕괴 구멍과 아래층 식품 매장을 향한다. 눈은 보이지 않아 정확한 주시점은 확인할 수 없다. 오른손의 손전등처럼 보이는 물체는 아래로 향하며 빛을 비추지는 않는다.",
        "built_space": "상층 바닥에서 남자 뒤로 구멍의 가까운 가장자리를 비스듬히 내려다본다. 큰 붕괴 개구부 하나, 왼쪽 에스컬레이터 한 대와 유리 난간, 왼쪽 창열, 맞은편 상층 통로, 노출 철근과 매달린 배선이 보인다. 아래층에는 전면 통조림 진열대, 중앙 오른쪽 진열대, 뒤쪽 낮은 판매대와 벽면 선반이 있다. 참조의 층간 구조와 잘 맞으며 남자는 남아 있는 상층 바닥에 위치한다.",
        "entities": "인물은 짧은 검은 머리의 성인 남성 한 명이며 얼굴과 출신은 뒷모습으로 확인할 수 없다. 어두운 녹색 외투와 갈색 계열 바지는 보이지만 참조의 비니·배낭·풍성한 후드 털은 보이지 않는다. 오른손에는 참조와 지시에 없는 손전등 형태의 물체가 있다. 아래층 식품 재고는 있으나 다른 사람은 보이지 않는다. 왼쪽 안내판의 ‘1F’, ‘화장품’ 등과 벽면 광고 문구가 읽힌다.",
        "hard_violations": [
         "안내판과 벽면에 읽을 수 있는 글자가 있어 이미지 전체의 판독 가능한 문자 금지 조건을 위반한다.",
         "오른손에 지시와 참조에 없는 손전등 형태의 소품이 추가되어 있다."
        ],
        "physics": "발은 화면 밖이지만 다리는 가장자리 안쪽의 온전한 바닥으로 이어지며 공중에 떠 있다는 징후는 없다. 오른손 물체는 손으로 잡고 있다. 진열대는 아래층 바닥에 놓이고, 늘어진 배선은 파손된 구조체에 연결되어 있다. 인체나 물체의 명백한 지지 불가능 상태는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "남자는 카메라에 등과 오른쪽 측면을 보인 채 고개를 숙여 붕괴 가장자리 너머의 아래층 식품 매장을 바라본다. 시선 방향이 숨은 관찰자와 아래층 공간을 연결하며, 아래층 사람은 드러내지 않는다. 손에는 조준하거나 사용하는 물체가 없다.",
        "built_space": "상층 바닥의 남자 뒤쪽에서 가까운 붕괴 가장자리를 비스듬히 보는 와이드숏이다. 붕괴 개구부 하나, 왼쪽 에스컬레이터 한 대와 양측 난간, 왼쪽 창열, 맞은편 상층 통로, 파손된 슬래브와 배선이 보인다. 아래층의 전면 통조림 진열대, 중앙 오른쪽 진열대, 뒤쪽 판매대와 벽면 선반 배치도 참조와 일치한다. 남자는 구멍 왼쪽의 남아 있는 바닥에 서 있다. 참조의 독립 안내판은 보이지 않지만 주요 건축 구조는 유지된다.",
        "entities": "성인 남성 한 명만 보이며, 부분적으로 보이는 옆얼굴은 동아시아계 외형이다. 얼굴 대부분과 머리카락은 가려져 정확한 얼굴 동일성이나 짧은 머리는 확인하기 어렵다. 비니, 털 장식 녹색 파카, 갈색 배낭, 카고 바지와 부츠는 인물 참조에 부합한다. 식품 재고와 통조림은 보이지만 개봉 여부는 판별되지 않는다. 아래층 인물과 명확히 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 부츠가 상층의 남아 있는 바닥에 닿아 몸을 지탱하고, 발 간격과 앞으로 숙인 머리는 가장자리에서 아래를 살피며 멈춘 자세로 가능하다. 배낭은 어깨끈으로 지지되고 손은 몸 옆에 자연스럽게 내려와 있다. 진열대와 잔해는 각 층 바닥에 놓이며 배선은 구조체에서 늘어진다. 지지 없는 인체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "붕괴 구멍을 향한 어두운 뒷모습과 장소는 맞지만, 읽을 수 있는 안내문·벽면 문구와 지시되지 않은 손전등이 명시적 제약을 위반한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "상층 가장자리에서 아래를 살피는 한수의 어두운 후면을 와이드숏으로 담고, 참조 공간·복장과 아래층 인물 비노출 조건을 충실히 지킨다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "남자의 등과 머리는 카메라 반대편의 붕괴 구멍과 아래층 식품 매장을 향한다. 눈은 보이지 않아 정확한 주시점은 확인할 수 없다. 오른손의 손전등처럼 보이는 물체는 아래로 향하며 빛을 비추지는 않는다.",
        "built_space": "상층 바닥에서 남자 뒤로 구멍의 가까운 가장자리를 비스듬히 내려다본다. 큰 붕괴 개구부 하나, 왼쪽 에스컬레이터 한 대와 유리 난간, 왼쪽 창열, 맞은편 상층 통로, 노출 철근과 매달린 배선이 보인다. 아래층에는 전면 통조림 진열대, 중앙 오른쪽 진열대, 뒤쪽 낮은 판매대와 벽면 선반이 있다. 참조의 층간 구조와 잘 맞으며 남자는 남아 있는 상층 바닥에 위치한다.",
        "entities": "인물은 짧은 검은 머리의 성인 남성 한 명이며 얼굴과 출신은 뒷모습으로 확인할 수 없다. 어두운 녹색 외투와 갈색 계열 바지는 보이지만 참조의 비니·배낭·풍성한 후드 털은 보이지 않는다. 오른손에는 참조와 지시에 없는 손전등 형태의 물체가 있다. 아래층 식품 재고는 있으나 다른 사람은 보이지 않는다. 왼쪽 안내판의 ‘1F’, ‘화장품’ 등과 벽면 광고 문구가 읽힌다.",
        "hard_violations": [
         "안내판과 벽면에 읽을 수 있는 글자가 있어 이미지 전체의 판독 가능한 문자 금지 조건을 위반한다.",
         "오른손에 지시와 참조에 없는 손전등 형태의 소품이 추가되어 있다."
        ],
        "physics": "발은 화면 밖이지만 다리는 가장자리 안쪽의 온전한 바닥으로 이어지며 공중에 떠 있다는 징후는 없다. 오른손 물체는 손으로 잡고 있다. 진열대는 아래층 바닥에 놓이고, 늘어진 배선은 파손된 구조체에 연결되어 있다. 인체나 물체의 명백한 지지 불가능 상태는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "남자는 카메라에 등과 오른쪽 측면을 보인 채 고개를 숙여 붕괴 가장자리 너머의 아래층 식품 매장을 바라본다. 시선 방향이 숨은 관찰자와 아래층 공간을 연결하며, 아래층 사람은 드러내지 않는다. 손에는 조준하거나 사용하는 물체가 없다.",
        "built_space": "상층 바닥의 남자 뒤쪽에서 가까운 붕괴 가장자리를 비스듬히 보는 와이드숏이다. 붕괴 개구부 하나, 왼쪽 에스컬레이터 한 대와 양측 난간, 왼쪽 창열, 맞은편 상층 통로, 파손된 슬래브와 배선이 보인다. 아래층의 전면 통조림 진열대, 중앙 오른쪽 진열대, 뒤쪽 판매대와 벽면 선반 배치도 참조와 일치한다. 남자는 구멍 왼쪽의 남아 있는 바닥에 서 있다. 참조의 독립 안내판은 보이지 않지만 주요 건축 구조는 유지된다.",
        "entities": "성인 남성 한 명만 보이며, 부분적으로 보이는 옆얼굴은 동아시아계 외형이다. 얼굴 대부분과 머리카락은 가려져 정확한 얼굴 동일성이나 짧은 머리는 확인하기 어렵다. 비니, 털 장식 녹색 파카, 갈색 배낭, 카고 바지와 부츠는 인물 참조에 부합한다. 식품 재고와 통조림은 보이지만 개봉 여부는 판별되지 않는다. 아래층 인물과 명확히 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 부츠가 상층의 남아 있는 바닥에 닿아 몸을 지탱하고, 발 간격과 앞으로 숙인 머리는 가장자리에서 아래를 살피며 멈춘 자세로 가능하다. 배낭은 어깨끈으로 지지되고 손은 몸 옆에 자연스럽게 내려와 있다. 진열대와 잔해는 각 층 바닥에 놓이며 배선은 구조체에서 늘어진다. 지지 없는 인체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] invented objects (오른손에 쥔 손전등)",
     "[gemini-pro] leaked text (기둥 및 벽면의 선명하고 읽을 수 있는 한국어 텍스트)",
     "[gpt-high] 안내판과 벽면에 읽을 수 있는 글자가 있어 이미지 전체의 판독 가능한 문자 금지 조건을 위반한다.",
     "[gpt-high] 오른손에 지시와 참조에 없는 손전등 형태의 소품이 추가되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 짙은 실루엣 연출이 다소 부족하고 일부 구조물(기둥)이 누락되었으나, 읽을 수 있는 텍스트나 발명된 사물 같은 치명적인 위반 사항이 없어 B보다 낫습니다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "실루엣과 배경 구조 묘사는 우수하나, 프롬프트에서 엄격히 금지한 '읽을 수 있는 텍스트'가 명확히 노출되었고 지시되지 않은 사물(손전등)을 들고 있어 치명적인 감점을 받았습니다.  ★위반: [gemini-pro] invented objects (오른손에 쥔 손전등) / [gemini-pro] leaked text (기둥 및 벽면의 선명하고 읽을 수 있는 한국어 텍스트) / [gpt-high] 안내판과 벽면에 읽을 수 있는 글자가 있어 이미지 전체의 판독 가능한 문자 금지 조건을 위반한다. / [gpt-high] 오른손에 지시와 참조에 없는 손전등 형태의 소품이 추가되어 있다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L240B01.png",
    "asset_id": "75e19812-b609-48b1-a66c-c19637b5c909",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:907936>",
    "asset_id": "99c75fca-3287-4d9d-9682-424368fc63be",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-5ac8-796d-8baa-8eee890374fe",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26__bgfirst_bg.png",
   "bg_asset_id": "46131388-6cc7-4fff-87df-cd5907e8b98a",
   "bg_record_key": "S66sh26::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S66sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:10:05.861472+00:00",
  "fingerprint": "506e4402f2b59830cf83dc18df5363ee676e0d430e31a04149e420c3113fda7f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S66sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S66sh26_sel.png",
  "source_sha256": "ba6732928ffe3f7ca25b9358642be02756a06f16ea66d0d7e043e9c26ea0b471",
  "file": "S66sh26_cine.png",
  "staged_sha256": "987910e9523518b016e095a4e630e45a64aa81c7271749375d4f96a70ec85e08",
  "latency_ms": 10668
 },
 "S66sh36::signage": {
  "fp": "58aa5cc5a108d959",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S66sh36": {
  "input_fingerprint": "23377bd6e11c92e5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 날카로운 눈빛에 짧은 머리를 한 한수가 밧줄을 쥔 채 내려다보는 굳은 전신.\n\nLOCATION (lock): At the upper edge of the department store's collapse shaft, inside the dim ruined sales floor where the rescue rope is anchored. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Rescue rope (Held by 한수 and extending into the opening) — Runs obliquely from his hand toward the lower-left edge; used as Link the revealed rescuer to the unseen climber without showing the POV character; Upper-floor opening rim (Collapsed and open) — A narrow portion of the upper edge is visible at the bottom of the low viewpoint; used as Ground the emerging subjective perspective and establish separation from 한수.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the dim interior ambience while allowing enough facial separation to reveal 한수's sharp expression without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the broken concrete floor, rubble, and edges of the collapse opening. Exclude the basement food shelves and canned goods as furnishings of this upper-level location.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rescue rope hangs through the collapse opening into the food department. The opened parasol remains on the lower level, where rain has dripped through from above. 한수: He stands at the upper rescue point, with short hair and a sharp gaze.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 한수 right now, so 한수's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 한수: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 날카로운 눈빛에 짧은 머리를 한 한수가 밧줄을 쥔 채 내려다보는 굳은 전신.\n\nLOCATION (lock): At the upper edge of the department store's collapse shaft, inside the dim ruined sales floor where the rescue rope is anchored. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Rescue rope (Held by 한수 and extending into the opening) — Runs obliquely from his hand toward the lower-left edge; used as Link the revealed rescuer to the unseen climber without showing the POV character; Upper-floor opening rim (Collapsed and open) — A narrow portion of the upper edge is visible at the bottom of the low viewpoint; used as Ground the emerging subjective perspective and establish separation from 한수.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the dim interior ambience while allowing enough facial separation to reveal 한수's sharp expression without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the broken concrete floor, rubble, and edges of the collapse opening. Exclude the basement food shelves and canned goods as furnishings of this upper-level location.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rescue rope hangs through the collapse opening into the food department. The opened parasol remains on the lower level, where rain has dripped through from above. 한수: He stands at the upper rescue point, with short hair and a sharp gaze.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 한수 right now, so 한수's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 한수: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 날카로운 눈빛에 짧은 머리를 한 한수가 밧줄을 쥔 채 내려다보는 굳은 전신.\n\nLOCATION (lock): At the upper edge of the department store's collapse shaft, inside the dim ruined sales floor where the rescue rope is anchored. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Rescue rope (Held by 한수 and extending into the opening) — Runs obliquely from his hand toward the lower-left edge; used as Link the revealed rescuer to the unseen climber without showing the POV character; Upper-floor opening rim (Collapsed and open) — A narrow portion of the upper edge is visible at the bottom of the low viewpoint; used as Ground the emerging subjective perspective and establish separation from 한수.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the dim interior ambience while allowing enough facial separation to reveal 한수's sharp expression without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the broken concrete floor, rubble, and edges of the collapse opening. Exclude the basement food shelves and canned goods as furnishings of this upper-level location.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rescue rope hangs through the collapse opening into the food department. The opened parasol remains on the lower level, where rain has dripped through from above. 한수: He stands at the upper rescue point, with short hair and a sharp gaze.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 한수 right now, so 한수's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 한수: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "한수는 카메라를 향해 시선을 아래로 내리깔고 있으며, 쥐고 있는 밧줄은 화면 좌측 하단으로 향함.",
    "built_space": "아래에서 위를 올려다보는 로우 앵글 뷰. 화면 하단에 무너진 바닥 가장자리가 있고, 인물 뒤로는 위층의 천장과 빈 공간이 보임.",
    "entities": "짧은 검은 머리의 한수(녹색 파카, 회색 후드, 갈색 바지)와 구조용 밧줄이 프롬프트 및 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "콘크리트 가장자리에 안정적으로 서서 한 손으로 밧줄을 쥐고 있으며, 밧줄의 장력과 중력 방향이 자연스러움."
   },
   {
    "label": "B",
    "direction": "한수는 정면을 향해 약간 아래로 시선을 두고 있으며, 밧줄은 좌측 하단으로 비스듬히 이어짐.",
    "built_space": "카메라가 붕괴된 구멍 위 허공에 떠 있는 구도. 인물 뒤로 에스컬레이터와 프롬프트에서 배제하라고 명시된 지하층 진열대가 보임.",
    "entities": "짧은 검은 머리의 한수(녹색 파카, 백팩 착용)와 구조용 밧줄이 존재함.",
    "hard_violations": [
     "[gemini-pro] 명시된 하단 시점(로우 앵글)을 무시하고 허공에서 건너편을 바라보는 불가능한 카메라 위치",
     "[gemini-pro] 명시적으로 배제하도록 지시된 하층부의 식품 진열대가 배경에 렌더링됨",
     "[gpt-high] 상층 구조 지점에 서 있어야 할 한수가 붕괴 구멍 안쪽에 배치되어 있으며, 그 위치에서 몸을 받치는 바닥이나 다른 지지물이 보이지 않는다."
    ],
    "physics": "바닥 가장자리에 서서 두 손으로 밧줄을 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 붕괴된 구멍 아래에서의 로우 앵글 시점과 인물이 카메라를 내려다보는 구도를 정확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 하단 시점을 무시하고 허공에서 바라보는 구도를 취했으며, 배제해야 할 지하층 배경이 그대로 나타나는 치명적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "한수는 카메라를 향해 시선을 아래로 내리깔고 있으며, 쥐고 있는 밧줄은 화면 좌측 하단으로 향함.",
        "built_space": "아래에서 위를 올려다보는 로우 앵글 뷰. 화면 하단에 무너진 바닥 가장자리가 있고, 인물 뒤로는 위층의 천장과 빈 공간이 보임.",
        "entities": "짧은 검은 머리의 한수(녹색 파카, 회색 후드, 갈색 바지)와 구조용 밧줄이 프롬프트 및 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "콘크리트 가장자리에 안정적으로 서서 한 손으로 밧줄을 쥐고 있으며, 밧줄의 장력과 중력 방향이 자연스러움."
       },
       {
        "label": "B",
        "direction": "한수는 정면을 향해 약간 아래로 시선을 두고 있으며, 밧줄은 좌측 하단으로 비스듬히 이어짐.",
        "built_space": "카메라가 붕괴된 구멍 위 허공에 떠 있는 구도. 인물 뒤로 에스컬레이터와 프롬프트에서 배제하라고 명시된 지하층 진열대가 보임.",
        "entities": "짧은 검은 머리의 한수(녹색 파카, 백팩 착용)와 구조용 밧줄이 존재함.",
        "hard_violations": [
         "명시된 하단 시점(로우 앵글)을 무시하고 허공에서 건너편을 바라보는 불가능한 카메라 위치",
         "명시적으로 배제하도록 지시된 하층부의 식품 진열대가 배경에 렌더링됨"
        ],
        "physics": "바닥 가장자리에 서서 두 손으로 밧줄을 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 붕괴된 구멍 아래에서의 로우 앵글 시점과 인물이 카메라를 내려다보는 구도를 정확하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 하단 시점을 무시하고 허공에서 바라보는 구도를 취했으며, 배제해야 할 지하층 배경이 그대로 나타나는 치명적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "한수는 카메라를 향해 시선을 아래로 내리깔고 있으며, 쥐고 있는 밧줄은 화면 좌측 하단으로 향함.",
        "built_space": "아래에서 위를 올려다보는 로우 앵글 뷰. 화면 하단에 무너진 바닥 가장자리가 있고, 인물 뒤로는 위층의 천장과 빈 공간이 보임.",
        "entities": "짧은 검은 머리의 한수(녹색 파카, 회색 후드, 갈색 바지)와 구조용 밧줄이 프롬프트 및 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "콘크리트 가장자리에 안정적으로 서서 한 손으로 밧줄을 쥐고 있으며, 밧줄의 장력과 중력 방향이 자연스러움."
       },
       {
        "label": "B",
        "direction": "한수는 정면을 향해 약간 아래로 시선을 두고 있으며, 밧줄은 좌측 하단으로 비스듬히 이어짐.",
        "built_space": "카메라가 붕괴된 구멍 위 허공에 떠 있는 구도. 인물 뒤로 에스컬레이터와 프롬프트에서 배제하라고 명시된 지하층 진열대가 보임.",
        "entities": "짧은 검은 머리의 한수(녹색 파카, 백팩 착용)와 구조용 밧줄이 존재함.",
        "hard_violations": [
         "명시된 하단 시점(로우 앵글)을 무시하고 허공에서 건너편을 바라보는 불가능한 카메라 위치",
         "명시적으로 배제하도록 지시된 하층부의 식품 진열대가 배경에 렌더링됨"
        ],
        "physics": "바닥 가장자리에 서서 두 손으로 밧줄을 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "아래에서 올려다보는 현우의 시점과 한수의 하향 시선, 왼쪽 아래로 이어지는 밧줄은 정확하지만, 하체가 가려져 요구한 굳은 전신 와이드 숏에는 미달한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "밧줄 방향과 폐허의 재질은 맞지만 한수가 상층 구조 지점이 아니라 구멍 안에 지지 없이 놓였고, 카메라도 아래에서 그를 올려다보는 시점이 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "한수는 카메라 쪽을 거의 수평으로 응시한다. 아래에서 올라오는 현우를 내려다보는 시선은 약하다. 손에 잡힌 밧줄은 화면 왼쪽 아래로 뻗어 지정된 화면상 방향은 맞는다.",
        "built_space": "왼쪽에 에스컬레이터 한 대, 중앙에 큰 붕괴 개구부 하나, 둘레의 파손된 슬래브와 기둥들, 뒤쪽 창과 난간이 보인다. 아래층에는 상품 진열대 여러 개가 남아 있다. 참고 장소의 특징은 상당 부분 유지했지만, 한수는 상층 바닥 위가 아니라 개구부 안쪽에 내려앉은 위치로 보인다. 카메라는 상층 바닥 너머로 아래층을 내려다보며, 하단의 바닥도 좁은 가장자리보다 넓게 차지한다.",
        "entities": "성인 동아시아계 남성 한 명이며 다른 사람은 없다. 짧은 검은 머리, 녹색 털장식 파카, 갈색 계열 바지와 배낭은 한수의 주요 외형과 부합한다. 참고의 니트모자는 없지만 짧은 머리를 드러내라는 본문에는 맞는다. 얼굴은 참고와 다소 다르다. 구조용 밧줄과 콘크리트 잔해가 보이며, 하체 끝과 신발은 가려져 전신은 아니다. 아래층 양산은 보이지 않으나 이를 확인할 구도는 아니다.",
        "hard_violations": [
         "상층 구조 지점에 서 있어야 할 한수가 붕괴 구멍 안쪽에 배치되어 있으며, 그 위치에서 몸을 받치는 바닥이나 다른 지지물이 보이지 않는다."
        ],
        "physics": "손은 밧줄을 실제로 움켜쥐고 있고 밧줄은 왼쪽 전경으로 팽팽하게 이어진다. 그러나 몸 아래는 아래층까지 열린 구멍이며, 한수의 발이 놓일 상층 바닥이나 발판은 보이지 않는다. 전경 가장자리 뒤로 다리가 사라지는데 그 높이와 위치에서 몸을 지탱하는 것이 없다. 밧줄을 잡은 자세 역시 매달린 몸을 지지하는 자세는 아니다."
       },
       {
        "label": "B",
        "direction": "한수의 고개와 눈은 아래쪽 카메라, 즉 구멍에서 올라오는 현우의 위치를 향한다. 손의 밧줄은 왼쪽 아래 화면 밖으로 이어져 보이지 않는 등반자와의 연결을 명확하게 만든다.",
        "built_space": "하단에 파손된 상층 슬래브 가장자리 하나가 가로놓이고, 그 뒤에 한수가 서 있다. 왼쪽 전경 기둥 하나와 뒤쪽 기둥들, 노출된 보와 손상된 천장이 보인다. 낮은 카메라에서 상층 가장자리를 올려다보는 구조는 맞고, 상층에 식품 진열대나 통조림을 옮겨 놓지 않았다. 다만 가장자리가 하단을 상당히 두껍게 차지하고 참고 장소 고유의 창과 에스컬레이터는 이 구도에서 확인되지 않는다.",
        "entities": "성인 동아시아계 남성 한 명으로, 짧은 검은 머리와 얼굴이 한수 참고에 비교적 가깝다. 녹색 털장식 파카와 갈색 계열 바지도 부합한다. 니트모자는 없지만 본문의 짧은 머리 지시를 따른다. 맨손으로 구조용 밧줄을 잡고 있으며 다른 인물이나 읽을 수 있는 글자는 없다. 배낭과 신발, 아래층 양산은 이 시점에서 확인되지 않는다. 무릎 아래가 가장자리에 가려져 요구한 전신은 보이지 않는다.",
        "hard_violations": [],
        "physics": "한수는 붕괴 가장자리 뒤의 남아 있는 상층 바닥에 서 있는 배치다. 발의 접촉점은 슬래브에 가려졌지만 몸과 다리의 위치는 그 바닥의 지지와 양립한다. 손이 밧줄을 감싸 쥐고, 당겨진 구간은 왼쪽 아래로 이어지며 남는 구간은 아래로 늘어진다. 공중에 뜬 인물이나 손 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "아래에서 올려다보는 현우의 시점과 한수의 하향 시선, 왼쪽 아래로 이어지는 밧줄은 정확하지만, 하체가 가려져 요구한 굳은 전신 와이드 숏에는 미달한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "밧줄 방향과 폐허의 재질은 맞지만 한수가 상층 구조 지점이 아니라 구멍 안에 지지 없이 놓였고, 카메라도 아래에서 그를 올려다보는 시점이 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "한수는 카메라 쪽을 거의 수평으로 응시한다. 아래에서 올라오는 현우를 내려다보는 시선은 약하다. 손에 잡힌 밧줄은 화면 왼쪽 아래로 뻗어 지정된 화면상 방향은 맞는다.",
        "built_space": "왼쪽에 에스컬레이터 한 대, 중앙에 큰 붕괴 개구부 하나, 둘레의 파손된 슬래브와 기둥들, 뒤쪽 창과 난간이 보인다. 아래층에는 상품 진열대 여러 개가 남아 있다. 참고 장소의 특징은 상당 부분 유지했지만, 한수는 상층 바닥 위가 아니라 개구부 안쪽에 내려앉은 위치로 보인다. 카메라는 상층 바닥 너머로 아래층을 내려다보며, 하단의 바닥도 좁은 가장자리보다 넓게 차지한다.",
        "entities": "성인 동아시아계 남성 한 명이며 다른 사람은 없다. 짧은 검은 머리, 녹색 털장식 파카, 갈색 계열 바지와 배낭은 한수의 주요 외형과 부합한다. 참고의 니트모자는 없지만 짧은 머리를 드러내라는 본문에는 맞는다. 얼굴은 참고와 다소 다르다. 구조용 밧줄과 콘크리트 잔해가 보이며, 하체 끝과 신발은 가려져 전신은 아니다. 아래층 양산은 보이지 않으나 이를 확인할 구도는 아니다.",
        "hard_violations": [
         "상층 구조 지점에 서 있어야 할 한수가 붕괴 구멍 안쪽에 배치되어 있으며, 그 위치에서 몸을 받치는 바닥이나 다른 지지물이 보이지 않는다."
        ],
        "physics": "손은 밧줄을 실제로 움켜쥐고 있고 밧줄은 왼쪽 전경으로 팽팽하게 이어진다. 그러나 몸 아래는 아래층까지 열린 구멍이며, 한수의 발이 놓일 상층 바닥이나 발판은 보이지 않는다. 전경 가장자리 뒤로 다리가 사라지는데 그 높이와 위치에서 몸을 지탱하는 것이 없다. 밧줄을 잡은 자세 역시 매달린 몸을 지지하는 자세는 아니다."
       },
       {
        "label": "A",
        "direction": "한수의 고개와 눈은 아래쪽 카메라, 즉 구멍에서 올라오는 현우의 위치를 향한다. 손의 밧줄은 왼쪽 아래 화면 밖으로 이어져 보이지 않는 등반자와의 연결을 명확하게 만든다.",
        "built_space": "하단에 파손된 상층 슬래브 가장자리 하나가 가로놓이고, 그 뒤에 한수가 서 있다. 왼쪽 전경 기둥 하나와 뒤쪽 기둥들, 노출된 보와 손상된 천장이 보인다. 낮은 카메라에서 상층 가장자리를 올려다보는 구조는 맞고, 상층에 식품 진열대나 통조림을 옮겨 놓지 않았다. 다만 가장자리가 하단을 상당히 두껍게 차지하고 참고 장소 고유의 창과 에스컬레이터는 이 구도에서 확인되지 않는다.",
        "entities": "성인 동아시아계 남성 한 명으로, 짧은 검은 머리와 얼굴이 한수 참고에 비교적 가깝다. 녹색 털장식 파카와 갈색 계열 바지도 부합한다. 니트모자는 없지만 본문의 짧은 머리 지시를 따른다. 맨손으로 구조용 밧줄을 잡고 있으며 다른 인물이나 읽을 수 있는 글자는 없다. 배낭과 신발, 아래층 양산은 이 시점에서 확인되지 않는다. 무릎 아래가 가장자리에 가려져 요구한 전신은 보이지 않는다.",
        "hard_violations": [],
        "physics": "한수는 붕괴 가장자리 뒤의 남아 있는 상층 바닥에 서 있는 배치다. 발의 접촉점은 슬래브에 가려졌지만 몸과 다리의 위치는 그 바닥의 지지와 양립한다. 손이 밧줄을 감싸 쥐고, 당겨진 구간은 왼쪽 아래로 이어지며 남는 구간은 아래로 늘어진다. 공중에 뜬 인물이나 손 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.714
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.464
   },
   "violations": {
    "B": [
     "[gemini-pro] 명시된 하단 시점(로우 앵글)을 무시하고 허공에서 건너편을 바라보는 불가능한 카메라 위치",
     "[gemini-pro] 명시적으로 배제하도록 지시된 하층부의 식품 진열대가 배경에 렌더링됨",
     "[gpt-high] 상층 구조 지점에 서 있어야 할 한수가 붕괴 구멍 안쪽에 배치되어 있으며, 그 위치에서 몸을 받치는 바닥이나 다른 지지물이 보이지 않는다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 464
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 붕괴된 구멍 아래에서의 로우 앵글 시점과 인물이 카메라를 내려다보는 구도를 정확하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 464,
    "verdict_ko": "지정된 하단 시점을 무시하고 허공에서 바라보는 구도를 취했으며, 배제해야 할 지하층 배경이 그대로 나타나는 치명적 오류가 있습니다.  ★위반: [gemini-pro] 명시된 하단 시점(로우 앵글)을 무시하고 허공에서 건너편을 바라보는 불가능한 카메라 위치 / [gemini-pro] 명시적으로 배제하도록 지시된 하층부의 식품 진열대가 배경에 렌더링됨 / [gpt-high] 상층 구조 지점에 서 있어야 할 한수가 붕괴 구멍 안쪽에 배치되어 있으며, 그 위치에서 몸을 받치는 바닥이나 다른 지지물이 보이지 않는다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26_sel.png",
    "asset_id": "90b51f3a-ef00-4740-a517-df8b5061dc7f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:907936>",
    "asset_id": "99c75fca-3287-4d9d-9682-424368fc63be",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-5e0a-70c2-986e-bdd704f411eb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S66sh26"
  }
 },
 "S66sh36::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:11:08.939005+00:00",
  "fingerprint": "c3f73a245b2fda9b7e5c2bb51c9d9038eb0c7c159799254abbf37896a3f405df",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S66sh36_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S66sh36_sel.png",
  "source_sha256": "c365f718d575b3b36ebc52154a809ba0a3651f05bfcf06ea896c2d4544c6a260",
  "file": "S66sh36_cine.png",
  "staged_sha256": "693b29e1eab717f9bd85f54e8214adedaf9973964c67dbb0fa6902f5d11fa919",
  "latency_ms": 10060
 },
 "S67sh40::signage": {
  "fp": "01d87f1ac80ec18e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::village_lodging_yard": {
  "input_fingerprint": "36b68fc8253e9dfe",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "village_lodging_yard",
    "tags": [
     "S67sh40"
    ]
   },
   "context_sig": "de29d4e7b39271bd"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 찰리와 B-200. 숙소를 내려다보는데... 여기저기 터지는 군용 헤드라이트 불빛!\n- 이때 어디선가 끌려 나오는 현우와 앰버.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 찰리와 B-200. 숙소를 내려다보는데... 여기저기 터지는 군용 헤드라이트 불빛!\n- 이때 어디선가 끌려 나오는 현우와 앰버.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_lodging_yard_928b09.png",
  "asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be",
  "input_asset_ids": [
   "a32fd8a9-4876-47cb-ad74-e319feef74f4"
  ],
  "origin_tag": "S67sh40",
  "place_text": "In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.",
  "origin_inputs": {
   "place_text": "In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.",
   "time_of_day_en": "night, bright moonlight",
   "conti_asset_id": "a32fd8a9-4876-47cb-ad74-e319feef74f4"
  }
 },
 "S67sh40::bgfirst_bg": {
  "input_fingerprint": "94e73786571ee72e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh40__bgfirst_bg.png",
  "asset_id": "6d5cb97b-e1e7-40f2-8bdc-880bfadc45a7",
  "input_asset_ids": [
   "a32fd8a9-4876-47cb-ad74-e319feef74f4",
   "69fb2d3b-78b7-46d7-89c8-296f2c05b2be"
  ]
 },
 "S67sh40": {
  "input_fingerprint": "46fb49a2523c4f1b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The village is lit by a bright full moon and military headlights, and the dump truck used for the infiltration remains present. Charlie stands with his previously dented, punctured body and impaired systems; B-200's gun-hands are still intact. 박철진: He holds a raised pistol in a hostage-threatening position. 현우: He is held captive in the village and retains his earlier fighting injuries.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The village is lit by a bright full moon and military headlights, and the dump truck used for the infiltration remains present. Charlie stands with his previously dented, punctured body and impaired systems; B-200's gun-hands are still intact. 박철진: He holds a raised pistol in a hostage-threatening position. 현우: He is held captive in the village and retains his earlier fighting injuries.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The village is lit by a bright full moon and military headlights, and the dump truck used for the infiltration remains present. Charlie stands with his previously dented, punctured body and impaired systems; B-200's gun-hands are still intact. 박철진: He holds a raised pistol in a hostage-threatening position. 현우: He is held captive in the village and retains his earlier fighting injuries.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh40__bgfirst_bg.png",
     "asset_id": "6d5cb97b-e1e7-40f2-8bdc-880bfadc45a7",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S67sh40.png",
     "asset_id": "a32fd8a9-4876-47cb-ad74-e319feef74f4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_lodging_yard_928b09.png",
     "asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:822027>",
     "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "박철진이 현우의 우측 관자놀이에 총구를 밀착하여 겨누고 있으며 두 사람의 시선은 측면을 향함.",
    "built_space": "야외 마을 배경의 조명과 건물은 일치하나, 중앙의 군용 트럭이 두 대로 중복 복제되어 배치됨.",
    "entities": "현우와 박철진의 인상착의는 참조와 일치하며, 총기 프롭은 참조 이미지의 개머리판이 누락된 형태임.",
    "hard_violations": [
     "[gemini-pro] 배경의 군용 트럭이 두 대로 복제됨 (단일 사물의 중복 오류)",
     "[gpt-high] 장소 참조의 군용 트럭 한 대를 유사한 두 대로 복제하여 배경에 추가했다."
    ],
    "physics": "박철진의 오른손이 총을 정상적으로 쥐고 현우의 머리에 맞대어 지지하고 있음."
   },
   {
    "label": "B",
    "direction": "박철진이 현우의 좌측 관자놀이에 총구를 겨누고, 현우는 정면을 응시함.",
    "built_space": "참조 이미지와 일치하는 단일 트럭과 주변 구조물이 올바르게 배치된 야외 공간.",
    "entities": "두 인물의 외모 및 부상 표현은 참조와 부합하나, 총기의 개머리판이 생략됨.",
    "hard_violations": [
     "[gemini-pro] 현우의 왼쪽 어깨에 올려진 손이 해부학적으로 불가능한 형태(박철진의 두 번째 오른손)를 띠고 있음 (물리적으로 불가능한 해부학)"
    ],
    "physics": "박철진의 오른손이 총을 쥐고 있으나, 어깨를 짚고 있는 추가적인 손의 인체 구조적 출처가 불가능함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경에 트럭이 복제된 치명적 오류가 있으나, 전경의 총기 겨눔 액션과 인물 묘사는 비교적 기준에 부합합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "총을 쥔 손 외에 어깨를 잡고 있는 해부학적으로 불가능한 여분의 손이 존재하여 치명적인 하드 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 현우의 우측 관자놀이에 총구를 밀착하여 겨누고 있으며 두 사람의 시선은 측면을 향함.",
        "built_space": "야외 마을 배경의 조명과 건물은 일치하나, 중앙의 군용 트럭이 두 대로 중복 복제되어 배치됨.",
        "entities": "현우와 박철진의 인상착의는 참조와 일치하며, 총기 프롭은 참조 이미지의 개머리판이 누락된 형태임.",
        "hard_violations": [
         "배경의 군용 트럭이 두 대로 복제됨 (단일 사물의 중복 오류)"
        ],
        "physics": "박철진의 오른손이 총을 정상적으로 쥐고 현우의 머리에 맞대어 지지하고 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 현우의 좌측 관자놀이에 총구를 겨누고, 현우는 정면을 응시함.",
        "built_space": "참조 이미지와 일치하는 단일 트럭과 주변 구조물이 올바르게 배치된 야외 공간.",
        "entities": "두 인물의 외모 및 부상 표현은 참조와 부합하나, 총기의 개머리판이 생략됨.",
        "hard_violations": [
         "현우의 왼쪽 어깨에 올려진 손이 해부학적으로 불가능한 형태(박철진의 두 번째 오른손)를 띠고 있음 (물리적으로 불가능한 해부학)"
        ],
        "physics": "박철진의 오른손이 총을 쥐고 있으나, 어깨를 짚고 있는 추가적인 손의 인체 구조적 출처가 불가능함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경에 트럭이 복제된 치명적 오류가 있으나, 전경의 총기 겨눔 액션과 인물 묘사는 비교적 기준에 부합합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "총을 쥔 손 외에 어깨를 잡고 있는 해부학적으로 불가능한 여분의 손이 존재하여 치명적인 하드 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 현우의 우측 관자놀이에 총구를 밀착하여 겨누고 있으며 두 사람의 시선은 측면을 향함.",
        "built_space": "야외 마을 배경의 조명과 건물은 일치하나, 중앙의 군용 트럭이 두 대로 중복 복제되어 배치됨.",
        "entities": "현우와 박철진의 인상착의는 참조와 일치하며, 총기 프롭은 참조 이미지의 개머리판이 누락된 형태임.",
        "hard_violations": [
         "배경의 군용 트럭이 두 대로 복제됨 (단일 사물의 중복 오류)"
        ],
        "physics": "박철진의 오른손이 총을 정상적으로 쥐고 현우의 머리에 맞대어 지지하고 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 현우의 좌측 관자놀이에 총구를 겨누고, 현우는 정면을 응시함.",
        "built_space": "참조 이미지와 일치하는 단일 트럭과 주변 구조물이 올바르게 배치된 야외 공간.",
        "entities": "두 인물의 외모 및 부상 표현은 참조와 부합하나, 총기의 개머리판이 생략됨.",
        "hard_violations": [
         "현우의 왼쪽 어깨에 올려진 손이 해부학적으로 불가능한 형태(박철진의 두 번째 오른손)를 띠고 있음 (물리적으로 불가능한 해부학)"
        ],
        "physics": "박철진의 오른손이 총을 쥐고 있으나, 어깨를 짚고 있는 추가적인 손의 인체 구조적 출처가 불가능함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "관자놀이에 닿는 총구와 박철진의 밀착 제압을 명확하게 구현하고 장소도 유지하지만, 참조 총기의 개머리판이 빠져 있다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "밀착 상체와 권총 위협은 구현했으나, 장소 참조의 군용 트럭을 두 대로 복제했으며 총구 접촉점도 관자놀이보다 눈꼬리에 치우친다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 손에서 오른쪽으로 뻗은 권총 총구가 현우의 귀 위쪽 관자놀이에 닿는다. 손과 머리 사이에 총기 측면이 드러나 위협 방향이 분명하다. 두 사람은 모두 화면 오른쪽 바깥을 바라보며, 그 시선의 대상은 화면에 없다.",
        "built_space": "두 사람의 머리와 상체가 전경을 채운다. 뒤에는 좌우의 낡은 기와집과 돌담, 왼쪽 불붙은 드럼통 하나, 오른쪽의 화염, 중앙 오른쪽 군용 트럭 한 대와 전조등 두 개가 보인다. 오른쪽 전신주와 보름달도 참조 장소의 배치에 부합한다. 인물들은 실외 길에 있으며 시설과 충돌하거나 불가능한 반사가 나타나지 않는다.",
        "entities": "보이는 인물은 현우와 박철진 두 명뿐이다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 회색 셔츠와 얼굴의 전투 상처를 갖췄다. 박철진은 짧게 정돈한 검은 머리의 중년 동아시아계 남성이며 검은 제복과 일부 드러난 붉은 완장이 참조에 부합한다. 한국계라는 설정과 외모상 충돌은 없으나 국적 자체는 영상으로 판별할 수 없다. 권총은 어두운 금속 재질과 측면 형상은 유사하지만 참조의 긴 개머리판이 없다. 덤프트럭과 다른 인물들의 상태는 이 근접 프레임에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "권총은 박철진의 손이 손잡이를 감싸 지지하고, 손목과 검은 소매의 팔이 자연스럽게 이어진다. 다른 팔은 현우의 어깨와 윗가슴을 감싸며 손이 셔츠 위에 닿아 실제 제압 동작으로 성립한다. 두 사람의 상체는 아래 프레임 밖 몸통으로 이어지며 공중에 떠 있는 징후는 없다. 발이 잘린 것은 근접 촬영에 따른 것으로 지지 결함은 아니다."
       },
       {
        "label": "B",
        "direction": "권총은 박철진의 손에서 화면 왼쪽의 현우 머리를 향한다. 총구가 닿는 지점은 눈꼬리 바로 옆으로, 지정된 관자놀이 접촉보다 낮고 앞쪽에 치우쳐 보인다. 총기 측면과 잡은 손은 명확하다. 두 사람의 시선은 모두 화면 오른쪽 바깥을 향하며 대상은 보이지 않는다.",
        "built_space": "현우가 왼쪽 전경, 박철진이 중앙에서 오른쪽 전경을 차지하는 밀착 상체 구도다. 낡은 기와집과 돌담, 오른쪽 전신주, 보름달, 도로의 콘크리트 장벽과 상자 더미가 보인다. 그러나 참조에서 군용 트럭 한 대가 있는 배경에 거의 같은 트럭 두 대가 나란히 놓여 전조등도 네 개로 늘었다. 고정된 장소의 차량 배치가 복제되어 바뀌었다.",
        "entities": "인물은 현우와 박철진 두 명으로 제한되어 있다. 현우의 검은 헝클어진 머리, 회색 셔츠와 얼굴 상처는 맞지만 참조보다 얼굴이 다소 성숙해 보인다. 박철진은 중년 동아시아계 남성의 얼굴, 짧은 검은 머리, 검은 제복과 붉은 완장을 갖춘다. 국적은 외모만으로 확정할 수 없다. 금속 권총은 참조와 비슷한 계열이지만 긴 개머리판이 없고 상부가 더 밝은 은색이다. 배경 군용 트럭은 한 대가 아니라 두 대다.",
        "hard_violations": [
         "장소 참조의 군용 트럭 한 대를 유사한 두 대로 복제하여 배경에 추가했다."
        ],
        "physics": "권총 손잡이는 박철진의 손이 쥐고 있고 손목과 제복 소매가 연결되어 있어 무기가 떠 있지 않다. 팔을 굽혀 옆 사람의 머리에 총구를 대는 동작은 신체적으로 가능하다. 두 사람의 몸통은 프레임 아래로 이어지며 부유나 불가능한 관절 배치는 보이지 않는다. 배경 트럭도 도로 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "관자놀이에 닿는 총구와 박철진의 밀착 제압을 명확하게 구현하고 장소도 유지하지만, 참조 총기의 개머리판이 빠져 있다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "밀착 상체와 권총 위협은 구현했으나, 장소 참조의 군용 트럭을 두 대로 복제했으며 총구 접촉점도 관자놀이보다 눈꼬리에 치우친다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 손에서 오른쪽으로 뻗은 권총 총구가 현우의 귀 위쪽 관자놀이에 닿는다. 손과 머리 사이에 총기 측면이 드러나 위협 방향이 분명하다. 두 사람은 모두 화면 오른쪽 바깥을 바라보며, 그 시선의 대상은 화면에 없다.",
        "built_space": "두 사람의 머리와 상체가 전경을 채운다. 뒤에는 좌우의 낡은 기와집과 돌담, 왼쪽 불붙은 드럼통 하나, 오른쪽의 화염, 중앙 오른쪽 군용 트럭 한 대와 전조등 두 개가 보인다. 오른쪽 전신주와 보름달도 참조 장소의 배치에 부합한다. 인물들은 실외 길에 있으며 시설과 충돌하거나 불가능한 반사가 나타나지 않는다.",
        "entities": "보이는 인물은 현우와 박철진 두 명뿐이다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 회색 셔츠와 얼굴의 전투 상처를 갖췄다. 박철진은 짧게 정돈한 검은 머리의 중년 동아시아계 남성이며 검은 제복과 일부 드러난 붉은 완장이 참조에 부합한다. 한국계라는 설정과 외모상 충돌은 없으나 국적 자체는 영상으로 판별할 수 없다. 권총은 어두운 금속 재질과 측면 형상은 유사하지만 참조의 긴 개머리판이 없다. 덤프트럭과 다른 인물들의 상태는 이 근접 프레임에서 확인되지 않는다.",
        "hard_violations": [],
        "physics": "권총은 박철진의 손이 손잡이를 감싸 지지하고, 손목과 검은 소매의 팔이 자연스럽게 이어진다. 다른 팔은 현우의 어깨와 윗가슴을 감싸며 손이 셔츠 위에 닿아 실제 제압 동작으로 성립한다. 두 사람의 상체는 아래 프레임 밖 몸통으로 이어지며 공중에 떠 있는 징후는 없다. 발이 잘린 것은 근접 촬영에 따른 것으로 지지 결함은 아니다."
       },
       {
        "label": "A",
        "direction": "권총은 박철진의 손에서 화면 왼쪽의 현우 머리를 향한다. 총구가 닿는 지점은 눈꼬리 바로 옆으로, 지정된 관자놀이 접촉보다 낮고 앞쪽에 치우쳐 보인다. 총기 측면과 잡은 손은 명확하다. 두 사람의 시선은 모두 화면 오른쪽 바깥을 향하며 대상은 보이지 않는다.",
        "built_space": "현우가 왼쪽 전경, 박철진이 중앙에서 오른쪽 전경을 차지하는 밀착 상체 구도다. 낡은 기와집과 돌담, 오른쪽 전신주, 보름달, 도로의 콘크리트 장벽과 상자 더미가 보인다. 그러나 참조에서 군용 트럭 한 대가 있는 배경에 거의 같은 트럭 두 대가 나란히 놓여 전조등도 네 개로 늘었다. 고정된 장소의 차량 배치가 복제되어 바뀌었다.",
        "entities": "인물은 현우와 박철진 두 명으로 제한되어 있다. 현우의 검은 헝클어진 머리, 회색 셔츠와 얼굴 상처는 맞지만 참조보다 얼굴이 다소 성숙해 보인다. 박철진은 중년 동아시아계 남성의 얼굴, 짧은 검은 머리, 검은 제복과 붉은 완장을 갖춘다. 국적은 외모만으로 확정할 수 없다. 금속 권총은 참조와 비슷한 계열이지만 긴 개머리판이 없고 상부가 더 밝은 은색이다. 배경 군용 트럭은 한 대가 아니라 두 대다.",
        "hard_violations": [
         "장소 참조의 군용 트럭 한 대를 유사한 두 대로 복제하여 배경에 추가했다."
        ],
        "physics": "권총 손잡이는 박철진의 손이 쥐고 있고 손목과 제복 소매가 연결되어 있어 무기가 떠 있지 않다. 팔을 굽혀 옆 사람의 머리에 총구를 대는 동작은 신체적으로 가능하다. 두 사람의 몸통은 프레임 아래로 이어지며 부유나 불가능한 관절 배치는 보이지 않는다. 배경 트럭도 도로 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.125,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 배경의 군용 트럭이 두 대로 복제됨 (단일 사물의 중복 오류)",
     "[gpt-high] 장소 참조의 군용 트럭 한 대를 유사한 두 대로 복제하여 배경에 추가했다."
    ],
    "B": [
     "[gemini-pro] 현우의 왼쪽 어깨에 올려진 손이 해부학적으로 불가능한 형태(박철진의 두 번째 오른손)를 띠고 있음 (물리적으로 불가능한 해부학)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1125,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1125,
    "verdict_ko": "배경에 트럭이 복제된 치명적 오류가 있으나, 전경의 총기 겨눔 액션과 인물 묘사는 비교적 기준에 부합합니다.  ★위반: [gemini-pro] 배경의 군용 트럭이 두 대로 복제됨 (단일 사물의 중복 오류) / [gpt-high] 장소 참조의 군용 트럭 한 대를 유사한 두 대로 복제하여 배경에 추가했다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "총을 쥔 손 외에 어깨를 잡고 있는 해부학적으로 불가능한 여분의 손이 존재하여 치명적인 하드 위반이 발생했습니다.  ★위반: [gemini-pro] 현우의 왼쪽 어깨에 올려진 손이 해부학적으로 불가능한 형태(박철진의 두 번째 오른손)를 띠고 있음 (물리적으로 불가능한 해부학)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_lodging_yard_928b09.png",
    "asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:822027>",
    "asset_id": "4cd26332-bd3b-46be-b250-493e44394a2f",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-5fb4-77e8-a413-82a75c4ef142",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh40__bgfirst_bg.png",
   "bg_asset_id": "6d5cb97b-e1e7-40f2-8bdc-880bfadc45a7",
   "bg_record_key": "S67sh40::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "village_lodging_yard",
   "groupbg_asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S67sh40::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:13:05.321652+00:00",
  "fingerprint": "26e79952f265360a078b2fa8281a22b75bc6b9773fdf59922cee2f0564aec70f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S67sh40_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S67sh40_sel.png",
  "source_sha256": "62cda6557eaed4cc083a5bc053c96aaff6fb5ce53094c7d55c745b05561142ed",
  "file": "S67sh40_cine.png",
  "staged_sha256": "bb451dba679d57b9a3dff5838cda5e79ab1589a26fa4880ebba9e70a8f8db274",
  "latency_ms": 10414
 },
 "S67sh49::signage": {
  "fp": "9e3dbed7e7b2d35a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S67sh49": {
  "input_fingerprint": "ffe843da5c273202",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The militia trucks are still upright and traveling away; Charlie is confined in the transport cage with heavy restraints around his neck and both hands, retaining his earlier battle damage. B-200 blocks the route with intact gun-hands aimed forward.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The militia trucks are still upright and traveling away; Charlie is confined in the transport cage with heavy restraints around his neck and both hands, retaining his earlier battle damage. B-200 blocks the route with intact gun-hands aimed forward.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The militia trucks are still upright and traveling away; Charlie is confined in the transport cage with heavy restraints around his neck and both hands, retaining his earlier battle damage. B-200 blocks the route with intact gun-hands aimed forward.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh49__bgfirst_bg.png",
     "asset_id": "5b89bd25-951b-4371-9526-4d76774415ae",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S67sh49.png",
     "asset_id": "f64f2986-6e3d-4497-9926-988ba8db7762",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
     "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "B-200은 도로 우측에서 왼쪽 허공을 겨누며, 트럭의 정면 진행 경로를 가로막지 않습니다.",
    "built_space": "폐허가 된 마을의 흙길이며, 왼쪽 편에 멀어지는 방향의 트럭 1대가 있습니다.",
    "entities": "B-200의 양팔이 권총 레퍼런스로 기괴하게 변형되었습니다. 찰리는 트럭 뒤칸에 손이 묶인 채 서 있습니다.",
    "hard_violations": [
     "[gemini-pro] 로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
     "[gemini-pro] 트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류",
     "[gpt-high] B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
    ],
    "physics": "합성된 거대한 권총 팔의 부품들이 물리적 지지 없이 매달려 있습니다."
   },
   {
    "label": "B",
    "direction": "B-200은 도로 중앙에서 정면을 겨눕니다. 오른쪽 트럭은 로봇을 향해, 왼쪽 트럭은 카메라를 향해 달립니다.",
    "built_space": "달빛이 비치는 마을 흙길에 2대의 트럭이 배치되어 있습니다.",
    "entities": "B-200의 형태는 캐릭터 레퍼런스와 일치합니다. 찰리는 목과 손이 묶인 채 우측 트럭에 앉아 있습니다. 총기 레퍼런스는 누락되었습니다.",
    "hard_violations": [
     "[gemini-pro] 출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반",
     "[gpt-high] 가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
    ],
    "physics": "모든 차량과 로봇의 몸체는 지면에 안정적으로 닿아 지지받고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "B-200의 외형과 인물 구속 상태는 잘 표현되었으나, 수송대 트럭들이 서로 반대 방향으로 달리는 치명적인 동선 오류가 발생했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "로봇의 팔에 권총 레퍼런스가 그대로 합성되는 심각한 형태 오류가 있으며, 트럭의 정면을 가로막지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200은 도로 우측에서 왼쪽 허공을 겨누며, 트럭의 정면 진행 경로를 가로막지 않습니다.",
        "built_space": "폐허가 된 마을의 흙길이며, 왼쪽 편에 멀어지는 방향의 트럭 1대가 있습니다.",
        "entities": "B-200의 양팔이 권총 레퍼런스로 기괴하게 변형되었습니다. 찰리는 트럭 뒤칸에 손이 묶인 채 서 있습니다.",
        "hard_violations": [
         "로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
         "트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류"
        ],
        "physics": "합성된 거대한 권총 팔의 부품들이 물리적 지지 없이 매달려 있습니다."
       },
       {
        "label": "B",
        "direction": "B-200은 도로 중앙에서 정면을 겨눕니다. 오른쪽 트럭은 로봇을 향해, 왼쪽 트럭은 카메라를 향해 달립니다.",
        "built_space": "달빛이 비치는 마을 흙길에 2대의 트럭이 배치되어 있습니다.",
        "entities": "B-200의 형태는 캐릭터 레퍼런스와 일치합니다. 찰리는 목과 손이 묶인 채 우측 트럭에 앉아 있습니다. 총기 레퍼런스는 누락되었습니다.",
        "hard_violations": [
         "출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반"
        ],
        "physics": "모든 차량과 로봇의 몸체는 지면에 안정적으로 닿아 지지받고 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "B-200의 외형과 인물 구속 상태는 잘 표현되었으나, 수송대 트럭들이 서로 반대 방향으로 달리는 치명적인 동선 오류가 발생했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "로봇의 팔에 권총 레퍼런스가 그대로 합성되는 심각한 형태 오류가 있으며, 트럭의 정면을 가로막지 못했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "B-200은 도로 우측에서 왼쪽 허공을 겨누며, 트럭의 정면 진행 경로를 가로막지 않습니다.",
        "built_space": "폐허가 된 마을의 흙길이며, 왼쪽 편에 멀어지는 방향의 트럭 1대가 있습니다.",
        "entities": "B-200의 양팔이 권총 레퍼런스로 기괴하게 변형되었습니다. 찰리는 트럭 뒤칸에 손이 묶인 채 서 있습니다.",
        "hard_violations": [
         "로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
         "트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류"
        ],
        "physics": "합성된 거대한 권총 팔의 부품들이 물리적 지지 없이 매달려 있습니다."
       },
       {
        "label": "B",
        "direction": "B-200은 도로 중앙에서 정면을 겨눕니다. 오른쪽 트럭은 로봇을 향해, 왼쪽 트럭은 카메라를 향해 달립니다.",
        "built_space": "달빛이 비치는 마을 흙길에 2대의 트럭이 배치되어 있습니다.",
        "entities": "B-200의 형태는 캐릭터 레퍼런스와 일치합니다. 찰리는 목과 손이 묶인 채 우측 트럭에 앉아 있습니다. 총기 레퍼런스는 누락되었습니다.",
        "hard_violations": [
         "출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반"
        ],
        "physics": "모든 차량과 로봇의 몸체는 지면에 안정적으로 닿아 지지받고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "장소와 B-200의 장갑·포신은 가깝지만, 포구가 트럭을 겨누지 않고 차량의 대향 진행도 성립하지 않으며 금지된 인물들이 추가되었다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전신 와이드와 트럭 쪽으로 향한 조준은 상대적으로 낫지만, 행렬 앞 차단 구도가 불명확하고 추가 인물 및 고사포 대신 확대된 권총형 무기가 요구를 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 머리와 양팔 포구는 화면 오른쪽을 향한다. 왼쪽 전경 트럭이나 오른쪽 아래 수송차를 정조준하는 축이 아니다. 왼쪽 트럭은 전면과 전조등을 카메라 쪽으로 드러내며 B-200에서 멀어지는 방향이고, 오른쪽 수송차는 후면을 보이며 왼쪽 안쪽을 향한다. 두 차량이 왼쪽 아래에서 중경의 B-200을 향해 함께 접근하는 관계가 성립하지 않는다.",
        "built_space": "흙길 양쪽의 파손된 목조 기와집, 오른쪽 돌담, 왼쪽 전경 불통 하나, 오른쪽 높은 횃불 바구니 하나와 안쪽의 여러 작은 화점이 장소 참조와 가깝다. 트럭 두 대는 전경 양쪽에 있고 B-200은 그 뒤 도로 중앙에 있다. 전신은 보이지만, 적어도 왼쪽 트럭의 진행 앞을 막는 배치는 아니다.",
        "entities": "B-200 한 기의 짙은 회색 장갑판, 육중한 관절, 양팔의 다연장 포신은 캐릭터 참조와 가깝다. 별도 소품 참조의 개머리판 달린 권총은 보이지 않는다. 왼쪽 트럭에 적어도 적재함 인물 네 명과 운전자 한 명, 오른쪽 철창에 성인 남성 한 명이 보여 B-200만 보이라는 제한을 어긴다. 철창 남성은 목과 손 주변의 사슬이 보이나 기존 전투 손상은 확정하기 어렵다. 차량은 모두 직립해 있고, 밝은 달밤이며 발포 섬광은 없다.",
        "hard_violations": [
         "가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
        ],
        "physics": "B-200은 양발을 흙길에 딛고 있고 포신은 팔의 기계 구조에 연결되어 지지된다. 트럭은 타이어로 노면에 서 있으며 먼지가 이동을 암시한다. 사람들의 하체는 차체에 가려지지만 적재함·운전석·철창 바닥이 지지할 수 있는 위치다. 지지 없이 떠 있는 몸이나 무기는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "B-200의 머리와 두 무기의 총구는 화면 왼쪽 트럭들 쪽을 향한다. 가까운 철창 수송차 쪽으로 겨누는 방향은 A보다 분명하고 카메라도 사선에서 벗어나 있다. 다만 트럭들은 후면을 보인 채 왼쪽 도로 안쪽으로 이어져, 오른쪽의 B-200을 향해 접근하기보다는 그 옆을 지나 멀어지는 행렬처럼 보인다.",
        "built_space": "B-200의 전신은 도로 오른쪽 중경에, 가장 가까운 수송차는 왼쪽 전경에 배치된다. 왼쪽으로 최소 네 대의 차량이 이어진다. 파손된 기와집과 돌담·산 능선은 있으나, 참조의 중앙 마을길과 달리 오른쪽에 긴 금속 가드레일과 전선 달린 전신주가 두드러진다. 행렬의 주행로는 B-200 왼쪽으로 열려 있어 도로를 정면 차단하는 관계가 약하다.",
        "entities": "B-200의 회색 장갑과 굵은 기계 다리는 유지되지만, 양손 고사포가 아니라 개머리판·권총 손잡이가 있는 거대한 총 두 정이 보인다. 이는 소품 참조의 형태에는 가깝지만 캐릭터 참조의 다연장 고사포 손과 다르다. 가까운 철창에 목과 손이 구속된 남성 한 명, 다음 트럭에 헬멧 쓴 인물 한 명이 보여 인물 제한을 위반한다. 기존 전투 손상은 명확하지 않다. 달빛 아래 차량들은 직립해 있고 발포 섬광이나 뚜렷하게 읽히는 문자는 없다.",
        "hard_violations": [
         "B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
        ],
        "physics": "B-200의 벌린 두 발은 노면에 확실히 닿아 체중을 지지한다. 총들은 팔 앞의 기계 연결부에 걸쳐 있어 완전히 떠 있다고 단정할 수는 없지만, 아래로 노출된 권총 손잡이를 손으로 잡아 조작하는 모습은 없다. 차량 바퀴는 지면에 닿아 있고 두 인물도 차량 내부 바닥이 지지할 수 있는 위치다. 트럭의 실제 주행을 보여주는 동작 단서는 약하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "장소와 B-200의 장갑·포신은 가깝지만, 포구가 트럭을 겨누지 않고 차량의 대향 진행도 성립하지 않으며 금지된 인물들이 추가되었다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전신 와이드와 트럭 쪽으로 향한 조준은 상대적으로 낫지만, 행렬 앞 차단 구도가 불명확하고 추가 인물 및 고사포 대신 확대된 권총형 무기가 요구를 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "B-200의 머리와 양팔 포구는 화면 오른쪽을 향한다. 왼쪽 전경 트럭이나 오른쪽 아래 수송차를 정조준하는 축이 아니다. 왼쪽 트럭은 전면과 전조등을 카메라 쪽으로 드러내며 B-200에서 멀어지는 방향이고, 오른쪽 수송차는 후면을 보이며 왼쪽 안쪽을 향한다. 두 차량이 왼쪽 아래에서 중경의 B-200을 향해 함께 접근하는 관계가 성립하지 않는다.",
        "built_space": "흙길 양쪽의 파손된 목조 기와집, 오른쪽 돌담, 왼쪽 전경 불통 하나, 오른쪽 높은 횃불 바구니 하나와 안쪽의 여러 작은 화점이 장소 참조와 가깝다. 트럭 두 대는 전경 양쪽에 있고 B-200은 그 뒤 도로 중앙에 있다. 전신은 보이지만, 적어도 왼쪽 트럭의 진행 앞을 막는 배치는 아니다.",
        "entities": "B-200 한 기의 짙은 회색 장갑판, 육중한 관절, 양팔의 다연장 포신은 캐릭터 참조와 가깝다. 별도 소품 참조의 개머리판 달린 권총은 보이지 않는다. 왼쪽 트럭에 적어도 적재함 인물 네 명과 운전자 한 명, 오른쪽 철창에 성인 남성 한 명이 보여 B-200만 보이라는 제한을 어긴다. 철창 남성은 목과 손 주변의 사슬이 보이나 기존 전투 손상은 확정하기 어렵다. 차량은 모두 직립해 있고, 밝은 달밤이며 발포 섬광은 없다.",
        "hard_violations": [
         "가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
        ],
        "physics": "B-200은 양발을 흙길에 딛고 있고 포신은 팔의 기계 구조에 연결되어 지지된다. 트럭은 타이어로 노면에 서 있으며 먼지가 이동을 암시한다. 사람들의 하체는 차체에 가려지지만 적재함·운전석·철창 바닥이 지지할 수 있는 위치다. 지지 없이 떠 있는 몸이나 무기는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "B-200의 머리와 두 무기의 총구는 화면 왼쪽 트럭들 쪽을 향한다. 가까운 철창 수송차 쪽으로 겨누는 방향은 A보다 분명하고 카메라도 사선에서 벗어나 있다. 다만 트럭들은 후면을 보인 채 왼쪽 도로 안쪽으로 이어져, 오른쪽의 B-200을 향해 접근하기보다는 그 옆을 지나 멀어지는 행렬처럼 보인다.",
        "built_space": "B-200의 전신은 도로 오른쪽 중경에, 가장 가까운 수송차는 왼쪽 전경에 배치된다. 왼쪽으로 최소 네 대의 차량이 이어진다. 파손된 기와집과 돌담·산 능선은 있으나, 참조의 중앙 마을길과 달리 오른쪽에 긴 금속 가드레일과 전선 달린 전신주가 두드러진다. 행렬의 주행로는 B-200 왼쪽으로 열려 있어 도로를 정면 차단하는 관계가 약하다.",
        "entities": "B-200의 회색 장갑과 굵은 기계 다리는 유지되지만, 양손 고사포가 아니라 개머리판·권총 손잡이가 있는 거대한 총 두 정이 보인다. 이는 소품 참조의 형태에는 가깝지만 캐릭터 참조의 다연장 고사포 손과 다르다. 가까운 철창에 목과 손이 구속된 남성 한 명, 다음 트럭에 헬멧 쓴 인물 한 명이 보여 인물 제한을 위반한다. 기존 전투 손상은 명확하지 않다. 달빛 아래 차량들은 직립해 있고 발포 섬광이나 뚜렷하게 읽히는 문자는 없다.",
        "hard_violations": [
         "B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
        ],
        "physics": "B-200의 벌린 두 발은 노면에 확실히 닿아 체중을 지지한다. 총들은 팔 앞의 기계 연결부에 걸쳐 있어 완전히 떠 있다고 단정할 수는 없지만, 아래로 노출된 권총 손잡이를 손으로 잡아 조작하는 모습은 없다. 차량 바퀴는 지면에 닿아 있고 두 인물도 차량 내부 바닥이 지지할 수 있는 위치다. 트럭의 실제 주행을 보여주는 동작 단서는 약하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
     "[gemini-pro] 트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류",
     "[gpt-high] B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
    ],
    "B": [
     "[gemini-pro] 출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반",
     "[gpt-high] 가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1417,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "B-200의 외형과 인물 구속 상태는 잘 표현되었으나, 수송대 트럭들이 서로 반대 방향으로 달리는 치명적인 동선 오류가 발생했습니다.  ★위반: [gemini-pro] 출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반 / [gpt-high] 가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "로봇의 팔에 권총 레퍼런스가 그대로 합성되는 심각한 형태 오류가 있으며, 트럭의 정면을 가로막지 못했습니다.  ★위반: [gemini-pro] 로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태 / [gemini-pro] 트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류 / [gpt-high] B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
    "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-6484-7923-ae3f-3125fbce483b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh49__bgfirst_bg.png",
   "bg_asset_id": "5b89bd25-951b-4371-9526-4d76774415ae",
   "bg_record_key": "S67sh49::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "village_truck_ambush",
   "groupbg_asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S67sh49::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:54:59.307848+00:00",
  "fingerprint": "f617e70ed8a71cdaf21f220dcb357f5b03cf4da9fd968663dcab6b10bfb7f8fd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S67sh49_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S67sh49_sel.png",
  "source_sha256": "7d90931d09c2346a987f7446f4a3f8b97db1462586b95c10e3fb0ad103f0e07f",
  "file": "S67sh49_cine.png",
  "staged_sha256": "b87c1738e6edd7ec528856af8f077542f15aeab2848b198de64135cb9196ae4e",
  "latency_ms": 10614
 },
 "S67sh70::signage": {
  "fp": "7726b4ee80764f48",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S67sh70": {
  "input_fingerprint": "6cd2f234ab993c9e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): B-200 has been destroyed, with only a component from his chest remaining, held against Charlie's chest. No articulated head, torso or limbs remain to form a whole-body pose.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Overturned trucks and the blast wreckage remain after the smoke clears; B-200 is destroyed, leaving the recovered chest component. Charlie holds that component with his newly shot and previously battered body, closing his eyes without an established permanent loss of power.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): B-200 has been destroyed, with only a component from his chest remaining, held against Charlie's chest. No articulated head, torso or limbs remain to form a whole-body pose.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Overturned trucks and the blast wreckage remain after the smoke clears; B-200 is destroyed, leaving the recovered chest component. Charlie holds that component with his newly shot and previously battered body, closing his eyes without an established permanent loss of power.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): B-200 has been destroyed, with only a component from his chest remaining, held against Charlie's chest. No articulated head, torso or limbs remain to form a whole-body pose.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Overturned trucks and the blast wreckage remain after the smoke clears; B-200 is destroyed, leaving the recovered chest component. Charlie holds that component with his newly shot and previously battered body, closing his eyes without an established permanent loss of power.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh70__bgfirst_bg.png",
     "asset_id": "24c611fd-00da-42ad-afe0-c4324cb61919",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S67sh70.png",
     "asset_id": "0909ba75-540e-4c7a-b173-2342ee20e201",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
     "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 약간 아래쪽을 향하고 있으며, 두 눈의 불빛은 꺼져 있음.",
    "built_space": "달빛이 비치는 야간의 파괴된 마을로, 폐허가 된 건물과 잔해가 프레임 뒤편에 위치함.",
    "entities": "찰리(로봇 외형)가 등장하지만 레퍼런스 이미지의 화려한 셔츠를 입고 있지 않으며, 가슴 앞의 부품이 찰리의 몸통과 구분되지 않음.",
    "hard_violations": [
     "[gemini-pro] 부품을 쥐고 있어야 할 손이 프레임에 존재하지 않음",
     "[gemini-pro] 가슴에 품은 부품이 캐릭터의 신체 장갑과 기형적으로 융합되어 물리적으로 불가능한 구조를 보임"
    ],
    "physics": "가슴 앞의 부품을 지지하거나 잡고 있는 손이 없어 부품의 지지 상태가 물리적으로 설명되지 않음."
   },
   {
    "label": "B",
    "direction": "고개를 약간 숙인 채 정면 아래를 향하고 있으며, 눈의 디지털 불빛이 완전히 꺼져 있음.",
    "built_space": "야간의 부서진 잔해와 차량이 있는 폭발 현장 배경이 올바르게 묘사됨.",
    "entities": "레퍼런스의 꽃무늬 셔츠를 입은 찰리가 등장하며, 내부 회로가 드러난 B-200의 가슴 부품을 명확히 안고 있음.",
    "hard_violations": [],
    "physics": "오른손이 기계 부품을 단단히 쥐고 가슴에 밀착시켜 안정적으로 지지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 의상(꽃무늬 셔츠)을 잘 반영했으며, 꺼진 눈과 부품을 쥔 손의 묘사가 프롬프트의 요구사항을 정확히 충족합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스의 의상이 누락되었고, 부품을 들고 있어야 할 손이 보이지 않으며 구조물이 기형적으로 융합되어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 약간 아래쪽을 향하고 있으며, 두 눈의 불빛은 꺼져 있음.",
        "built_space": "달빛이 비치는 야간의 파괴된 마을로, 폐허가 된 건물과 잔해가 프레임 뒤편에 위치함.",
        "entities": "찰리(로봇 외형)가 등장하지만 레퍼런스 이미지의 화려한 셔츠를 입고 있지 않으며, 가슴 앞의 부품이 찰리의 몸통과 구분되지 않음.",
        "hard_violations": [
         "부품을 쥐고 있어야 할 손이 프레임에 존재하지 않음",
         "가슴에 품은 부품이 캐릭터의 신체 장갑과 기형적으로 융합되어 물리적으로 불가능한 구조를 보임"
        ],
        "physics": "가슴 앞의 부품을 지지하거나 잡고 있는 손이 없어 부품의 지지 상태가 물리적으로 설명되지 않음."
       },
       {
        "label": "B",
        "direction": "고개를 약간 숙인 채 정면 아래를 향하고 있으며, 눈의 디지털 불빛이 완전히 꺼져 있음.",
        "built_space": "야간의 부서진 잔해와 차량이 있는 폭발 현장 배경이 올바르게 묘사됨.",
        "entities": "레퍼런스의 꽃무늬 셔츠를 입은 찰리가 등장하며, 내부 회로가 드러난 B-200의 가슴 부품을 명확히 안고 있음.",
        "hard_violations": [],
        "physics": "오른손이 기계 부품을 단단히 쥐고 가슴에 밀착시켜 안정적으로 지지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 의상(꽃무늬 셔츠)을 잘 반영했으며, 꺼진 눈과 부품을 쥔 손의 묘사가 프롬프트의 요구사항을 정확히 충족합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스의 의상이 누락되었고, 부품을 들고 있어야 할 손이 보이지 않으며 구조물이 기형적으로 융합되어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 약간 아래쪽을 향하고 있으며, 두 눈의 불빛은 꺼져 있음.",
        "built_space": "달빛이 비치는 야간의 파괴된 마을로, 폐허가 된 건물과 잔해가 프레임 뒤편에 위치함.",
        "entities": "찰리(로봇 외형)가 등장하지만 레퍼런스 이미지의 화려한 셔츠를 입고 있지 않으며, 가슴 앞의 부품이 찰리의 몸통과 구분되지 않음.",
        "hard_violations": [
         "부품을 쥐고 있어야 할 손이 프레임에 존재하지 않음",
         "가슴에 품은 부품이 캐릭터의 신체 장갑과 기형적으로 융합되어 물리적으로 불가능한 구조를 보임"
        ],
        "physics": "가슴 앞의 부품을 지지하거나 잡고 있는 손이 없어 부품의 지지 상태가 물리적으로 설명되지 않음."
       },
       {
        "label": "B",
        "direction": "고개를 약간 숙인 채 정면 아래를 향하고 있으며, 눈의 디지털 불빛이 완전히 꺼져 있음.",
        "built_space": "야간의 부서진 잔해와 차량이 있는 폭발 현장 배경이 올바르게 묘사됨.",
        "entities": "레퍼런스의 꽃무늬 셔츠를 입은 찰리가 등장하며, 내부 회로가 드러난 B-200의 가슴 부품을 명확히 안고 있음.",
        "hard_violations": [],
        "physics": "오른손이 기계 부품을 단단히 쥐고 가슴에 밀착시켜 안정적으로 지지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "꺼진 두 눈과 가슴에 밀착해 움켜쥔 잔해, 참고의 꽃무늬 천은 잘 반영했지만, 얼굴 클로즈업보다 넓고 부품이 하단 일부가 아니라 화면의 큰 부분을 차지한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "달밤의 전복 트럭과 꺼진 눈은 맞지만, 얼굴보다 상반신과 배경을 강조한 넓은 구도이며 부품도 가슴 깊숙이 품기보다 아래쪽에 받쳐 들고 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개를 정면에서 조금 숙여 가슴에 안은 부품 쪽으로 얼굴을 향한다. 두 눈의 발광은 보이지 않으며 렌즈를 응시하는 인상도 약하다. 손가락은 부품 앞면을 감싸 몸 쪽으로 붙잡고 있다. 무기나 조준 대상은 없다.",
        "built_space": "왼쪽 배경에 기울어진 파손 트럭 한 대, 양옆에 무너진 기와지붕과 목조 건물, 오른쪽 지면에 금속 잔해와 타이어가 보인다. 참고의 폐허가 된 한국 마을길과 재료는 대체로 맞지만 정확히 같은 지점인지는 확인하기 어렵다. 찰리의 머리부터 넓은 어깨와 팔까지 들어와 요구된 얼굴 클로즈업보다 넓다. 반사면이나 중복된 고정 설비는 보이지 않는다.",
        "entities": "등장 인물은 비인간 로봇 찰리 한 명이다. 샌드 베이지 장갑, 육중한 팔, 흰 각진 마스크와 꽃무늬 천은 참고와 부합한다. 참고의 밀짚모자는 없고 얼굴 판의 세부 형상은 다르다. 양쪽 눈은 어둡고 장갑에는 긁힘과 오염이 있다. 손에 든 것은 배선과 회로가 노출된 분리된 기계 부품으로, 완전한 B-200은 없다. 다만 부품이 하단 가장자리에 일부만 보이지 않고 크게 드러난다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "부품 앞면을 큰 기계 손이 움켜쥐고 전완이 가슴 앞으로 이어져 있어 부품의 무게를 팔과 몸통이 받치는 관계가 보인다. 머리는 목 구조에 연결되어 있고 공중에 떠 있는 신체나 물체는 없다. 하체는 프레임 밖이므로 자세나 지면 접촉은 판단하지 않는다. 눈의 소등만으로 찰리를 의식 없는 상태라고 볼 근거는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 고개를 아래쪽으로 숙여 양팔 사이의 흉부 부품을 향한다. 두 눈은 발광하지 않는다. 손과 팔은 부품을 위로 제시하기보다 아래에서 안아 받치지만, 부품이 가슴 중심보다 낮게 놓여 깊이 끌어안은 관계는 약하다. 무기나 조준 대상은 없다.",
        "built_space": "좌우 배경에 옆으로 넘어진 트럭이 최소 두 대 보이고, 오른쪽에는 도로 난간, 멀리에는 기와지붕과 탑 형태의 건물이 보인다. 달과 산, 파괴된 도로는 참고의 야간 분위기와 맞지만 참고의 가까운 한옥 벽체 대신 차량과 열린 도로가 크게 부각된다. 머리부터 상반신과 양팔, 넓은 배경까지 보여 얼굴 클로즈업과 거리가 크다. 불가능한 반사나 명백한 설비 중복은 없다.",
        "entities": "찰리 한 명만 보이며 베이지 장갑과 큰 팔, 흰 마스크형 얼굴, 어두운 두 눈은 조건에 맞는다. 참고의 밀짚모자와 꽃무늬 천은 모두 없다. 얼굴의 점과 선, 장갑의 파손 흔적은 보인다. 안고 있는 것은 어깨 장갑처럼 넓게 벌어진 외장이 붙은 흉부 잔해로 보이며 별도의 머리나 팔다리는 확인되지 않는다. 잔해는 하단에 조금 걸치는 수준보다 훨씬 크고 온전하게 노출된다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "흉부 잔해는 가로지른 전완과 오른쪽에 보이는 기계 손에 의해 받쳐져 있어 떠 있지 않다. 머리와 팔도 몸통에 정상적으로 연결되어 있다. 하체와 착석 여부는 프레임 밖이라 확인할 수 없지만, 보이는 범위에서 지지 없는 물체나 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "꺼진 두 눈과 가슴에 밀착해 움켜쥔 잔해, 참고의 꽃무늬 천은 잘 반영했지만, 얼굴 클로즈업보다 넓고 부품이 하단 일부가 아니라 화면의 큰 부분을 차지한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "달밤의 전복 트럭과 꺼진 눈은 맞지만, 얼굴보다 상반신과 배경을 강조한 넓은 구도이며 부품도 가슴 깊숙이 품기보다 아래쪽에 받쳐 들고 있다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 고개를 정면에서 조금 숙여 가슴에 안은 부품 쪽으로 얼굴을 향한다. 두 눈의 발광은 보이지 않으며 렌즈를 응시하는 인상도 약하다. 손가락은 부품 앞면을 감싸 몸 쪽으로 붙잡고 있다. 무기나 조준 대상은 없다.",
        "built_space": "왼쪽 배경에 기울어진 파손 트럭 한 대, 양옆에 무너진 기와지붕과 목조 건물, 오른쪽 지면에 금속 잔해와 타이어가 보인다. 참고의 폐허가 된 한국 마을길과 재료는 대체로 맞지만 정확히 같은 지점인지는 확인하기 어렵다. 찰리의 머리부터 넓은 어깨와 팔까지 들어와 요구된 얼굴 클로즈업보다 넓다. 반사면이나 중복된 고정 설비는 보이지 않는다.",
        "entities": "등장 인물은 비인간 로봇 찰리 한 명이다. 샌드 베이지 장갑, 육중한 팔, 흰 각진 마스크와 꽃무늬 천은 참고와 부합한다. 참고의 밀짚모자는 없고 얼굴 판의 세부 형상은 다르다. 양쪽 눈은 어둡고 장갑에는 긁힘과 오염이 있다. 손에 든 것은 배선과 회로가 노출된 분리된 기계 부품으로, 완전한 B-200은 없다. 다만 부품이 하단 가장자리에 일부만 보이지 않고 크게 드러난다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "부품 앞면을 큰 기계 손이 움켜쥐고 전완이 가슴 앞으로 이어져 있어 부품의 무게를 팔과 몸통이 받치는 관계가 보인다. 머리는 목 구조에 연결되어 있고 공중에 떠 있는 신체나 물체는 없다. 하체는 프레임 밖이므로 자세나 지면 접촉은 판단하지 않는다. 눈의 소등만으로 찰리를 의식 없는 상태라고 볼 근거는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 고개를 아래쪽으로 숙여 양팔 사이의 흉부 부품을 향한다. 두 눈은 발광하지 않는다. 손과 팔은 부품을 위로 제시하기보다 아래에서 안아 받치지만, 부품이 가슴 중심보다 낮게 놓여 깊이 끌어안은 관계는 약하다. 무기나 조준 대상은 없다.",
        "built_space": "좌우 배경에 옆으로 넘어진 트럭이 최소 두 대 보이고, 오른쪽에는 도로 난간, 멀리에는 기와지붕과 탑 형태의 건물이 보인다. 달과 산, 파괴된 도로는 참고의 야간 분위기와 맞지만 참고의 가까운 한옥 벽체 대신 차량과 열린 도로가 크게 부각된다. 머리부터 상반신과 양팔, 넓은 배경까지 보여 얼굴 클로즈업과 거리가 크다. 불가능한 반사나 명백한 설비 중복은 없다.",
        "entities": "찰리 한 명만 보이며 베이지 장갑과 큰 팔, 흰 마스크형 얼굴, 어두운 두 눈은 조건에 맞는다. 참고의 밀짚모자와 꽃무늬 천은 모두 없다. 얼굴의 점과 선, 장갑의 파손 흔적은 보인다. 안고 있는 것은 어깨 장갑처럼 넓게 벌어진 외장이 붙은 흉부 잔해로 보이며 별도의 머리나 팔다리는 확인되지 않는다. 잔해는 하단에 조금 걸치는 수준보다 훨씬 크고 온전하게 노출된다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "흉부 잔해는 가로지른 전완과 오른쪽에 보이는 기계 손에 의해 받쳐져 있어 떠 있지 않다. 머리와 팔도 몸통에 정상적으로 연결되어 있다. 하체와 착석 여부는 프레임 밖이라 확인할 수 없지만, 보이는 범위에서 지지 없는 물체나 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.095,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.845,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 부품을 쥐고 있어야 할 손이 프레임에 존재하지 않음",
     "[gemini-pro] 가슴에 품은 부품이 캐릭터의 신체 장갑과 기형적으로 융합되어 물리적으로 불가능한 구조를 보임"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 845
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "레퍼런스의 의상(꽃무늬 셔츠)을 잘 반영했으며, 꺼진 눈과 부품을 쥔 손의 묘사가 프롬프트의 요구사항을 정확히 충족합니다."
   },
   {
    "label": "A",
    "score": 845,
    "verdict_ko": "레퍼런스의 의상이 누락되었고, 부품을 들고 있어야 할 손이 보이지 않으며 구조물이 기형적으로 융합되어 감점되었습니다.  ★위반: [gemini-pro] 부품을 쥐고 있어야 할 손이 프레임에 존재하지 않음 / [gemini-pro] 가슴에 품은 부품이 캐릭터의 신체 장갑과 기형적으로 융합되어 물리적으로 불가능한 구조를 보임"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
    "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-6961-71cb-85b8-a76bc014561c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh70__bgfirst_bg.png",
   "bg_asset_id": "24c611fd-00da-42ad-afe0-c4324cb61919",
   "bg_record_key": "S67sh70::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "village_truck_ambush",
   "groupbg_asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S67sh70::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:14:13.795831+00:00",
  "fingerprint": "6440eb96181532b30a7fb6ef9d0d7ba8df394f896523052ba06ba369d749176c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S67sh70_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S67sh70_sel.png",
  "source_sha256": "188c67c587905a1a7396fbbda217dc6194fbd42bc0c56aeb67de63946add3ab6",
  "file": "S67sh70_cine.png",
  "staged_sha256": "b2f2d4331c84ae41c9e3594d8a714b95a01b4559db567db88b2310d0f860c7cb",
  "latency_ms": 10621
 },
 "S68sh3::signage": {
  "fp": "4a83faa4cbfa3b76",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S68sh3": {
  "input_fingerprint": "052f65896b07a6fc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양 갈래 흙길에서 서로 다른 방향을 향해 달려가며 바퀴 뒤로 흙먼지를 일으키고 있는 두 낡은 트럭의 뒷모습 풀샷.\n\nLOCATION (lock): At a fork in the rural dirt road outside the village, where the two truck groups separate in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: 현우's truck on the left branch in the upper-left of the frame, midground, moves toward left branch extending away from the fork; 수빈's truck on the right branch in the upper-right of the frame, midground, moves toward right branch followed by the other vehicles.\n- KEY BACKGROUND ELEMENTS: 현우's truck (Old, roofless, and moving away along one branch) — Its rear and right side are visible as it recedes toward the upper left; used as Establish the branch the camera will subsequently follow; 수빈's truck and following vehicles (수빈's old roofless truck separates onto the other branch, with vehicles following it) — Rear profiles recede toward the upper right along the same branch; used as Make the separation of the two traveling groups spatially unambiguous; Forked dirt road (Both branches are being used, with dust rising behind the trucks' wheels) — The shared approach enters from the lower center and divides toward the upper corners; used as Provide a readable ground plan for the farewell and subsequent camera pursuit.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight keeps both branching routes and the departing trucks legible with restrained contrast and no added atmospheric treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Two old roofless trucks separate at the fork, with more vehicles following one branch. Charlie rides in the cargo bed with his battle damage unrepaired and the recovered B-200 chest component in his possession.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양 갈래 흙길에서 서로 다른 방향을 향해 달려가며 바퀴 뒤로 흙먼지를 일으키고 있는 두 낡은 트럭의 뒷모습 풀샷.\n\nLOCATION (lock): At a fork in the rural dirt road outside the village, where the two truck groups separate in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: 현우's truck on the left branch in the upper-left of the frame, midground, moves toward left branch extending away from the fork; 수빈's truck on the right branch in the upper-right of the frame, midground, moves toward right branch followed by the other vehicles.\n- KEY BACKGROUND ELEMENTS: 현우's truck (Old, roofless, and moving away along one branch) — Its rear and right side are visible as it recedes toward the upper left; used as Establish the branch the camera will subsequently follow; 수빈's truck and following vehicles (수빈's old roofless truck separates onto the other branch, with vehicles following it) — Rear profiles recede toward the upper right along the same branch; used as Make the separation of the two traveling groups spatially unambiguous; Forked dirt road (Both branches are being used, with dust rising behind the trucks' wheels) — The shared approach enters from the lower center and divides toward the upper corners; used as Provide a readable ground plan for the farewell and subsequent camera pursuit.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight keeps both branching routes and the departing trucks legible with restrained contrast and no added atmospheric treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Two old roofless trucks separate at the fork, with more vehicles following one branch. Charlie rides in the cargo bed with his battle damage unrepaired and the recovered B-200 chest component in his possession.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양 갈래 흙길에서 서로 다른 방향을 향해 달려가며 바퀴 뒤로 흙먼지를 일으키고 있는 두 낡은 트럭의 뒷모습 풀샷.\n\nLOCATION (lock): At a fork in the rural dirt road outside the village, where the two truck groups separate in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: 현우's truck on the left branch in the upper-left of the frame, midground, moves toward left branch extending away from the fork; 수빈's truck on the right branch in the upper-right of the frame, midground, moves toward right branch followed by the other vehicles.\n- KEY BACKGROUND ELEMENTS: 현우's truck (Old, roofless, and moving away along one branch) — Its rear and right side are visible as it recedes toward the upper left; used as Establish the branch the camera will subsequently follow; 수빈's truck and following vehicles (수빈's old roofless truck separates onto the other branch, with vehicles following it) — Rear profiles recede toward the upper right along the same branch; used as Make the separation of the two traveling groups spatially unambiguous; Forked dirt road (Both branches are being used, with dust rising behind the trucks' wheels) — The shared approach enters from the lower center and divides toward the upper corners; used as Provide a readable ground plan for the farewell and subsequent camera pursuit.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight keeps both branching routes and the departing trucks legible with restrained contrast and no added atmospheric treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Two old roofless trucks separate at the fork, with more vehicles following one branch. Charlie rides in the cargo bed with his battle damage unrepaired and the recovered B-200 chest component in his possession.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "두 트럭과 차량들이 갈래길을 따라 각각 멀어지는 방향으로 이동 중임.",
    "built_space": "화면 하단 중앙에서 시작되어 좌상단과 우상단으로 갈라지는 흙길이 명확하게 구현됨.",
    "entities": "왼쪽 트럭 짐칸에 짐이 덮여 있으나 찰리의 모습은 보이지 않음. 오른쪽 길의 오토바이에 사람이 탑승해 있음.",
    "hard_violations": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
    ],
    "physics": "차량들이 땅에 닿아 주행하며 바퀴 뒤로 흙먼지가 자연스럽게 발생함."
   },
   {
    "label": "B",
    "direction": "두 트럭과 오토바이 무리가 갈래길을 따라 양쪽으로 멀어지는 방향을 향하고 있음.",
    "built_space": "흙길이 중앙에서 좌우로 뻗어나가는 구조가 프레임에 맞게 배치됨.",
    "entities": "왼쪽 트럭의 짐칸이 완전히 비어 있어 찰리가 누락됨. 오른쪽 길을 따르는 다수의 오토바이에 사람들이 탑승해 있음.",
    "hard_violations": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
    ],
    "physics": "차량과 오토바이들이 노면 위를 주행하며 흙먼지를 일으키고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "사람을 포함하지 말라는 지침을 어기고 오토바이 탑승자를 생성했으며, 짐칸에 명시된 찰리(Charlie)가 누락되어 치명적 감점 요소가 됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "사람을 묘사하지 말라는 명시적 지침을 위반하고 다수의 오토바이 탑승자를 화면에 포함시켰으며, 짐칸이 비어 있어 실패함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 트럭과 차량들이 갈래길을 따라 각각 멀어지는 방향으로 이동 중임.",
        "built_space": "화면 하단 중앙에서 시작되어 좌상단과 우상단으로 갈라지는 흙길이 명확하게 구현됨.",
        "entities": "왼쪽 트럭 짐칸에 짐이 덮여 있으나 찰리의 모습은 보이지 않음. 오른쪽 길의 오토바이에 사람이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량들이 땅에 닿아 주행하며 바퀴 뒤로 흙먼지가 자연스럽게 발생함."
       },
       {
        "label": "B",
        "direction": "두 트럭과 오토바이 무리가 갈래길을 따라 양쪽으로 멀어지는 방향을 향하고 있음.",
        "built_space": "흙길이 중앙에서 좌우로 뻗어나가는 구조가 프레임에 맞게 배치됨.",
        "entities": "왼쪽 트럭의 짐칸이 완전히 비어 있어 찰리가 누락됨. 오른쪽 길을 따르는 다수의 오토바이에 사람들이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량과 오토바이들이 노면 위를 주행하며 흙먼지를 일으키고 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "사람을 포함하지 말라는 지침을 어기고 오토바이 탑승자를 생성했으며, 짐칸에 명시된 찰리(Charlie)가 누락되어 치명적 감점 요소가 됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "사람을 묘사하지 말라는 명시적 지침을 위반하고 다수의 오토바이 탑승자를 화면에 포함시켰으며, 짐칸이 비어 있어 실패함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "두 트럭과 차량들이 갈래길을 따라 각각 멀어지는 방향으로 이동 중임.",
        "built_space": "화면 하단 중앙에서 시작되어 좌상단과 우상단으로 갈라지는 흙길이 명확하게 구현됨.",
        "entities": "왼쪽 트럭 짐칸에 짐이 덮여 있으나 찰리의 모습은 보이지 않음. 오른쪽 길의 오토바이에 사람이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량들이 땅에 닿아 주행하며 바퀴 뒤로 흙먼지가 자연스럽게 발생함."
       },
       {
        "label": "B",
        "direction": "두 트럭과 오토바이 무리가 갈래길을 따라 양쪽으로 멀어지는 방향을 향하고 있음.",
        "built_space": "흙길이 중앙에서 좌우로 뻗어나가는 구조가 프레임에 맞게 배치됨.",
        "entities": "왼쪽 트럭의 짐칸이 완전히 비어 있어 찰리가 누락됨. 오른쪽 길을 따르는 다수의 오토바이에 사람들이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량과 오토바이들이 노면 위를 주행하며 흙먼지를 일으키고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 트럭의 좌우 상단 배치와 갈림길 풀샷은 더 정확하지만, 금지된 사람들이 등장하고 오른쪽 추가 차량들이 수빈의 트럭을 뒤따르지 않아 부적합하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "갈라지는 주행과 흙먼지는 구현했지만, 사람 등장 금지를 위반하고 트럭들이 화면 중앙에 내려와 있으며 추가 차량도 선행한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 트럭은 왼쪽 위 갈래로, 오른쪽 트럭은 오른쪽 위 갈래로 카메라에서 멀어진다. 두 차량의 후면이 보이고 왼쪽 트럭의 측면도 드러난다. 오른쪽 도로의 추가 트럭들과 오토바이들은 같은 방향으로 향하지만, 주된 오른쪽 트럭보다 앞에 있어 '뒤따르는 차량' 관계와 반대다. 인물들의 정확한 시선은 판독하기 어렵다.",
        "built_space": "아래 중앙의 공통 흙길 하나가 좌우 두 갈래로 나뉘며, 두 주된 트럭은 각각 화면 왼쪽 위와 오른쪽 위의 중경에 놓인다. 오른쪽 갈래에 추가 트럭 두 대와 오토바이 세 대가 보인다. 주변은 경작지와 낮은 산이며 실내 설비나 반사는 없다. 낮의 농촌 갈림길이라는 공간 조건과 와이드 구도는 잘 읽힌다.",
        "entities": "낡고 녹슨 화물 트럭 두 대, 양 갈래 흙길, 바퀴 뒤 흙먼지가 보인다. 적재함은 개방되어 있고 운전석 상부는 프레임이 두드러지지만 지붕 제거 여부는 완전히 명확하지 않다. 오토바이 탑승자 세 명과 오른쪽 주 트럭 운전석의 사람 형상이 보여 무인 화면 조건에 어긋난다. 인물의 국적·나이·성별은 이 크기에서 확인할 수 없다. 찰리의 손상 상태와 B-200 부품은 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
        ],
        "physics": "트럭과 오토바이는 바퀴로 흙길에 지지되어 있고, 탑승자들은 좌석에 앉아 있다. 먼지는 바퀴 부근에서 뒤쪽으로 퍼져 주행 동작과 부합한다. 지지 없이 떠 있는 차체나 인물은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 트럭은 왼쪽 위로 굽은 도로를 따라, 오른쪽 트럭은 오른쪽 위 도로를 따라 멀어진다. 두 트럭의 후면과 측면이 보인다. 오른쪽 갈래의 승용차, 오토바이, 먼 트럭은 모두 주된 오른쪽 트럭보다 앞에 있으므로 요구된 후속 차량 배치가 아니다. 오토바이 탑승자는 진행 방향을 향한다.",
        "built_space": "아래 중앙에서 들어오는 흙길 하나가 수풀을 사이에 두고 두 갈래로 나뉜다. 주된 트럭 두 대는 각 갈래에 있지만 화면 상단보다는 중앙 높이에 크게 놓인다. 오른쪽 갈래에는 승용차 한 대, 오토바이 한 대, 추가 트럭 한 대가 보인다. 배경에는 수풀, 산, 작은 건물과 전신주가 있으며 반사나 고정 설비의 중복 문제는 없다.",
        "entities": "낡은 주 트럭 두 대와 갈림길, 바퀴 뒤 흙먼지가 보인다. 적재함은 열려 있지만 운전석에는 지붕 외피가 남아 보여 '지붕 없는 트럭'과 맞지 않는다. 왼쪽 적재함에는 천과 여러 짐이 실려 있으나 찰리나 B-200으로 식별할 수 없다. 오른쪽 오토바이에는 최소 한 명의 사람이 명확히 보이며, 그 사람의 국적·나이·성별은 판독할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
        ],
        "physics": "차량들은 바퀴로 노면에 지지되고, 오토바이 탑승자는 좌석에 앉아 있다. 왼쪽 트럭의 짐과 천은 적재함 바닥과 난간에 놓여 지지된다. 바퀴 뒤 먼지와 진행 방향은 자연스럽고, 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두 트럭의 좌우 상단 배치와 갈림길 풀샷은 더 정확하지만, 금지된 사람들이 등장하고 오른쪽 추가 차량들이 수빈의 트럭을 뒤따르지 않아 부적합하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "갈라지는 주행과 흙먼지는 구현했지만, 사람 등장 금지를 위반하고 트럭들이 화면 중앙에 내려와 있으며 추가 차량도 선행한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 트럭은 왼쪽 위 갈래로, 오른쪽 트럭은 오른쪽 위 갈래로 카메라에서 멀어진다. 두 차량의 후면이 보이고 왼쪽 트럭의 측면도 드러난다. 오른쪽 도로의 추가 트럭들과 오토바이들은 같은 방향으로 향하지만, 주된 오른쪽 트럭보다 앞에 있어 '뒤따르는 차량' 관계와 반대다. 인물들의 정확한 시선은 판독하기 어렵다.",
        "built_space": "아래 중앙의 공통 흙길 하나가 좌우 두 갈래로 나뉘며, 두 주된 트럭은 각각 화면 왼쪽 위와 오른쪽 위의 중경에 놓인다. 오른쪽 갈래에 추가 트럭 두 대와 오토바이 세 대가 보인다. 주변은 경작지와 낮은 산이며 실내 설비나 반사는 없다. 낮의 농촌 갈림길이라는 공간 조건과 와이드 구도는 잘 읽힌다.",
        "entities": "낡고 녹슨 화물 트럭 두 대, 양 갈래 흙길, 바퀴 뒤 흙먼지가 보인다. 적재함은 개방되어 있고 운전석 상부는 프레임이 두드러지지만 지붕 제거 여부는 완전히 명확하지 않다. 오토바이 탑승자 세 명과 오른쪽 주 트럭 운전석의 사람 형상이 보여 무인 화면 조건에 어긋난다. 인물의 국적·나이·성별은 이 크기에서 확인할 수 없다. 찰리의 손상 상태와 B-200 부품은 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
        ],
        "physics": "트럭과 오토바이는 바퀴로 흙길에 지지되어 있고, 탑승자들은 좌석에 앉아 있다. 먼지는 바퀴 부근에서 뒤쪽으로 퍼져 주행 동작과 부합한다. 지지 없이 떠 있는 차체나 인물은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 트럭은 왼쪽 위로 굽은 도로를 따라, 오른쪽 트럭은 오른쪽 위 도로를 따라 멀어진다. 두 트럭의 후면과 측면이 보인다. 오른쪽 갈래의 승용차, 오토바이, 먼 트럭은 모두 주된 오른쪽 트럭보다 앞에 있으므로 요구된 후속 차량 배치가 아니다. 오토바이 탑승자는 진행 방향을 향한다.",
        "built_space": "아래 중앙에서 들어오는 흙길 하나가 수풀을 사이에 두고 두 갈래로 나뉜다. 주된 트럭 두 대는 각 갈래에 있지만 화면 상단보다는 중앙 높이에 크게 놓인다. 오른쪽 갈래에는 승용차 한 대, 오토바이 한 대, 추가 트럭 한 대가 보인다. 배경에는 수풀, 산, 작은 건물과 전신주가 있으며 반사나 고정 설비의 중복 문제는 없다.",
        "entities": "낡은 주 트럭 두 대와 갈림길, 바퀴 뒤 흙먼지가 보인다. 적재함은 열려 있지만 운전석에는 지붕 외피가 남아 보여 '지붕 없는 트럭'과 맞지 않는다. 왼쪽 적재함에는 천과 여러 짐이 실려 있으나 찰리나 B-200으로 식별할 수 없다. 오른쪽 오토바이에는 최소 한 명의 사람이 명확히 보이며, 그 사람의 국적·나이·성별은 판독할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
        ],
        "physics": "차량들은 바퀴로 노면에 지지되고, 오토바이 탑승자는 좌석에 앉아 있다. 왼쪽 트럭의 짐과 천은 적재함 바닥과 난간에 놓여 지지된다. 바퀴 뒤 먼지와 진행 방향은 자연스럽고, 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
    ],
    "B": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "사람을 포함하지 말라는 지침을 어기고 오토바이 탑승자를 생성했으며, 짐칸에 명시된 찰리(Charlie)가 누락되어 치명적 감점 요소가 됨.  ★위반: [gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함 / [gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "사람을 묘사하지 말라는 명시적 지침을 위반하고 다수의 오토바이 탑승자를 화면에 포함시켰으며, 짐칸이 비어 있어 실패함.  ★위반: [gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함 / [gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-6e38-739c-a0ce-d5aa8bc177b5",
  "ref_mode": "lane(map_marker): 스케치",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S68sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:13:06.518654+00:00",
  "fingerprint": "0214ba487f2d4ea1b6711587d3826031db40e3a15cf7c280b253d8e87b068db3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S68sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S68sh3_sel.png",
  "source_sha256": "f7f360c88b06c2faaba0f604c26c7f8e244f6e5dc1c7e6b172d44e8ce62639be",
  "file": "S68sh3_cine.png",
  "staged_sha256": "48556c0842a3d5a93b9b4c69c15c26b18f13a740db6cb6f092f0a59119e1753b",
  "latency_ms": 10034
 },
 "S68sh7::signage": {
  "fp": "28960e5541b3014d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S68sh7": {
  "input_fingerprint": "4d2843e5d7243844",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie sits in the roofless truck's cargo bed with the recovered B-200 chest component and his unrepaired battle damage. A baby bird formerly sheltered in the warehouse now perches on his finger.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie sits in the roofless truck's cargo bed with the recovered B-200 chest component and his unrepaired battle damage. A baby bird formerly sheltered in the warehouse now perches on his finger.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie sits in the roofless truck's cargo bed with the recovered B-200 chest component and his unrepaired battle damage. A baby bird formerly sheltered in the warehouse now perches on his finger.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S68sh7__bgfirst_bg.png",
     "asset_id": "16f6020b-ef17-4edc-b91e-0f09780d6056",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S68sh7.png",
     "asset_id": "aa8d3a60-5597-4fad-886f-90bc9f48a604",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:833571>",
     "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_bed_9f9871.png",
     "asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:833571>",
     "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "고개를 숙이고 지그시 감은 눈이 손가락 위에 있는 아기 새를 향하고 있음.",
    "built_space": "트럭 적재함 측면이 왼쪽에 보이고, 그 뒤로 참조 사진과 동일한 시골 도로와 산이 넓게 펼쳐져 있음.",
    "entities": "찰리는 낡은 금속 마스크 얼굴을 가졌으나 참조의 점/선 패턴과 의상(셔츠, 모자)이 누락됨. 아기 새는 참조와 일치함.",
    "hard_violations": [],
    "physics": "아기 새가 찰리의 금속 손가락 위에 발을 딛고 안정적으로 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "고개를 숙이고 감은 눈이 손가락 위의 아기 새를 향하고 있음.",
    "built_space": "낡은 트럭 적재함 내부를 배경으로 하나, 지정된 시골 도로의 모습은 명확하게 보이지 않음.",
    "entities": "찰리의 얼굴이 해골 형태로 완전히 변형되었고 의상이 누락됨. 아기 새는 참조와 일치함.",
    "hard_violations": [
     "[gpt-high] 적재함 외벽과 후미등이 찰리 뒤에 놓여 찰리가 적재함 바깥에 있는 배치로 보이며, 적재함 안에 앉아 이동한다는 명시적 위치를 위반한다."
    ],
    "physics": "아기 새가 손가락 위에 앉아 있으나 발끝이 금속 표면에서 약간 떠 있는 것처럼 보임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "참조 이미지의 시골 도로 배경을 매우 정확하게 구현하였으나, 지정된 프레이밍보다 배경이 넓게 보이고 캐릭터의 의상이 누락되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 얼굴이 해골 형태로 크게 왜곡되어 참조 이미지의 정체성을 잃었으며, 지정된 도로 배경이 제대로 나타나지 않았습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "고개를 숙이고 지그시 감은 눈이 손가락 위에 있는 아기 새를 향하고 있음.",
        "built_space": "트럭 적재함 측면이 왼쪽에 보이고, 그 뒤로 참조 사진과 동일한 시골 도로와 산이 넓게 펼쳐져 있음.",
        "entities": "찰리는 낡은 금속 마스크 얼굴을 가졌으나 참조의 점/선 패턴과 의상(셔츠, 모자)이 누락됨. 아기 새는 참조와 일치함.",
        "hard_violations": [],
        "physics": "아기 새가 찰리의 금속 손가락 위에 발을 딛고 안정적으로 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "고개를 숙이고 감은 눈이 손가락 위의 아기 새를 향하고 있음.",
        "built_space": "낡은 트럭 적재함 내부를 배경으로 하나, 지정된 시골 도로의 모습은 명확하게 보이지 않음.",
        "entities": "찰리의 얼굴이 해골 형태로 완전히 변형되었고 의상이 누락됨. 아기 새는 참조와 일치함.",
        "hard_violations": [],
        "physics": "아기 새가 손가락 위에 앉아 있으나 발끝이 금속 표면에서 약간 떠 있는 것처럼 보임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "참조 이미지의 시골 도로 배경을 매우 정확하게 구현하였으나, 지정된 프레이밍보다 배경이 넓게 보이고 캐릭터의 의상이 누락되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 얼굴이 해골 형태로 크게 왜곡되어 참조 이미지의 정체성을 잃었으며, 지정된 도로 배경이 제대로 나타나지 않았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "고개를 숙이고 지그시 감은 눈이 손가락 위에 있는 아기 새를 향하고 있음.",
        "built_space": "트럭 적재함 측면이 왼쪽에 보이고, 그 뒤로 참조 사진과 동일한 시골 도로와 산이 넓게 펼쳐져 있음.",
        "entities": "찰리는 낡은 금속 마스크 얼굴을 가졌으나 참조의 점/선 패턴과 의상(셔츠, 모자)이 누락됨. 아기 새는 참조와 일치함.",
        "hard_violations": [],
        "physics": "아기 새가 찰리의 금속 손가락 위에 발을 딛고 안정적으로 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "고개를 숙이고 감은 눈이 손가락 위의 아기 새를 향하고 있음.",
        "built_space": "낡은 트럭 적재함 내부를 배경으로 하나, 지정된 시골 도로의 모습은 명확하게 보이지 않음.",
        "entities": "찰리의 얼굴이 해골 형태로 완전히 변형되었고 의상이 누락됨. 아기 새는 참조와 일치함.",
        "hard_violations": [],
        "physics": "아기 새가 손가락 위에 앉아 있으나 발끝이 금속 표면에서 약간 떠 있는 것처럼 보임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "눈을 감고 손가락에 새를 올린 순간은 구현했지만, 찰리가 적재함 바깥에 놓인 공간 배치와 해골형으로 바뀐 얼굴이 핵심 지시를 어긴다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "적재함 안에서 새를 향해 고개를 숙이고 눈을 감은 낡은 금속 얼굴의 클로즈업이 더 충실하나, 적재함과 풍경의 노출은 요구보다 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 눈을 감고 얼굴을 약간 아래로 기울였으며, 새는 얼굴의 왼쪽 아래 손가락에 있다. 다만 얼굴은 새 쪽으로 충분히 돌아가기보다 정면을 향한다. 새의 부리는 오른쪽의 찰리 얼굴을 향한다.",
        "built_space": "찰리 뒤로 가까운 가로 철판 한 면, 그 너머 적재함 바닥과 맞은편 벽 한 면, 왼쪽 측벽 한 면이 보인다. 가까운 철판 아래 왼쪽에는 후미등도 보여, 카메라와 찰리는 적재함 외벽 바깥쪽에 놓인 것으로 읽힌다. 이는 적재함 안에 앉아 있다는 배치와 맞지 않는다. 적재함도 어깨 아래의 좁고 비스듬한 일부가 아니라 배경 대부분을 차지한다.",
        "entities": "찰리 한 개체와 작은 새 한 마리가 보인다. 찰리의 베이지 금속 장갑과 긁힘은 맞지만, 얼굴은 참고의 흰 각진 마스크 대신 치열과 큰 비강 구멍이 있는 해골형이며 노란 발광선도 추가되었다. 참고의 밀짚모자와 화려한 목 주변 천은 보이지 않는다. 새는 작은 부리, 접힌 날개, 가느다란 발을 갖췄으나 참고보다 몸통이 길고 덜 둥글다. 읽을 수 있는 글자는 없다. 회수한 가슴 부품은 확인되지 않으며 하체는 프레임 밖이다.",
        "hard_violations": [
         "적재함 외벽과 후미등이 찰리 뒤에 놓여 찰리가 적재함 바깥에 있는 배치로 보이며, 적재함 안에 앉아 이동한다는 명시적 위치를 위반한다."
        ],
        "physics": "새의 발가락이 찰리의 펼친 금속 손가락에 접촉해 체중을 지탱한다. 손은 손목과 팔로 연결되고 머리는 기계식 목이 받친다. 공중에 떠 있는 개체는 없다. 골반과 좌면은 잘려 있어 앉은 자세의 지지는 확인할 수 없으며, 보이지 않는 지지를 부유로 판단하지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 눈을 감은 채 고개를 아래 왼쪽으로 숙여 손가락 위 새를 향한다. 새는 부리를 아래 오른쪽으로 기울여 자신의 발과 찰리의 손가락 쪽을 보고 있다. 새를 내려다보다 눈을 감는 순간의 방향 관계가 자연스럽다.",
        "built_space": "찰리 뒤로 적재함의 끝벽 한 면, 왼쪽 측벽 한 면과 바닥이 연결되어 보이고, 찰리는 그 내부에 놓인다. 녹슨 밝은 회색 철판과 바닥, 뒤로 이어지는 농촌 도로·전봇대·산은 장소 참고와 부합한다. 지붕 없이 하늘이 열려 있다. 다만 적재함의 두 벽과 바닥까지 상당 부분 드러나므로 어깨 아래 좁은 사선 조각만 남기라는 배경 제한에는 미달한다.",
        "entities": "찰리 한 개체와 아기 새 한 마리가 보인다. 찰리는 닳고 긁힌 베이지 장갑과 밝은 각진 마스크 얼굴을 지녀 A보다 참고 정체성에 가깝지만, 얼굴 판금의 세부 형태는 다르고 참고의 모자와 화려한 천은 없다. 새는 둥근 솜털 몸통, 작은 부리, 짧게 접힌 날개와 가는 발을 갖춰 참고와 잘 맞는다. 읽을 수 있는 글자나 추가 인물은 없다. 회수한 가슴 부품과 하체의 상태는 이 크롭으로 확정할 수 없다.",
        "hard_violations": [],
        "physics": "새의 두 발이 수평으로 뻗은 금속 손가락을 딛고 있으며 발톱도 표면에 걸려 있다. 손가락은 손과 팔로 연결되고 숙인 머리는 목의 기계 구조로 지지된다. 몸통은 적재함 안에 있지만 골반과 좌면은 화면 밖이어서 착석 접촉 자체는 확인되지 않는다. 보이는 범위에 지지 없는 부유나 불가능한 관절 배치는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "눈을 감고 손가락에 새를 올린 순간은 구현했지만, 찰리가 적재함 바깥에 놓인 공간 배치와 해골형으로 바뀐 얼굴이 핵심 지시를 어긴다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "적재함 안에서 새를 향해 고개를 숙이고 눈을 감은 낡은 금속 얼굴의 클로즈업이 더 충실하나, 적재함과 풍경의 노출은 요구보다 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 눈을 감고 얼굴을 약간 아래로 기울였으며, 새는 얼굴의 왼쪽 아래 손가락에 있다. 다만 얼굴은 새 쪽으로 충분히 돌아가기보다 정면을 향한다. 새의 부리는 오른쪽의 찰리 얼굴을 향한다.",
        "built_space": "찰리 뒤로 가까운 가로 철판 한 면, 그 너머 적재함 바닥과 맞은편 벽 한 면, 왼쪽 측벽 한 면이 보인다. 가까운 철판 아래 왼쪽에는 후미등도 보여, 카메라와 찰리는 적재함 외벽 바깥쪽에 놓인 것으로 읽힌다. 이는 적재함 안에 앉아 있다는 배치와 맞지 않는다. 적재함도 어깨 아래의 좁고 비스듬한 일부가 아니라 배경 대부분을 차지한다.",
        "entities": "찰리 한 개체와 작은 새 한 마리가 보인다. 찰리의 베이지 금속 장갑과 긁힘은 맞지만, 얼굴은 참고의 흰 각진 마스크 대신 치열과 큰 비강 구멍이 있는 해골형이며 노란 발광선도 추가되었다. 참고의 밀짚모자와 화려한 목 주변 천은 보이지 않는다. 새는 작은 부리, 접힌 날개, 가느다란 발을 갖췄으나 참고보다 몸통이 길고 덜 둥글다. 읽을 수 있는 글자는 없다. 회수한 가슴 부품은 확인되지 않으며 하체는 프레임 밖이다.",
        "hard_violations": [
         "적재함 외벽과 후미등이 찰리 뒤에 놓여 찰리가 적재함 바깥에 있는 배치로 보이며, 적재함 안에 앉아 이동한다는 명시적 위치를 위반한다."
        ],
        "physics": "새의 발가락이 찰리의 펼친 금속 손가락에 접촉해 체중을 지탱한다. 손은 손목과 팔로 연결되고 머리는 기계식 목이 받친다. 공중에 떠 있는 개체는 없다. 골반과 좌면은 잘려 있어 앉은 자세의 지지는 확인할 수 없으며, 보이지 않는 지지를 부유로 판단하지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 눈을 감은 채 고개를 아래 왼쪽으로 숙여 손가락 위 새를 향한다. 새는 부리를 아래 오른쪽으로 기울여 자신의 발과 찰리의 손가락 쪽을 보고 있다. 새를 내려다보다 눈을 감는 순간의 방향 관계가 자연스럽다.",
        "built_space": "찰리 뒤로 적재함의 끝벽 한 면, 왼쪽 측벽 한 면과 바닥이 연결되어 보이고, 찰리는 그 내부에 놓인다. 녹슨 밝은 회색 철판과 바닥, 뒤로 이어지는 농촌 도로·전봇대·산은 장소 참고와 부합한다. 지붕 없이 하늘이 열려 있다. 다만 적재함의 두 벽과 바닥까지 상당 부분 드러나므로 어깨 아래 좁은 사선 조각만 남기라는 배경 제한에는 미달한다.",
        "entities": "찰리 한 개체와 아기 새 한 마리가 보인다. 찰리는 닳고 긁힌 베이지 장갑과 밝은 각진 마스크 얼굴을 지녀 A보다 참고 정체성에 가깝지만, 얼굴 판금의 세부 형태는 다르고 참고의 모자와 화려한 천은 없다. 새는 둥근 솜털 몸통, 작은 부리, 짧게 접힌 날개와 가는 발을 갖춰 참고와 잘 맞는다. 읽을 수 있는 글자나 추가 인물은 없다. 회수한 가슴 부품과 하체의 상태는 이 크롭으로 확정할 수 없다.",
        "hard_violations": [],
        "physics": "새의 두 발이 수평으로 뻗은 금속 손가락을 딛고 있으며 발톱도 표면에 걸려 있다. 손가락은 손과 팔로 연결되고 숙인 머리는 목의 기계 구조로 지지된다. 몸통은 적재함 안에 있지만 골반과 좌면은 화면 밖이어서 착석 접촉 자체는 확인되지 않는다. 보이는 범위에 지지 없는 부유나 불가능한 관절 배치는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.042
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.792
   },
   "violations": {
    "B": [
     "[gpt-high] 적재함 외벽과 후미등이 찰리 뒤에 놓여 찰리가 적재함 바깥에 있는 배치로 보이며, 적재함 안에 앉아 이동한다는 명시적 위치를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 792
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "참조 이미지의 시골 도로 배경을 매우 정확하게 구현하였으나, 지정된 프레이밍보다 배경이 넓게 보이고 캐릭터의 의상이 누락되었습니다."
   },
   {
    "label": "B",
    "score": 792,
    "verdict_ko": "찰리의 얼굴이 해골 형태로 크게 왜곡되어 참조 이미지의 정체성을 잃었으며, 지정된 도로 배경이 제대로 나타나지 않았습니다.  ★위반: [gpt-high] 적재함 외벽과 후미등이 찰리 뒤에 놓여 찰리가 적재함 바깥에 있는 배치로 보이며, 적재함 안에 앉아 이동한다는 명시적 위치를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_bed_9f9871.png",
    "asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:833571>",
    "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-6fd8-7e00-a453-084bbc22c0c8",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S68sh7__bgfirst_bg.png",
   "bg_asset_id": "16f6020b-ef17-4edc-b91e-0f09780d6056",
   "bg_record_key": "S68sh7::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "open_truck_bed",
   "groupbg_asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S68sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:15:32.992163+00:00",
  "fingerprint": "7157e96bff0498bf533c3de15dfcfd4c2ae88ae59c4e38e379aa5b6b60c6868e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S68sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S68sh7_sel.png",
  "source_sha256": "e3d2333653aa94441252242310398c76bd10b2f2c6fd6ebd7fc4f0be12573fea",
  "file": "S68sh7_cine.png",
  "staged_sha256": "d1bcba7c207704564dc626929b257facbf972c89bd1a56633d25a6dfef3e6649",
  "latency_ms": 10793
 },
 "S69sh3::signage": {
  "fp": "12dedf02e3cfd39a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::b540692c4ba06c93": {
  "subjects": [],
  "subject_text": "현우의 트럭 적재함\n지붕 없이 위로 열린 낡은 적재 공간. 낮은 측벽이 바닥을 둘러싸고, 한쪽에는 옮겨 펼칠 수 있는 그늘막이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L233",
  "scope_role": "location_exterior",
  "scope_sha": "df9b8f1b489c5ca2"
 },
 "S69sh3::bgfirst_bg": {
  "input_fingerprint": "20bd23281715627f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S69sh3__bgfirst_bg.png",
  "asset_id": "ccf56b84-5714-4267-b1b6-cfb4a6ff27b2",
  "input_asset_ids": [
   "acde8000-023e-40d0-b3ff-468a1735a66a",
   "c94c53e3-76a2-4f78-b4ae-95204318d3c2"
  ]
 },
 "S69sh3": {
  "input_fingerprint": "d2c3f1297ec978b1",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is asleep in the truck's rear cargo bed with his head grouped closely beside Amber's and Raul's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is asleep in the truck's rear cargo bed with her head grouped closely beside Hyunwoo's and Raul's heads beneath the shade Charlie provides. The source does not specify her torso's orientation or the positions of her arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Raul is asleep in the truck's rear cargo bed with his head grouped closely beside Hyunwoo's and Amber's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The roofless truck is parked beside rocks for shade, with a movable shade cloth over the cargo-bed resting area. Charlie adjusts the cloth while retaining his unrepaired battle damage and the recovered B-200 chest component in his possession. 현우: He is asleep in the rear cargo bed, still bearing his recent injuries. 앰버: She is asleep in the rear cargo bed; no treatment of her recent head injury has been established. 라울: He is asleep in the rear cargo bed.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is asleep in the truck's rear cargo bed with his head grouped closely beside Amber's and Raul's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is asleep in the truck's rear cargo bed with her head grouped closely beside Hyunwoo's and Raul's heads beneath the shade Charlie provides. The source does not specify her torso's orientation or the positions of her arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Raul is asleep in the truck's rear cargo bed with his head grouped closely beside Hyunwoo's and Amber's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The roofless truck is parked beside rocks for shade, with a movable shade cloth over the cargo-bed resting area. Charlie adjusts the cloth while retaining his unrepaired battle damage and the recovered B-200 chest component in his possession. 현우: He is asleep in the rear cargo bed, still bearing his recent injuries. 앰버: She is asleep in the rear cargo bed; no treatment of her recent head injury has been established. 라울: He is asleep in the rear cargo bed.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is asleep in the truck's rear cargo bed with his head grouped closely beside Amber's and Raul's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is asleep in the truck's rear cargo bed with her head grouped closely beside Hyunwoo's and Raul's heads beneath the shade Charlie provides. The source does not specify her torso's orientation or the positions of her arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Raul is asleep in the truck's rear cargo bed with his head grouped closely beside Hyunwoo's and Amber's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The roofless truck is parked beside rocks for shade, with a movable shade cloth over the cargo-bed resting area. Charlie adjusts the cloth while retaining his unrepaired battle damage and the recovered B-200 chest component in his possession. 현우: He is asleep in the rear cargo bed, still bearing his recent injuries. 앰버: She is asleep in the rear cargo bed; no treatment of her recent head injury has been established. 라울: He is asleep in the rear cargo bed.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S69sh3__bgfirst_bg.png",
     "asset_id": "ccf56b84-5714-4267-b1b6-cfb4a6ff27b2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S69sh3.png",
     "asset_id": "acde8000-023e-40d0-b3ff-468a1735a66a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_open_roof_truck_sel.png",
     "asset_id": "c94c53e3-76a2-4f78-b4ae-95204318d3c2",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:889172>",
     "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 시선을 아래로 향한 채 양팔을 벌려 천을 펼치고 있음.",
    "built_space": "트럭 적재함 내부에 일행이 누워 있고, 뒤로는 일반적인 바위 절벽이 보임.",
    "entities": "현우와 라울이 잠들어 있으나, 앰버는 복장 레퍼런스와 달리 두꺼운 붕대를 감고 있고 찰리는 데미지 없이 깨끗한 상태임.",
    "hard_violations": [
     "[gpt-high] 머리 부상에 대한 치료가 설정되지 않았다고 명시된 앰버에게 흰 의료용 머리 붕대를 새로 추가해, 설정에 없는 치료 상태를 만들었다."
    ],
    "physics": "천은 찰리의 양손에 의해 지탱되고, 일행은 바닥에 엎드리거나 누워 중력에 맞게 위치함."
   },
   {
    "label": "B",
    "direction": "찰리의 얼굴과 시선은 아래쪽의 잠든 아이들을 향하고 있음.",
    "built_space": "트럭 적재함 안쪽에 세 명이 누워 있으며, 배경의 바위 틈새가 로케이션 레퍼런스의 구조와 일치함.",
    "entities": "현우, 앰버(머리 위 방독면과 멜빵바지 착용), 라울이 모여 잠들어 있으며 찰리는 가슴의 원형 부품과 그을린 전투 데미지를 지니고 있음.",
    "hard_violations": [],
    "physics": "찰리의 양손이 천을 들어 올리고 있으며, 잠든 인물들은 적재함 바닥에 안정적으로 누워 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "로케이션 레퍼런스의 바위 형태를 정확히 반영했으며, 앰버의 마스크와 복장, 찰리의 전투 데미지 및 가슴 부품 설정을 충실히 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "치료받지 않았다는 설정과 달리 앰버의 머리에 붕대가 있고 레퍼런스 복장이 누락되었으며, 찰리의 파손 흔적도 보이지 않음."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 시선은 아래쪽의 잠든 아이들을 향하고 있음.",
        "built_space": "트럭 적재함 안쪽에 세 명이 누워 있으며, 배경의 바위 틈새가 로케이션 레퍼런스의 구조와 일치함.",
        "entities": "현우, 앰버(머리 위 방독면과 멜빵바지 착용), 라울이 모여 잠들어 있으며 찰리는 가슴의 원형 부품과 그을린 전투 데미지를 지니고 있음.",
        "hard_violations": [],
        "physics": "찰리의 양손이 천을 들어 올리고 있으며, 잠든 인물들은 적재함 바닥에 안정적으로 누워 있음."
       },
       {
        "label": "A",
        "direction": "찰리는 시선을 아래로 향한 채 양팔을 벌려 천을 펼치고 있음.",
        "built_space": "트럭 적재함 내부에 일행이 누워 있고, 뒤로는 일반적인 바위 절벽이 보임.",
        "entities": "현우와 라울이 잠들어 있으나, 앰버는 복장 레퍼런스와 달리 두꺼운 붕대를 감고 있고 찰리는 데미지 없이 깨끗한 상태임.",
        "hard_violations": [],
        "physics": "천은 찰리의 양손에 의해 지탱되고, 일행은 바닥에 엎드리거나 누워 중력에 맞게 위치함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "로케이션 레퍼런스의 바위 형태를 정확히 반영했으며, 앰버의 마스크와 복장, 찰리의 전투 데미지 및 가슴 부품 설정을 충실히 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "치료받지 않았다는 설정과 달리 앰버의 머리에 붕대가 있고 레퍼런스 복장이 누락되었으며, 찰리의 파손 흔적도 보이지 않음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 시선은 아래쪽의 잠든 아이들을 향하고 있음.",
        "built_space": "트럭 적재함 안쪽에 세 명이 누워 있으며, 배경의 바위 틈새가 로케이션 레퍼런스의 구조와 일치함.",
        "entities": "현우, 앰버(머리 위 방독면과 멜빵바지 착용), 라울이 모여 잠들어 있으며 찰리는 가슴의 원형 부품과 그을린 전투 데미지를 지니고 있음.",
        "hard_violations": [],
        "physics": "찰리의 양손이 천을 들어 올리고 있으며, 잠든 인물들은 적재함 바닥에 안정적으로 누워 있음."
       },
       {
        "label": "A",
        "direction": "찰리는 시선을 아래로 향한 채 양팔을 벌려 천을 펼치고 있음.",
        "built_space": "트럭 적재함 내부에 일행이 누워 있고, 뒤로는 일반적인 바위 절벽이 보임.",
        "entities": "현우와 라울이 잠들어 있으나, 앰버는 복장 레퍼런스와 달리 두꺼운 붕대를 감고 있고 찰리는 데미지 없이 깨끗한 상태임.",
        "hard_violations": [],
        "physics": "천은 찰리의 양손에 의해 지탱되고, 일행은 바닥에 엎드리거나 누워 중력에 맞게 위치함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "양손으로 천을 펼치는 찰리의 상체와 얼굴, 그 아래 잠든 세 사람 및 기준 바위 지형을 충실히 담았지만, 현우의 머리가 조금 떨어져 있고 찰리의 기준 의상이 빠졌다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "찰리의 역동적인 상체와 서로 가까운 세 머리는 잘 구현했지만, 치료가 설정되지 않은 앰버에게 머리 붕대를 추가해 상태 연속성을 위반했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 아래쪽의 잠든 일행을 향한다. 양팔은 좌우 위로 벌어져 각각 같은 천의 가장자리를 붙잡고 있다. 천은 화면 상단을 비스듬히 가로지르며 세 사람의 얼굴 앞이 아니라 머리 위를 덮고, 찰리의 얼굴도 드러나 있다. 세 사람은 눈을 감고 있다.",
        "built_space": "트럭 한 대의 열린 적재함이 보이며, 왼쪽에 운전실 뒤창과 보호 프레임 하나, 아래쪽에 가까운 측벽 하나, 뒤쪽과 오른쪽에 적재함 경계가 보인다. 카메라는 적재함 바깥에서 측면을 바라본다. 찰리는 반대쪽 측벽 너머에 있고 세 사람은 적재함 안에 누워 있다. 배경의 왼쪽 큰 바위 면, 중앙의 좁은 틈과 걸친 바위, 오른쪽 둥근 바위가 장소 기준과 가깝다. 현우의 머리는 서로 붙어 있는 앰버·라울의 머리에서 다소 떨어져 있다.",
        "entities": "찰리 한 명과 잠든 세 사람만 보인다. 찰리의 베이지 장갑판, 육중한 팔, 흰 기계식 얼굴과 전투 손상은 부합하지만 기준의 밀짚모자와 화려한 겉옷은 없다. 별도로 회수한 흉부 부품은 식별되지 않는다. 현우는 검은 머리의 앳된 동아시아계 남성으로 회색 셔츠와 얼굴 상처가 보인다. 앰버는 금발의 어린 여자아이로 고글, 남색 상의와 갈색 작업복이 기준에 가깝고 머리 붕대는 없다. 라울은 어두운 피부와 뒤로 묶은 머리의 어린아이로 회색 상의를 입었으나 얼굴 상당 부분이 가려져 세부 동일성 확인은 제한된다. 천과 낡은 금속 적재함은 실제 재질로 보이며, 작은 차량 표찰의 문구는 판독되지 않는다.",
        "hard_violations": [],
        "physics": "천은 찰리의 두 손과 왼쪽 차량 프레임의 묶인 지점에 지지되며, 사이 부분은 중력에 따라 처진다. 현우의 머리는 접은 손과 천 더미에, 앰버의 머리는 적재함의 깔개와 팔 부근에, 라울의 머리는 앰버 뒤쪽의 받침에 기대 있다. 잠든 몸과 팔은 적재함 바닥 및 깔개에 놓여 있고 공중에 들린 팔다리는 없다. 찰리의 하체 지지점은 측벽에 가려져 있지만 상체가 떠 있다고 볼 근거는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 얼굴을 아래 왼쪽의 잠든 일행 쪽으로 기울이고 양손을 좌우 위로 벌려 천을 잡고 있다. 하나의 천이 두 손을 연결하며 화면 상단에서 일행 위로 펼쳐지고, 찰리와 세 사람의 얼굴을 가리지 않는다. 세 사람 모두 눈을 감고 머리를 서로 가까이 두었다.",
        "built_space": "트럭 한 대의 열린 적재함을 외부에서 사선으로 내려다본다. 오른쪽 뒤에 운전실과 뒤창 하나, 오른쪽 끝에 측면 거울 하나가 보이고, 적재함의 가까운 측벽과 맞은편 측벽이 공간을 구획한다. 세 사람은 적재함 안에 누워 있고 찰리는 맞은편 측벽 너머에 있다. 배경에는 기준과 유사한 좁은 바위 틈과 걸친 바위가 있으나, 오른쪽으로 하늘과 먼 산이 열려 기준보다 주변 경관의 비중이 크다.",
        "entities": "찰리와 현우·앰버·라울에 해당하는 네 인물만 보인다. 찰리의 베이지 장갑과 흰 얼굴, 긁힌 표면은 부합하지만 모자와 화려한 겉옷이 없다. 가슴에는 끈으로 고정된 별도 직사각형 부품이 있으나 지정된 회수 부품인지 외형만으로 확정할 수 없다. 현우의 검은 머리, 회색 셔츠와 얼굴 상처는 부합한다. 앰버는 금발의 어린 여자아이지만 기준의 고글 대신 흰 머리 붕대를 착용했다. 라울은 어두운 피부, 뒤로 묶은 머리와 어린 얼굴이 부합하며 회색 계열 옷을 입고 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "머리 부상에 대한 치료가 설정되지 않았다고 명시된 앰버에게 흰 의료용 머리 붕대를 새로 추가해, 설정에 없는 치료 상태를 만들었다."
        ],
        "physics": "천은 찰리의 두 손에 잡혀 있으며 손 사이가 완만하게 처지고 주름이 이어진다. 가슴의 별도 부품은 몸통을 두른 끈으로 지지된다. 세 사람의 머리와 몸통, 보이는 손과 팔은 깔개와 적재함 바닥에 충분히 받쳐져 있어 수면 자세가 자연스럽다. 찰리의 골반 아래는 적재함 벽에 가려져 지면 접촉을 확인할 수 없지만, 지지 없이 공중에 떠 있는 모습은 아니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "양손으로 천을 펼치는 찰리의 상체와 얼굴, 그 아래 잠든 세 사람 및 기준 바위 지형을 충실히 담았지만, 현우의 머리가 조금 떨어져 있고 찰리의 기준 의상이 빠졌다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "찰리의 역동적인 상체와 서로 가까운 세 머리는 잘 구현했지만, 치료가 설정되지 않은 앰버에게 머리 붕대를 추가해 상태 연속성을 위반했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 아래쪽의 잠든 일행을 향한다. 양팔은 좌우 위로 벌어져 각각 같은 천의 가장자리를 붙잡고 있다. 천은 화면 상단을 비스듬히 가로지르며 세 사람의 얼굴 앞이 아니라 머리 위를 덮고, 찰리의 얼굴도 드러나 있다. 세 사람은 눈을 감고 있다.",
        "built_space": "트럭 한 대의 열린 적재함이 보이며, 왼쪽에 운전실 뒤창과 보호 프레임 하나, 아래쪽에 가까운 측벽 하나, 뒤쪽과 오른쪽에 적재함 경계가 보인다. 카메라는 적재함 바깥에서 측면을 바라본다. 찰리는 반대쪽 측벽 너머에 있고 세 사람은 적재함 안에 누워 있다. 배경의 왼쪽 큰 바위 면, 중앙의 좁은 틈과 걸친 바위, 오른쪽 둥근 바위가 장소 기준과 가깝다. 현우의 머리는 서로 붙어 있는 앰버·라울의 머리에서 다소 떨어져 있다.",
        "entities": "찰리 한 명과 잠든 세 사람만 보인다. 찰리의 베이지 장갑판, 육중한 팔, 흰 기계식 얼굴과 전투 손상은 부합하지만 기준의 밀짚모자와 화려한 겉옷은 없다. 별도로 회수한 흉부 부품은 식별되지 않는다. 현우는 검은 머리의 앳된 동아시아계 남성으로 회색 셔츠와 얼굴 상처가 보인다. 앰버는 금발의 어린 여자아이로 고글, 남색 상의와 갈색 작업복이 기준에 가깝고 머리 붕대는 없다. 라울은 어두운 피부와 뒤로 묶은 머리의 어린아이로 회색 상의를 입었으나 얼굴 상당 부분이 가려져 세부 동일성 확인은 제한된다. 천과 낡은 금속 적재함은 실제 재질로 보이며, 작은 차량 표찰의 문구는 판독되지 않는다.",
        "hard_violations": [],
        "physics": "천은 찰리의 두 손과 왼쪽 차량 프레임의 묶인 지점에 지지되며, 사이 부분은 중력에 따라 처진다. 현우의 머리는 접은 손과 천 더미에, 앰버의 머리는 적재함의 깔개와 팔 부근에, 라울의 머리는 앰버 뒤쪽의 받침에 기대 있다. 잠든 몸과 팔은 적재함 바닥 및 깔개에 놓여 있고 공중에 들린 팔다리는 없다. 찰리의 하체 지지점은 측벽에 가려져 있지만 상체가 떠 있다고 볼 근거는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 얼굴을 아래 왼쪽의 잠든 일행 쪽으로 기울이고 양손을 좌우 위로 벌려 천을 잡고 있다. 하나의 천이 두 손을 연결하며 화면 상단에서 일행 위로 펼쳐지고, 찰리와 세 사람의 얼굴을 가리지 않는다. 세 사람 모두 눈을 감고 머리를 서로 가까이 두었다.",
        "built_space": "트럭 한 대의 열린 적재함을 외부에서 사선으로 내려다본다. 오른쪽 뒤에 운전실과 뒤창 하나, 오른쪽 끝에 측면 거울 하나가 보이고, 적재함의 가까운 측벽과 맞은편 측벽이 공간을 구획한다. 세 사람은 적재함 안에 누워 있고 찰리는 맞은편 측벽 너머에 있다. 배경에는 기준과 유사한 좁은 바위 틈과 걸친 바위가 있으나, 오른쪽으로 하늘과 먼 산이 열려 기준보다 주변 경관의 비중이 크다.",
        "entities": "찰리와 현우·앰버·라울에 해당하는 네 인물만 보인다. 찰리의 베이지 장갑과 흰 얼굴, 긁힌 표면은 부합하지만 모자와 화려한 겉옷이 없다. 가슴에는 끈으로 고정된 별도 직사각형 부품이 있으나 지정된 회수 부품인지 외형만으로 확정할 수 없다. 현우의 검은 머리, 회색 셔츠와 얼굴 상처는 부합한다. 앰버는 금발의 어린 여자아이지만 기준의 고글 대신 흰 머리 붕대를 착용했다. 라울은 어두운 피부, 뒤로 묶은 머리와 어린 얼굴이 부합하며 회색 계열 옷을 입고 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "머리 부상에 대한 치료가 설정되지 않았다고 명시된 앰버에게 흰 의료용 머리 붕대를 새로 추가해, 설정에 없는 치료 상태를 만들었다."
        ],
        "physics": "천은 찰리의 두 손에 잡혀 있으며 손 사이가 완만하게 처지고 주름이 이어진다. 가슴의 별도 부품은 몸통을 두른 끈으로 지지된다. 세 사람의 머리와 몸통, 보이는 손과 팔은 깔개와 적재함 바닥에 충분히 받쳐져 있어 수면 자세가 자연스럽다. 찰리의 골반 아래는 적재함 벽에 가려져 지면 접촉을 확인할 수 없지만, 지지 없이 공중에 떠 있는 모습은 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.196,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.946,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gpt-high] 머리 부상에 대한 치료가 설정되지 않았다고 명시된 앰버에게 흰 의료용 머리 붕대를 새로 추가해, 설정에 없는 치료 상태를 만들었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 946
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "로케이션 레퍼런스의 바위 형태를 정확히 반영했으며, 앰버의 마스크와 복장, 찰리의 전투 데미지 및 가슴 부품 설정을 충실히 구현함."
   },
   {
    "label": "A",
    "score": 946,
    "verdict_ko": "치료받지 않았다는 설정과 달리 앰버의 머리에 붕대가 있고 레퍼런스 복장이 누락되었으며, 찰리의 파손 흔적도 보이지 않음.  ★위반: [gpt-high] 머리 부상에 대한 치료가 설정되지 않았다고 명시된 앰버에게 흰 의료용 머리 붕대를 새로 추가해, 설정에 없는 치료 상태를 만들었다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_open_roof_truck_sel.png",
    "asset_id": "c94c53e3-76a2-4f78-b4ae-95204318d3c2",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:889172>",
    "asset_id": "b33e860b-3c16-469c-94fd-03efcf7a5c67",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-74ab-7f98-a2c8-074f5d8a513d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S69sh3__bgfirst_bg.png",
   "bg_asset_id": "ccf56b84-5714-4267-b1b6-cfb4a6ff27b2",
   "bg_record_key": "S69sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "seed_bg"
  },
  "ref_mode": "seed-bg+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S69sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:16:53.104228+00:00",
  "fingerprint": "db0f8ec0822854dfdefe163b0618ac63f5a66883bac99a2f93f2ce53031eb7fc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S69sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S69sh3_sel.png",
  "source_sha256": "66f536a1bcbf30e00e4150b0463ef540270f8bae83ab950fd4bbfa864ad99c9b",
  "file": "S69sh3_cine.png",
  "staged_sha256": "e9a3a1be9066e53afbf16c2ef6b0866fad77ae7aec3971da7557f509a81a3909",
  "latency_ms": 11786
 },
 "S70sh5::signage": {
  "fp": "6185ff30a3ef5644",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5210e323d6019fba": {
  "subjects": [],
  "subject_text": "군 병원 복도\n여러 병실의 출입문이 이어지는 실내 복도. 천장 조명 아래로 긴 통로가 뻗어 있고 병실 안쪽을 볼 수 있는 출입부가 있다.",
  "identity": "canonical",
  "scope_id": "L244",
  "scope_role": "location_interior",
  "scope_sha": "d5e2db056cd48c88"
 },
 "groupbg::military_hospital_room": {
  "input_fingerprint": "746b93870f51bf8b",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "military_hospital_room",
    "tags": [
     "S70sh5",
     "S70sh9"
    ]
   },
   "context_sig": "3be9ca213d18297f"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n군 병원 복도: 차갑고 통제된 군 의료 시설의 하얀 통로. (특징: 깨끗하지만 차가운 색감의 병원 복도; 제복을 입은 장교들의 걸음; 환자들이 있는 병실 문들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 병실에 누워있는 사람은... 한 쪽 눈, 팔과 다리 깁스를 하고 치료 중인 박철진이다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n군 병원 복도: 차갑고 통제된 군 의료 시설의 하얀 통로. (특징: 깨끗하지만 차가운 색감의 병원 복도; 제복을 입은 장교들의 걸음; 환자들이 있는 병실 문들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 병실에 누워있는 사람은... 한 쪽 눈, 팔과 다리 깁스를 하고 치료 중인 박철진이다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_military_hospital_room_03987e.png",
  "asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa",
  "input_asset_ids": [
   "e18af667-e876-4f68-a303-2e9caf00a411"
  ],
  "origin_tag": "S70sh5",
  "place_text": "Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.",
  "origin_inputs": {
   "place_text": "Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.",
   "time_of_day_en": "night",
   "conti_asset_id": "e18af667-e876-4f68-a303-2e9caf00a411"
  }
 },
 "S70sh5::bgfirst_bg": {
  "input_fingerprint": "a2a594571d2e0ef7",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5__bgfirst_bg.png",
  "asset_id": "a0d10e4e-ddb6-42d2-9c3e-f1455cd58d3a",
  "input_asset_ids": [
   "e18af667-e876-4f68-a303-2e9caf00a411",
   "83fc438f-e6e6-45b5-9493-51fe5f5a3faa"
  ]
 },
 "S70sh5": {
  "input_fingerprint": "2c678252156bee52",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night in a military hospital ward containing beds for wounded patients. 박철진: He is reclined on a hospital bed under treatment for severe injuries, with one eye covered and an arm and a leg immobilized in casts. He is only faintly conscious. 국방장관: He has reached the ward and stands looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night in a military hospital ward containing beds for wounded patients. 박철진: He is reclined on a hospital bed under treatment for severe injuries, with one eye covered and an arm and a leg immobilized in casts. He is only faintly conscious. 국방장관: He has reached the ward and stands looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night in a military hospital ward containing beds for wounded patients. 박철진: He is reclined on a hospital bed under treatment for severe injuries, with one eye covered and an arm and a leg immobilized in casts. He is only faintly conscious. 국방장관: He has reached the ward and stands looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5__bgfirst_bg.png",
     "asset_id": "a0d10e4e-ddb6-42d2-9c3e-f1455cd58d3a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S70sh5.png",
     "asset_id": "e18af667-e876-4f68-a303-2e9caf00a411",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1269105>",
     "asset_id": "ccc1d001-b248-4f0d-929f-9d10d042e801",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:835328>",
     "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_military_hospital_room_03987e.png",
     "asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1269105>",
     "asset_id": "ccc1d001-b248-4f0d-929f-9d10d042e801",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:835328>",
     "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "장관이 환자의 얼굴을 내려다보고, 환자는 장관을 올려다봄.",
    "built_space": "병실 내부. 전경 침대, IV 펌프, 뒤쪽의 커튼과 문, 배경 창문까지 모든 구조물이 올바르게 구성됨.",
    "entities": "장관(정장, 얼굴 특징 일치), 박철진(환자복, 한쪽 눈 안대, 오른팔과 오른다리 깁스, 얼굴 일치).",
    "hard_violations": [],
    "physics": "환자는 침대에 정상적으로 누워 있고, 장관은 서 있음. 불안정한 자세 없음."
   },
   {
    "label": "B",
    "direction": "장관의 시선은 환자의 얼굴을 향하며, 환자 역시 장관을 올려다봄.",
    "built_space": "병실 내부. 전경의 침대와 IV 폴, 배경의 커튼, 문, 창문이 참조 이미지와 동일한 위치와 비례로 배치됨.",
    "entities": "장관(정장, 얼굴 특징 일치), 박철진(한쪽 눈 안대, 오른팔과 왼발 깁스, 얼굴 일치).",
    "hard_violations": [],
    "physics": "환자는 침대에 누워 체중이 지탱되며, 장관은 바닥에 굳건히 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "공간과 인물 외형을 정확히 구현했으며, 특히 장관의 멸시하는 듯한 차가운 표정을 텍스트 지시대로 훌륭히 연출함."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "구도와 배경 요소는 잘 맞췄으나, 장관의 표정이 차갑거나 멸시하기보다는 걱정스러워 보여 핵심 감정 연출 지시를 놓침."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "장관의 시선은 환자의 얼굴을 향하며, 환자 역시 장관을 올려다봄.",
        "built_space": "병실 내부. 전경의 침대와 IV 폴, 배경의 커튼, 문, 창문이 참조 이미지와 동일한 위치와 비례로 배치됨.",
        "entities": "장관(정장, 얼굴 특징 일치), 박철진(한쪽 눈 안대, 오른팔과 왼발 깁스, 얼굴 일치).",
        "hard_violations": [],
        "physics": "환자는 침대에 누워 체중이 지탱되며, 장관은 바닥에 굳건히 서 있음."
       },
       {
        "label": "A",
        "direction": "장관이 환자의 얼굴을 내려다보고, 환자는 장관을 올려다봄.",
        "built_space": "병실 내부. 전경 침대, IV 펌프, 뒤쪽의 커튼과 문, 배경 창문까지 모든 구조물이 올바르게 구성됨.",
        "entities": "장관(정장, 얼굴 특징 일치), 박철진(환자복, 한쪽 눈 안대, 오른팔과 오른다리 깁스, 얼굴 일치).",
        "hard_violations": [],
        "physics": "환자는 침대에 정상적으로 누워 있고, 장관은 서 있음. 불안정한 자세 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "공간과 인물 외형을 정확히 구현했으며, 특히 장관의 멸시하는 듯한 차가운 표정을 텍스트 지시대로 훌륭히 연출함."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "구도와 배경 요소는 잘 맞췄으나, 장관의 표정이 차갑거나 멸시하기보다는 걱정스러워 보여 핵심 감정 연출 지시를 놓침."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "장관의 시선은 환자의 얼굴을 향하며, 환자 역시 장관을 올려다봄.",
        "built_space": "병실 내부. 전경의 침대와 IV 폴, 배경의 커튼, 문, 창문이 참조 이미지와 동일한 위치와 비례로 배치됨.",
        "entities": "장관(정장, 얼굴 특징 일치), 박철진(한쪽 눈 안대, 오른팔과 왼발 깁스, 얼굴 일치).",
        "hard_violations": [],
        "physics": "환자는 침대에 누워 체중이 지탱되며, 장관은 바닥에 굳건히 서 있음."
       },
       {
        "label": "A",
        "direction": "장관이 환자의 얼굴을 내려다보고, 환자는 장관을 올려다봄.",
        "built_space": "병실 내부. 전경 침대, IV 펌프, 뒤쪽의 커튼과 문, 배경 창문까지 모든 구조물이 올바르게 구성됨.",
        "entities": "장관(정장, 얼굴 특징 일치), 박철진(환자복, 한쪽 눈 안대, 오른팔과 오른다리 깁스, 얼굴 일치).",
        "hard_violations": [],
        "physics": "환자는 침대에 정상적으로 누워 있고, 장관은 서 있음. 불안정한 자세 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "내려다보는 시선과 야간 병실은 맞지만, 장관의 허벅지와 병상 대부분까지 담아 상체 중심 미디엄 숏보다 넓고 환자를 향한 멸시도 상대적으로 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "장관의 상체를 더 크게 두고 낮은 전경의 박철진을 차갑게 내려다보게 해 핵심 구도에 더 충실하지만, 환자 하반신까지 담은 범위는 다소 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장관은 고개와 눈을 화면 왼쪽 아래 박철진의 얼굴로 향한다. 박철진의 드러난 눈은 위쪽 장관 방향을 보고 있다. 장관의 찌푸린 표정은 냉담하지만 멸시보다는 심각한 관찰로도 읽힌다. 무기나 방향을 판정할 휴대 소품은 없다.",
        "built_space": "전경 병상 한 개와 뒤쪽 빈 병상 한 개가 보인다. 전경 왼쪽에는 벽 부착 의료 패널과 하부 조명 각 한 개, 수액대와 펌프 각 한 개가 있고 뒤 병상에도 별도의 수액 장치가 있다. 회녹색 커튼, 미닫이문, 천장 조명 두 개와 환기구, 밤 창문, 뒤쪽 모니터와 오른쪽 서랍 카트가 참조 병실 배치에 부합한다. 장관은 환자의 옆구리 쪽 통로에 서고 카메라는 반대편 난간 너머에서 머리 쪽과 가까운 측면을 본다. 다만 장관의 허벅지와 병상 길이 대부분을 보여 상체 중심 구도보다 넓다. 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 지정된 두 남성뿐이다. 박철진은 짧은 검은 머리와 중년 동아시아계 남성의 외형, 얼굴 상처를 보이며 참조 인물과 대체로 닮았다. 먼 쪽 눈에는 흰 덮개가 일부 보이고 팔과 다리에 각각 고정 붕대 또는 석고가 있다. 참조의 손상된 군복 대신 올리브색 상의를 입었다. 장관은 단정한 검은 머리의 성인 동아시아계 남성이며 짙은 정장, 흰 셔츠, 무늬 넥타이와 타이바가 참조와 부합한다. 얼굴은 참조보다 다소 각지고 나이 들어 보인다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "박철진의 머리와 목은 베개에, 등과 하체는 기울어진 병상과 매트리스에 지지된다. 굽힌 깁스 팔은 몸통 위 이불에 놓여 있고 깁스 다리는 병상 위에 누워 있어 공중에 떠 있지 않다. 장관의 발은 가려져 있지만 몸통과 팔의 자세는 병상 옆에 서 있는 사람으로 자연스럽다. 수액 용기는 수액대 고리에, 펌프는 기둥에 고정되어 있다."
       },
       {
        "label": "B",
        "direction": "장관은 몸을 환자 쪽으로 약간 기울이고 시선을 왼쪽 아래 박철진의 얼굴에 고정한다. 좁힌 눈과 다문 입이 차가운 평가와 멸시에 더 가깝다. 박철진의 얼굴은 위쪽 장관을 향하지만 가까운 눈은 패치로 가려지고 반대 눈은 측면 각도 때문에 명확하게 확인되지 않는다. 다른 조준 대상이나 휴대 소품은 없다.",
        "built_space": "전경의 점유 병상 한 개와 장관 뒤로 일부 보이는 빈 병상 한 개가 있다. 왼쪽 벽 의료 패널과 하부 조명 각 한 개, 전경 수액대·용기·펌프 각 한 개가 보이며 커튼과 미닫이문, 천장 조명 두 개, 환기구, 밤 창문, 뒤쪽 모니터와 오른쪽 서랍 카트가 참조 장소와 맞는다. 장관은 침대 옆 통로에 서고 환자의 머리와 어깨는 왼쪽 낮은 전경에 놓인다. 장관은 머리부터 재킷 하단까지 크게 잡혀 A보다 상체 중심 미디엄 숏에 가깝다. 병상 하반부까지 보이는 범위는 여전히 다소 넓다. 불가능한 반사나 명백한 설비 중복은 없다.",
        "entities": "지정된 두 인물만 보인다. 장관은 참조와 유사한 얼굴, 단정한 짧은 검은 머리, 짙은 정장과 흰 셔츠, 무늬 넥타이 및 타이바를 갖췄다. 박철진은 짧은 검은 머리의 중년 동아시아계 남성으로, 얼굴은 측면이라 정확한 동일성 확인이 제한된다. 한쪽 눈을 덮은 흰 패치와 팔·다리의 흰 고정 처치는 명확하다. 다만 얼굴 상처는 참조보다 약하고 의상은 참조 군복 대신 무늬 있는 환자복이다. 야간 창밖과 중립적인 병실 조명이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "박철진의 머리와 등은 베개와 올린 침대 등판에 기대어 있고 골반과 다리는 매트리스에 놓여 있다. 깁스 팔과 손은 가까운 침상 위에, 다른 손은 복부를 덮은 이불 위에 받쳐져 있다. 깁스 다리 역시 침상에 지지되어 있다. 장관은 침대 옆에 자연스럽게 서 있고 손을 아래로 내리고 있다. 수액 용기와 펌프는 각각 고리와 기둥에 지지되며, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "내려다보는 시선과 야간 병실은 맞지만, 장관의 허벅지와 병상 대부분까지 담아 상체 중심 미디엄 숏보다 넓고 환자를 향한 멸시도 상대적으로 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "장관의 상체를 더 크게 두고 낮은 전경의 박철진을 차갑게 내려다보게 해 핵심 구도에 더 충실하지만, 환자 하반신까지 담은 범위는 다소 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "장관은 고개와 눈을 화면 왼쪽 아래 박철진의 얼굴로 향한다. 박철진의 드러난 눈은 위쪽 장관 방향을 보고 있다. 장관의 찌푸린 표정은 냉담하지만 멸시보다는 심각한 관찰로도 읽힌다. 무기나 방향을 판정할 휴대 소품은 없다.",
        "built_space": "전경 병상 한 개와 뒤쪽 빈 병상 한 개가 보인다. 전경 왼쪽에는 벽 부착 의료 패널과 하부 조명 각 한 개, 수액대와 펌프 각 한 개가 있고 뒤 병상에도 별도의 수액 장치가 있다. 회녹색 커튼, 미닫이문, 천장 조명 두 개와 환기구, 밤 창문, 뒤쪽 모니터와 오른쪽 서랍 카트가 참조 병실 배치에 부합한다. 장관은 환자의 옆구리 쪽 통로에 서고 카메라는 반대편 난간 너머에서 머리 쪽과 가까운 측면을 본다. 다만 장관의 허벅지와 병상 길이 대부분을 보여 상체 중심 구도보다 넓다. 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 지정된 두 남성뿐이다. 박철진은 짧은 검은 머리와 중년 동아시아계 남성의 외형, 얼굴 상처를 보이며 참조 인물과 대체로 닮았다. 먼 쪽 눈에는 흰 덮개가 일부 보이고 팔과 다리에 각각 고정 붕대 또는 석고가 있다. 참조의 손상된 군복 대신 올리브색 상의를 입었다. 장관은 단정한 검은 머리의 성인 동아시아계 남성이며 짙은 정장, 흰 셔츠, 무늬 넥타이와 타이바가 참조와 부합한다. 얼굴은 참조보다 다소 각지고 나이 들어 보인다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "박철진의 머리와 목은 베개에, 등과 하체는 기울어진 병상과 매트리스에 지지된다. 굽힌 깁스 팔은 몸통 위 이불에 놓여 있고 깁스 다리는 병상 위에 누워 있어 공중에 떠 있지 않다. 장관의 발은 가려져 있지만 몸통과 팔의 자세는 병상 옆에 서 있는 사람으로 자연스럽다. 수액 용기는 수액대 고리에, 펌프는 기둥에 고정되어 있다."
       },
       {
        "label": "A",
        "direction": "장관은 몸을 환자 쪽으로 약간 기울이고 시선을 왼쪽 아래 박철진의 얼굴에 고정한다. 좁힌 눈과 다문 입이 차가운 평가와 멸시에 더 가깝다. 박철진의 얼굴은 위쪽 장관을 향하지만 가까운 눈은 패치로 가려지고 반대 눈은 측면 각도 때문에 명확하게 확인되지 않는다. 다른 조준 대상이나 휴대 소품은 없다.",
        "built_space": "전경의 점유 병상 한 개와 장관 뒤로 일부 보이는 빈 병상 한 개가 있다. 왼쪽 벽 의료 패널과 하부 조명 각 한 개, 전경 수액대·용기·펌프 각 한 개가 보이며 커튼과 미닫이문, 천장 조명 두 개, 환기구, 밤 창문, 뒤쪽 모니터와 오른쪽 서랍 카트가 참조 장소와 맞는다. 장관은 침대 옆 통로에 서고 환자의 머리와 어깨는 왼쪽 낮은 전경에 놓인다. 장관은 머리부터 재킷 하단까지 크게 잡혀 A보다 상체 중심 미디엄 숏에 가깝다. 병상 하반부까지 보이는 범위는 여전히 다소 넓다. 불가능한 반사나 명백한 설비 중복은 없다.",
        "entities": "지정된 두 인물만 보인다. 장관은 참조와 유사한 얼굴, 단정한 짧은 검은 머리, 짙은 정장과 흰 셔츠, 무늬 넥타이 및 타이바를 갖췄다. 박철진은 짧은 검은 머리의 중년 동아시아계 남성으로, 얼굴은 측면이라 정확한 동일성 확인이 제한된다. 한쪽 눈을 덮은 흰 패치와 팔·다리의 흰 고정 처치는 명확하다. 다만 얼굴 상처는 참조보다 약하고 의상은 참조 군복 대신 무늬 있는 환자복이다. 야간 창밖과 중립적인 병실 조명이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "박철진의 머리와 등은 베개와 올린 침대 등판에 기대어 있고 골반과 다리는 매트리스에 놓여 있다. 깁스 팔과 손은 가까운 침상 위에, 다른 손은 복부를 덮은 이불 위에 받쳐져 있다. 깁스 다리 역시 침상에 지지되어 있다. 장관은 침대 옆에 자연스럽게 서 있고 손을 아래로 내리고 있다. 수액 용기와 펌프는 각각 고리와 기둥에 지지되며, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.714,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.714,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1714
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "공간과 인물 외형을 정확히 구현했으며, 특히 장관의 멸시하는 듯한 차가운 표정을 텍스트 지시대로 훌륭히 연출함."
   },
   {
    "label": "A",
    "score": 1714,
    "verdict_ko": "구도와 배경 요소는 잘 맞췄으나, 장관의 표정이 차갑거나 멸시하기보다는 걱정스러워 보여 핵심 감정 연출 지시를 놓침."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_military_hospital_room_03987e.png",
    "asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1269105>",
    "asset_id": "ccc1d001-b248-4f0d-929f-9d10d042e801",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:835328>",
    "asset_id": "581db87a-927d-4d6e-8f59-4ccecac9881e",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-781a-7261-9822-07e83bf2e7ca",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5__bgfirst_bg.png",
   "bg_asset_id": "a0d10e4e-ddb6-42d2-9c3e-f1455cd58d3a",
   "bg_record_key": "S70sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "military_hospital_room",
   "groupbg_asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S70sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:18:24.850183+00:00",
  "fingerprint": "2211c17d952576a4d55a8f2f7f1ad8bfdc54a2b7d6a7193cafcc77784684076d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S70sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S70sh5_sel.png",
  "source_sha256": "e33b11befd07f32cb67d2b5376b776f79ecacb6f5e644a954384761c7b50cb56",
  "file": "S70sh5_cine.png",
  "staged_sha256": "936185361d60a771f89aae8faa5e125c7b9ceea4dfed54ace041fd8d6066f8bc",
  "latency_ms": 10350
 },
 "S70sh9::signage": {
  "fp": "ba8a0135d3930382",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S70sh9": {
  "input_fingerprint": "95da48ac994cc1f0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 붉게 충혈된 눈에서 뺨을 타고 눈물 한 방울이 흘러내리고 있는 찰나의 박철진의 억울하고 처절한 얼굴.\n\nLOCATION (lock): On the patient's bed inside the military-hospital room, under the ward's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment) — Only a narrow portion beside his head remains visible; used as A subdued edge establishes his recumbent position without distracting from the tear.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient hospital illumination preserves the reddened eye and the small highlight on the tear without specifying an unsupported fixture or light color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting remains a military hospital ward at night. 박철진: He remains bedridden with severe injuries to his face, arm and leg, with one eye covered and his arm and leg immobilized in casts. His remaining eye is bloodshot and filling with tears.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 붉게 충혈된 눈에서 뺨을 타고 눈물 한 방울이 흘러내리고 있는 찰나의 박철진의 억울하고 처절한 얼굴.\n\nLOCATION (lock): On the patient's bed inside the military-hospital room, under the ward's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment) — Only a narrow portion beside his head remains visible; used as A subdued edge establishes his recumbent position without distracting from the tear.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient hospital illumination preserves the reddened eye and the small highlight on the tear without specifying an unsupported fixture or light color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting remains a military hospital ward at night. 박철진: He remains bedridden with severe injuries to his face, arm and leg, with one eye covered and his arm and leg immobilized in casts. His remaining eye is bloodshot and filling with tears.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 붉게 충혈된 눈에서 뺨을 타고 눈물 한 방울이 흘러내리고 있는 찰나의 박철진의 억울하고 처절한 얼굴.\n\nLOCATION (lock): On the patient's bed inside the military-hospital room, under the ward's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment) — Only a narrow portion beside his head remains visible; used as A subdued edge establishes his recumbent position without distracting from the tear.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient hospital illumination preserves the reddened eye and the small highlight on the tear without specifying an unsupported fixture or light color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting remains a military hospital ward at night. 박철진: He remains bedridden with severe injuries to his face, arm and leg, with one eye covered and his arm and leg immobilized in casts. His remaining eye is bloodshot and filling with tears.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 위쪽을 향함.",
    "built_space": "병원 침대, 하얀 베개, 좌측의 침대 난간이 보임.",
    "entities": "박철진. 붉게 충혈된 눈과 눈물은 있으나, 레퍼런스의 안대와 얼굴 상처가 전혀 없음.",
    "hard_violations": [],
    "physics": "머리는 베개에 물리적으로 지지되며 눈물은 뺨을 타고 흐름."
   },
   {
    "label": "B",
    "direction": "시선은 위쪽을 향함.",
    "built_space": "병원 침대, 하얀 베개, 좌측의 침대 난간이 보임.",
    "entities": "박철진. 얼굴 상처와 충혈된 눈, 눈물은 있으나, 필수인 안대가 없음.",
    "hard_violations": [],
    "physics": "머리는 베개에 지지되며 눈물은 뺨을 따라 자연스럽게 흐름."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 상처와 눈물은 표현되었으나 핵심 설정인 안대(가려진 눈)가 누락되어 캐릭터 상태 구현에 실패함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "안대와 얼굴 상처가 모두 누락되어 레퍼런스 및 프롬프트의 캐릭터 외형 지침을 전혀 충족하지 못함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 위쪽을 향함.",
        "built_space": "병원 침대, 하얀 베개, 좌측의 침대 난간이 보임.",
        "entities": "박철진. 붉게 충혈된 눈과 눈물은 있으나, 레퍼런스의 안대와 얼굴 상처가 전혀 없음.",
        "hard_violations": [],
        "physics": "머리는 베개에 물리적으로 지지되며 눈물은 뺨을 타고 흐름."
       },
       {
        "label": "B",
        "direction": "시선은 위쪽을 향함.",
        "built_space": "병원 침대, 하얀 베개, 좌측의 침대 난간이 보임.",
        "entities": "박철진. 얼굴 상처와 충혈된 눈, 눈물은 있으나, 필수인 안대가 없음.",
        "hard_violations": [],
        "physics": "머리는 베개에 지지되며 눈물은 뺨을 따라 자연스럽게 흐름."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 상처와 눈물은 표현되었으나 핵심 설정인 안대(가려진 눈)가 누락되어 캐릭터 상태 구현에 실패함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "안대와 얼굴 상처가 모두 누락되어 레퍼런스 및 프롬프트의 캐릭터 외형 지침을 전혀 충족하지 못함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 위쪽을 향함.",
        "built_space": "병원 침대, 하얀 베개, 좌측의 침대 난간이 보임.",
        "entities": "박철진. 붉게 충혈된 눈과 눈물은 있으나, 레퍼런스의 안대와 얼굴 상처가 전혀 없음.",
        "hard_violations": [],
        "physics": "머리는 베개에 물리적으로 지지되며 눈물은 뺨을 타고 흐름."
       },
       {
        "label": "B",
        "direction": "시선은 위쪽을 향함.",
        "built_space": "병원 침대, 하얀 베개, 좌측의 침대 난간이 보임.",
        "entities": "박철진. 얼굴 상처와 충혈된 눈, 눈물은 있으나, 필수인 안대가 없음.",
        "hard_violations": [],
        "physics": "머리는 베개에 지지되며 눈물은 뺨을 따라 자연스럽게 흐름."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "얼굴을 꽉 채운 클로즈업과 충혈된 눈, 뺨의 눈물이 핵심 순간에 더 충실하지만, 필수인 한쪽 눈 가림이 없고 참고 사진의 심한 얼굴 상처가 약해졌다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "위를 보는 눈과 흘러내리는 눈물은 명확하지만, 가슴과 침대 머리판까지 넓게 보여 얼굴 중심 구도에서 멀어지며 한쪽 눈 가림도 누락했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈은 렌즈보다 화면 왼쪽 위, 천장 쪽의 화면 밖 지점을 향한다. 위를 보는 방향은 지정과 맞지만, 가리지 않은 눈 하나가 아니라 양쪽 눈이 모두 드러난다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "흰 베개 하나가 머리 뒤를 받치고, 화면 왼쪽에 침대 난간 일부 하나와 흐릿한 병실 커튼이 보인다. 얼굴이 화면 대부분을 차지하고 침대 구조물은 가장자리로 물러나 있어 요구한 얼굴 중심 클로즈업에 가깝다. 중복 설비나 반사는 없다. 좁은 구도 안에서 병실 재질과 낮은 조도는 참고 사진과 대체로 이어진다.",
        "entities": "중년 동아시아계 남성 한 명이며 짧은 검은 머리와 회갈색 둥근 목 티셔츠는 지정 및 참고 사진과 대체로 맞는다. 충혈된 눈과 뺨을 따라 빛나는 눈물 자국이 보이고, 긴장된 눈썹과 벌어진 입술이 고통을 전달한다. 그러나 필수 눈 가림이 없고 참고 환자의 뚜렷한 뺨 열상이 훨씬 약하다. 팔과 다리의 깁스는 구도 밖이므로 평가하지 않는다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수는 베개에 닿아 있고 목과 어깨는 침대에 기대어 누운 자세로 이어진다. 공중에 뜬 신체나 지지 없는 물체는 없다. 눈물은 피부 표면을 따라 흐르는 젖은 자국과 작은 방울로 표현되어 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "양쪽 눈의 시선이 화면 위쪽 천장 방향으로 향하며 렌즈를 직접 보지 않는다. 위를 보라는 자세는 맞지만 한쪽 눈을 가린 상태는 아니다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "베개 하나, 왼쪽 침대 난간 일부 하나, 상단의 구멍 뚫린 머리판 하나가 보이며 뒤로 커튼과 회색 벽이 놓인다. 각 부품의 배치와 크기는 누운 환자에게 가능한 구조다. 다만 머리판과 베개가 넓게 드러나고 가슴까지 포함되어, 머리 옆 침대의 좁은 부분만 남기라는 구도보다 배경 비중이 크다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명과 참고 사진에 가까운 회갈색 티셔츠가 보인다. 눈의 충혈과 뺨으로 내려오는 한 줄기 눈물 끝의 방울은 선명하다. 표정은 억울하고 처절하기보다는 비교적 절제되어 있다. 한쪽 눈의 덮개가 없고 참고 사진의 심한 얼굴 상처도 대부분 사라졌다. 팔과 다리의 깁스는 프레임 밖이며 다른 사람과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목은 베개에 기대고 어깨와 몸통은 침대에 놓여 있어 자세의 지지가 확인된다. 지지 없이 떠 있는 신체나 물체는 없다. 눈물은 눈 아래에서 뺨 표면을 따라 내려와 방울을 이루며 자연스럽게 붙어 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "얼굴을 꽉 채운 클로즈업과 충혈된 눈, 뺨의 눈물이 핵심 순간에 더 충실하지만, 필수인 한쪽 눈 가림이 없고 참고 사진의 심한 얼굴 상처가 약해졌다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "위를 보는 눈과 흘러내리는 눈물은 명확하지만, 가슴과 침대 머리판까지 넓게 보여 얼굴 중심 구도에서 멀어지며 한쪽 눈 가림도 누락했다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "두 눈은 렌즈보다 화면 왼쪽 위, 천장 쪽의 화면 밖 지점을 향한다. 위를 보는 방향은 지정과 맞지만, 가리지 않은 눈 하나가 아니라 양쪽 눈이 모두 드러난다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "흰 베개 하나가 머리 뒤를 받치고, 화면 왼쪽에 침대 난간 일부 하나와 흐릿한 병실 커튼이 보인다. 얼굴이 화면 대부분을 차지하고 침대 구조물은 가장자리로 물러나 있어 요구한 얼굴 중심 클로즈업에 가깝다. 중복 설비나 반사는 없다. 좁은 구도 안에서 병실 재질과 낮은 조도는 참고 사진과 대체로 이어진다.",
        "entities": "중년 동아시아계 남성 한 명이며 짧은 검은 머리와 회갈색 둥근 목 티셔츠는 지정 및 참고 사진과 대체로 맞는다. 충혈된 눈과 뺨을 따라 빛나는 눈물 자국이 보이고, 긴장된 눈썹과 벌어진 입술이 고통을 전달한다. 그러나 필수 눈 가림이 없고 참고 환자의 뚜렷한 뺨 열상이 훨씬 약하다. 팔과 다리의 깁스는 구도 밖이므로 평가하지 않는다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수는 베개에 닿아 있고 목과 어깨는 침대에 기대어 누운 자세로 이어진다. 공중에 뜬 신체나 지지 없는 물체는 없다. 눈물은 피부 표면을 따라 흐르는 젖은 자국과 작은 방울로 표현되어 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "양쪽 눈의 시선이 화면 위쪽 천장 방향으로 향하며 렌즈를 직접 보지 않는다. 위를 보라는 자세는 맞지만 한쪽 눈을 가린 상태는 아니다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "베개 하나, 왼쪽 침대 난간 일부 하나, 상단의 구멍 뚫린 머리판 하나가 보이며 뒤로 커튼과 회색 벽이 놓인다. 각 부품의 배치와 크기는 누운 환자에게 가능한 구조다. 다만 머리판과 베개가 넓게 드러나고 가슴까지 포함되어, 머리 옆 침대의 좁은 부분만 남기라는 구도보다 배경 비중이 크다. 중복 설비나 불가능한 반사는 없다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명과 참고 사진에 가까운 회갈색 티셔츠가 보인다. 눈의 충혈과 뺨으로 내려오는 한 줄기 눈물 끝의 방울은 선명하다. 표정은 억울하고 처절하기보다는 비교적 절제되어 있다. 한쪽 눈의 덮개가 없고 참고 사진의 심한 얼굴 상처도 대부분 사라졌다. 팔과 다리의 깁스는 프레임 밖이며 다른 사람과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목은 베개에 기대고 어깨와 몸통은 침대에 놓여 있어 자세의 지지가 확인된다. 지지 없이 떠 있는 신체나 물체는 없다. 눈물은 눈 아래에서 뺨 표면을 따라 내려와 방울을 이루며 자연스럽게 붙어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.583,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.583,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1583
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "얼굴 상처와 눈물은 표현되었으나 핵심 설정인 안대(가려진 눈)가 누락되어 캐릭터 상태 구현에 실패함."
   },
   {
    "label": "A",
    "score": 1583,
    "verdict_ko": "안대와 얼굴 상처가 모두 누락되어 레퍼런스 및 프롬프트의 캐릭터 외형 지침을 전혀 충족하지 못함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5_sel.png",
    "asset_id": "050ed65a-d195-4421-a2ce-7a77f7f64ec7",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-7cf0-7d30-8626-453d14832fa4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S70sh5"
  },
  "locked_char_refs_excluded": [
   "박철진(C16)"
  ]
 },
 "S70sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:19:18.148410+00:00",
  "fingerprint": "334e3aac144d33f13899f907284bb5e2779408eafbe4cc5c0efcd18c7f4c804e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S70sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S70sh9_sel.png",
  "source_sha256": "2df791ef6dd4980ac54448ae4004631849336a9e8a53c1efebc51ebc98b57a73",
  "file": "S70sh9_cine.png",
  "staged_sha256": "a563f76e58ca10dd176501bd02cb2b3b45b375be99bfd9691741bf4fff430e44",
  "latency_ms": 9545
 },
 "S71sh4::signage": {
  "fp": "21b31e69af2b58e2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::open_truck_cab": {
  "input_fingerprint": "b375775d773c387c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "open_truck_cab",
    "tags": [
     "S63sh17",
     "S71sh17",
     "S71sh4",
     "S72sh43",
     "S72sh58"
    ]
   },
   "context_sig": "db8201d83d85b000"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해남 지동현의 연구소 1층: 오래 방치되어 먼지가 쌓인 텅 빈 로비와 나선형 구조의 실내. (특징: 끼익 소리를 내며 열리는 낡은 거대 철문; 비어 있는 1층 실내 공간과 뽀얀 바닥 먼지; 바닥에 떨어져 유리가 깨진 액자; 위층으로 이어지는 나선형 철제 계단) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 홀로 트럭을 운전하는 현우. 보조석에는 수빈이 타고 있고.\n- 지붕 하나 없는 낡은 트럭들이다.\n- 71. 도로 위, 현우의 트럭 – 새벽\n- / 도로 위. 차 안 - N\n- 드르르르르륵-! 문이 열리면-!\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해남 지동현의 연구소 1층: 오래 방치되어 먼지가 쌓인 텅 빈 로비와 나선형 구조의 실내. (특징: 끼익 소리를 내며 열리는 낡은 거대 철문; 비어 있는 1층 실내 공간과 뽀얀 바닥 먼지; 바닥에 떨어져 유리가 깨진 액자; 위층으로 이어지는 나선형 철제 계단) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 홀로 트럭을 운전하는 현우. 보조석에는 수빈이 타고 있고.\n- 지붕 하나 없는 낡은 트럭들이다.\n- 71. 도로 위, 현우의 트럭 – 새벽\n- / 도로 위. 차 안 - N\n- 드르르르르륵-! 문이 열리면-!\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_cab_cebf8e.png",
  "asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824",
  "input_asset_ids": [
   "2d015eed-4450-4662-acb3-9b08a2be5734"
  ],
  "origin_tag": "S71sh4",
  "place_text": "At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.",
  "origin_inputs": {
   "place_text": "At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.",
   "time_of_day_en": "dawn",
   "conti_asset_id": "2d015eed-4450-4662-acb3-9b08a2be5734"
  }
 },
 "S71sh4::bgfirst_bg": {
  "input_fingerprint": "5a5a48233e11ce51",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh4__bgfirst_bg.png",
  "asset_id": "1e25c307-2725-4133-9fb5-567004d17fb4",
  "input_asset_ids": [
   "2d015eed-4450-4662-acb3-9b08a2be5734",
   "cf509bfd-c4f2-4e5c-850c-88fea08ed824"
  ]
 },
 "S71sh4": {
  "input_fingerprint": "11b2c9e303156738",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck is traveling south at dawn; the violet sketch is drawn on the reverse of a sheet of paper. Charlie remains damaged, with punctured and dented bodywork, an exposed chest opening and malfunctioning sensors, and retains B-200's chest component. 앰버: She holds the sheet bearing the flower drawing. The injury to the back of her head from the rifle-butt blow has not been treated on screen.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck is traveling south at dawn; the violet sketch is drawn on the reverse of a sheet of paper. Charlie remains damaged, with punctured and dented bodywork, an exposed chest opening and malfunctioning sensors, and retains B-200's chest component. 앰버: She holds the sheet bearing the flower drawing. The injury to the back of her head from the rifle-butt blow has not been treated on screen.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck is traveling south at dawn; the violet sketch is drawn on the reverse of a sheet of paper. Charlie remains damaged, with punctured and dented bodywork, an exposed chest opening and malfunctioning sensors, and retains B-200's chest component. 앰버: She holds the sheet bearing the flower drawing. The injury to the back of her head from the rifle-butt blow has not been treated on screen.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh4__bgfirst_bg.png",
     "asset_id": "1e25c307-2725-4133-9fb5-567004d17fb4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S71sh4.png",
     "asset_id": "2d015eed-4450-4662-acb3-9b08a2be5734",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 찰리의 꽃 그림: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1316405>",
     "asset_id": "108fb164-c05c-4d1c-9144-293628b1d420",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_cab_cebf8e.png",
     "asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 찰리의 꽃 그림: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1316405>",
     "asset_id": "108fb164-c05c-4d1c-9144-293628b1d420",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 눈썹을 찌푸리며 손에 든 종이 쪽을 내려다봅니다. 꽃 그림 면은 앰버보다 카메라 쪽을 향하며, 요구한 비스듬한 제시보다는 정면에 가깝습니다. 오른쪽에서 들어온 성인의 손도 종이 오른쪽 가장자리를 붙잡고 있어, 이미 전달받은 뒤라기보다 전달 중인 순간으로 보입니다.",
    "built_space": "왼쪽 문과 창틀 하나, 사이드미러 하나, 앰버 뒤 좌석과 오른쪽 빈 좌석의 등받이 두 개, 뒤쪽 열린 적재함과 녹슨 프레임이 보입니다. 앰버는 좌석 앞에 정상적으로 앉아 있고 대시보드는 화면 아래를 비스듬하게 가로지릅니다. 참조 트럭의 낡은 금속과 배치는 대체로 유지되며, 명백한 중복 설비나 불가능한 반사는 보이지 않습니다.",
    "entities": "금발의 어린 여자아이 한 명은 앰버의 나이대와 외형에 대체로 부합하지만, 혼혈 배경 자체는 외모만으로 확정할 수 없습니다. 남색 티셔츠는 맞지만 멜빵 작업복은 참조의 갈색 대신 청색이며 머리 위 고글도 없습니다. 종이는 가로로 접힌 회색빛 종이이고 꽃은 작은 무채색 선화여서, 참조의 세로형 종이와 제비꽃 묶음에서 크게 다릅니다. 오른쪽의 굵은 맨팔과 큰 손은 앰버의 신체가 아닌 추가 성인의 것입니다. 읽을 수 있는 글자는 없습니다.",
    "hard_violations": [
     "앰버만 등장해야 하는 화면에 다른 성인의 손과 팔을 추가했습니다. 앰버의 두 손이 별도로 종이를 잡고 있어 해당 팔을 앰버의 것으로 볼 수 없습니다."
    ],
    "physics": "앰버의 몸은 좌석에 지지되고 두 손이 종이 양쪽을 잡습니다. 종이의 접힘과 처짐은 손의 지지와 양립합니다. 추가 성인의 손도 화면 밖으로 이어지는 팔에 연결되어 있어 부유 신체는 아니지만, 허용되지 않은 인물의 신체라는 별도 문제가 있습니다."
   },
   {
    "label": "A",
    "direction": "앰버는 종이를 가슴 앞에 든 채 화면 오른쪽 밖을 바라봅니다. 시선의 상대는 보이지 않지만, 그림을 받은 뒤 건넨 상대에게 의문을 표하는 순간으로 자연스럽게 읽힙니다. 꽃 그림 면은 지시대로 카메라를 향해 약간 비스듬히 놓이며, 종이를 읽는 척하면서 반대 면을 보는 동작은 아닙니다.",
    "built_space": "왼쪽 문과 창틀 하나, 사이드미러 하나, 앰버가 사용하는 좌석과 오른쪽 빈 좌석의 등받이 두 개가 보입니다. 뒤에는 참조와 같은 녹슨 개방형 프레임과 적재함이 있고, 하단에는 대시보드 가장자리가 사선으로 놓입니다. 좌석에 앉은 앰버와 구조물의 크기 관계가 자연스럽고 종이도 대시보드에 가려지지 않습니다. 사이드미러의 하늘과 노면 반사에 명백한 광학적 모순은 없습니다.",
    "entities": "금발, 큰 눈, 둥근 얼굴의 열 살 안팎 여자아이로 참조 앰버와 잘 부합합니다. 혼혈 배경은 영상만으로 단정할 수 없지만 참조 얼굴과의 유사성은 높습니다. 남색 티셔츠와 갈색 작업복은 일치하며, 머리 위 고글은 빠졌습니다. 거친 가장자리의 밝은 종이에 보라색 제비꽃과 초록 잎이 그려져 있어 참조 소품에 가깝고, 선과 색은 종이 표면에 그려진 것으로 보입니다. 앰버 외 인물이나 신체 일부, 읽을 수 있는 글자는 없습니다. 뒤통수 상처는 이 각도에서 확인되지 않습니다.",
    "hard_violations": [],
    "physics": "앰버는 등받이 앞 좌석에 앉아 있고 안전벨트가 어깨에서 몸통으로 이어집니다. 두 손은 종이 좌우 가장자리를 실제로 집고 있으며 손목과 팔뚝도 아이의 체격에 맞습니다. 종이는 그립 사이에서 약간 휘어지고, 지지 없이 떠 있는 물체나 신체는 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "상체 구도와 당혹스러운 표정은 맞지만, 앰버 외 성인의 팔을 추가한 중대 위반이 있고 작업복과 꽃 그림도 참조에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "앰버만 보이는 상체 미디엄 숏에서 종이를 받은 뒤의 어리둥절함, 참조 꽃 그림과 작업복, 새벽 트럭 공간을 충실하게 구현합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 눈썹을 찌푸리며 손에 든 종이 쪽을 내려다봅니다. 꽃 그림 면은 앰버보다 카메라 쪽을 향하며, 요구한 비스듬한 제시보다는 정면에 가깝습니다. 오른쪽에서 들어온 성인의 손도 종이 오른쪽 가장자리를 붙잡고 있어, 이미 전달받은 뒤라기보다 전달 중인 순간으로 보입니다.",
        "built_space": "왼쪽 문과 창틀 하나, 사이드미러 하나, 앰버 뒤 좌석과 오른쪽 빈 좌석의 등받이 두 개, 뒤쪽 열린 적재함과 녹슨 프레임이 보입니다. 앰버는 좌석 앞에 정상적으로 앉아 있고 대시보드는 화면 아래를 비스듬하게 가로지릅니다. 참조 트럭의 낡은 금속과 배치는 대체로 유지되며, 명백한 중복 설비나 불가능한 반사는 보이지 않습니다.",
        "entities": "금발의 어린 여자아이 한 명은 앰버의 나이대와 외형에 대체로 부합하지만, 혼혈 배경 자체는 외모만으로 확정할 수 없습니다. 남색 티셔츠는 맞지만 멜빵 작업복은 참조의 갈색 대신 청색이며 머리 위 고글도 없습니다. 종이는 가로로 접힌 회색빛 종이이고 꽃은 작은 무채색 선화여서, 참조의 세로형 종이와 제비꽃 묶음에서 크게 다릅니다. 오른쪽의 굵은 맨팔과 큰 손은 앰버의 신체가 아닌 추가 성인의 것입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "앰버만 등장해야 하는 화면에 다른 성인의 손과 팔을 추가했습니다. 앰버의 두 손이 별도로 종이를 잡고 있어 해당 팔을 앰버의 것으로 볼 수 없습니다."
        ],
        "physics": "앰버의 몸은 좌석에 지지되고 두 손이 종이 양쪽을 잡습니다. 종이의 접힘과 처짐은 손의 지지와 양립합니다. 추가 성인의 손도 화면 밖으로 이어지는 팔에 연결되어 있어 부유 신체는 아니지만, 허용되지 않은 인물의 신체라는 별도 문제가 있습니다."
       },
       {
        "label": "B",
        "direction": "앰버는 종이를 가슴 앞에 든 채 화면 오른쪽 밖을 바라봅니다. 시선의 상대는 보이지 않지만, 그림을 받은 뒤 건넨 상대에게 의문을 표하는 순간으로 자연스럽게 읽힙니다. 꽃 그림 면은 지시대로 카메라를 향해 약간 비스듬히 놓이며, 종이를 읽는 척하면서 반대 면을 보는 동작은 아닙니다.",
        "built_space": "왼쪽 문과 창틀 하나, 사이드미러 하나, 앰버가 사용하는 좌석과 오른쪽 빈 좌석의 등받이 두 개가 보입니다. 뒤에는 참조와 같은 녹슨 개방형 프레임과 적재함이 있고, 하단에는 대시보드 가장자리가 사선으로 놓입니다. 좌석에 앉은 앰버와 구조물의 크기 관계가 자연스럽고 종이도 대시보드에 가려지지 않습니다. 사이드미러의 하늘과 노면 반사에 명백한 광학적 모순은 없습니다.",
        "entities": "금발, 큰 눈, 둥근 얼굴의 열 살 안팎 여자아이로 참조 앰버와 잘 부합합니다. 혼혈 배경은 영상만으로 단정할 수 없지만 참조 얼굴과의 유사성은 높습니다. 남색 티셔츠와 갈색 작업복은 일치하며, 머리 위 고글은 빠졌습니다. 거친 가장자리의 밝은 종이에 보라색 제비꽃과 초록 잎이 그려져 있어 참조 소품에 가깝고, 선과 색은 종이 표면에 그려진 것으로 보입니다. 앰버 외 인물이나 신체 일부, 읽을 수 있는 글자는 없습니다. 뒤통수 상처는 이 각도에서 확인되지 않습니다.",
        "hard_violations": [],
        "physics": "앰버는 등받이 앞 좌석에 앉아 있고 안전벨트가 어깨에서 몸통으로 이어집니다. 두 손은 종이 좌우 가장자리를 실제로 집고 있으며 손목과 팔뚝도 아이의 체격에 맞습니다. 종이는 그립 사이에서 약간 휘어지고, 지지 없이 떠 있는 물체나 신체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상체 구도와 당혹스러운 표정은 맞지만, 앰버 외 성인의 팔을 추가한 중대 위반이 있고 작업복과 꽃 그림도 참조에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "앰버만 보이는 상체 미디엄 숏에서 종이를 받은 뒤의 어리둥절함, 참조 꽃 그림과 작업복, 새벽 트럭 공간을 충실하게 구현합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 눈썹을 찌푸리며 손에 든 종이 쪽을 내려다봅니다. 꽃 그림 면은 앰버보다 카메라 쪽을 향하며, 요구한 비스듬한 제시보다는 정면에 가깝습니다. 오른쪽에서 들어온 성인의 손도 종이 오른쪽 가장자리를 붙잡고 있어, 이미 전달받은 뒤라기보다 전달 중인 순간으로 보입니다.",
        "built_space": "왼쪽 문과 창틀 하나, 사이드미러 하나, 앰버 뒤 좌석과 오른쪽 빈 좌석의 등받이 두 개, 뒤쪽 열린 적재함과 녹슨 프레임이 보입니다. 앰버는 좌석 앞에 정상적으로 앉아 있고 대시보드는 화면 아래를 비스듬하게 가로지릅니다. 참조 트럭의 낡은 금속과 배치는 대체로 유지되며, 명백한 중복 설비나 불가능한 반사는 보이지 않습니다.",
        "entities": "금발의 어린 여자아이 한 명은 앰버의 나이대와 외형에 대체로 부합하지만, 혼혈 배경 자체는 외모만으로 확정할 수 없습니다. 남색 티셔츠는 맞지만 멜빵 작업복은 참조의 갈색 대신 청색이며 머리 위 고글도 없습니다. 종이는 가로로 접힌 회색빛 종이이고 꽃은 작은 무채색 선화여서, 참조의 세로형 종이와 제비꽃 묶음에서 크게 다릅니다. 오른쪽의 굵은 맨팔과 큰 손은 앰버의 신체가 아닌 추가 성인의 것입니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "앰버만 등장해야 하는 화면에 다른 성인의 손과 팔을 추가했습니다. 앰버의 두 손이 별도로 종이를 잡고 있어 해당 팔을 앰버의 것으로 볼 수 없습니다."
        ],
        "physics": "앰버의 몸은 좌석에 지지되고 두 손이 종이 양쪽을 잡습니다. 종이의 접힘과 처짐은 손의 지지와 양립합니다. 추가 성인의 손도 화면 밖으로 이어지는 팔에 연결되어 있어 부유 신체는 아니지만, 허용되지 않은 인물의 신체라는 별도 문제가 있습니다."
       },
       {
        "label": "A",
        "direction": "앰버는 종이를 가슴 앞에 든 채 화면 오른쪽 밖을 바라봅니다. 시선의 상대는 보이지 않지만, 그림을 받은 뒤 건넨 상대에게 의문을 표하는 순간으로 자연스럽게 읽힙니다. 꽃 그림 면은 지시대로 카메라를 향해 약간 비스듬히 놓이며, 종이를 읽는 척하면서 반대 면을 보는 동작은 아닙니다.",
        "built_space": "왼쪽 문과 창틀 하나, 사이드미러 하나, 앰버가 사용하는 좌석과 오른쪽 빈 좌석의 등받이 두 개가 보입니다. 뒤에는 참조와 같은 녹슨 개방형 프레임과 적재함이 있고, 하단에는 대시보드 가장자리가 사선으로 놓입니다. 좌석에 앉은 앰버와 구조물의 크기 관계가 자연스럽고 종이도 대시보드에 가려지지 않습니다. 사이드미러의 하늘과 노면 반사에 명백한 광학적 모순은 없습니다.",
        "entities": "금발, 큰 눈, 둥근 얼굴의 열 살 안팎 여자아이로 참조 앰버와 잘 부합합니다. 혼혈 배경은 영상만으로 단정할 수 없지만 참조 얼굴과의 유사성은 높습니다. 남색 티셔츠와 갈색 작업복은 일치하며, 머리 위 고글은 빠졌습니다. 거친 가장자리의 밝은 종이에 보라색 제비꽃과 초록 잎이 그려져 있어 참조 소품에 가깝고, 선과 색은 종이 표면에 그려진 것으로 보입니다. 앰버 외 인물이나 신체 일부, 읽을 수 있는 글자는 없습니다. 뒤통수 상처는 이 각도에서 확인되지 않습니다.",
        "hard_violations": [],
        "physics": "앰버는 등받이 앞 좌석에 앉아 있고 안전벨트가 어깨에서 몸통으로 이어집니다. 두 손은 종이 좌우 가장자리를 실제로 집고 있으며 손목과 팔뚝도 아이의 체격에 맞습니다. 종이는 그립 사이에서 약간 휘어지고, 지지 없이 떠 있는 물체나 신체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 9
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "상체 구도와 당혹스러운 표정은 맞지만, 앰버 외 성인의 팔을 추가한 중대 위반이 있고 작업복과 꽃 그림도 참조에서 벗어납니다."
   },
   {
    "label": "A",
    "score": 9,
    "verdict_ko": "앰버만 보이는 상체 미디엄 숏에서 종이를 받은 뒤의 어리둥절함, 참조 꽃 그림과 작업복, 새벽 트럭 공간을 충실하게 구현합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_cab_cebf8e.png",
    "asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 찰리의 꽃 그림: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:1316405>",
    "asset_id": "108fb164-c05c-4d1c-9144-293628b1d420",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-7e94-7fac-afe0-1412cb06b306",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh4__bgfirst_bg.png",
   "bg_asset_id": "1e25c307-2725-4133-9fb5-567004d17fb4",
   "bg_record_key": "S71sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "open_truck_cab",
   "groupbg_asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S71sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:20:24.727592+00:00",
  "fingerprint": "3f20cdb1dd85cbbd6a9e925ada2fce8ef8efd019b035baf834d9c367055d4a1b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S71sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S71sh4_sel.png",
  "source_sha256": "a7167876efdb9dd44488b047adfe897554f589339976d7bc3a580b7a00522091",
  "file": "S71sh4_cine.png",
  "staged_sha256": "8ac5ae6013399cbf927e45803b80325178babc93d1b836fe065fd7950a5c1a52",
  "latency_ms": 11176
 },
 "S71sh7::signage": {
  "fp": "9c51bf276a36526b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::coastal_search_hill": {
  "input_fingerprint": "7a41bf6375577d7c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "coastal_search_hill",
    "tags": [
     "S71sh7"
    ]
   },
   "context_sig": "81a953d4a0d2dc9c"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 언덕 위에 서서 능선을 쳐다보는 현우.\n- 현우가 능선 너머로 제비꽃을 찾는 사이, 앰버와 찰리, 라울은 흙으로 장난을 치고 있다.\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 언덕 위에 서서 능선을 쳐다보는 현우.\n- 현우가 능선 너머로 제비꽃을 찾는 사이, 앰버와 찰리, 라울은 흙으로 장난을 치고 있다.\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_coastal_search_hill_7092b9.png",
  "asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21",
  "input_asset_ids": [
   "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1"
  ],
  "origin_tag": "S71sh7",
  "place_text": "On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.",
  "origin_inputs": {
   "place_text": "On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.",
   "time_of_day_en": "dawn",
   "conti_asset_id": "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1"
  }
 },
 "S71sh7::bgfirst_bg": {
  "input_fingerprint": "6dca32eb5c7fcfc8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh7__bgfirst_bg.png",
  "asset_id": "0cca7227-2e69-4db8-ab19-8acea477fee5",
  "input_asset_ids": [
   "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1",
   "ab08fa93-ee50-4572-91f7-b41dbd747f21"
  ]
 },
 "S71sh7": {
  "input_fingerprint": "eeddafd87fd8f91c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped near the coastal hills, and the flower sketch remains available without any stated disposal. Charlie's damaged bodywork, exposed chest opening and B-200 chest component persist. 앰버: She has tearful eyes and is wiping away tears; her earlier head injury remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped near the coastal hills, and the flower sketch remains available without any stated disposal. Charlie's damaged bodywork, exposed chest opening and B-200 chest component persist. 앰버: She has tearful eyes and is wiping away tears; her earlier head injury remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped near the coastal hills, and the flower sketch remains available without any stated disposal. Charlie's damaged bodywork, exposed chest opening and B-200 chest component persist. 앰버: She has tearful eyes and is wiping away tears; her earlier head injury remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh7__bgfirst_bg.png",
     "asset_id": "0cca7227-2e69-4db8-ab19-8acea477fee5",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S71sh7.png",
     "asset_id": "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_coastal_search_hill_7092b9.png",
     "asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "앰버의 시선이 화면 우측의 찰리를 향해 위로 향함.",
    "built_space": "새벽의 해안가 언덕, 배경에 바다와 픽업트럭이 배치됨.",
    "entities": "앰버(머리 부상 묘사 누락됨), 찰리, 그리고 우측 근경에 명시되지 않은 세 번째 기계 인물의 어깨 일부가 등장함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 세 번째 피사체(화면 우측 근경의 기계 어깨)가 추가됨."
    ],
    "physics": "앰버의 손이 얼굴 뺨에 닿아 눈물을 닦고 있으며, 캐릭터들이 서 있음."
   },
   {
    "label": "B",
    "direction": "앰버의 시선이 화면 상단에 위치한 찰리의 기계 팔 방향으로 위를 향함.",
    "built_space": "새벽의 해안가 언덕, 배경에 바다와 수풀이 배치됨.",
    "entities": "앰버(머리에 긁힌 부상과 핏자국 묘사됨, 눈물 흘림), 찰리의 기계 팔 일부.",
    "hard_violations": [],
    "physics": "찰리의 기계 팔이 앰버의 머리 위쪽을 자연스럽게 짚고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 프레이밍을 정확히 따랐으며 앰버의 머리 부상과 눈물 등 디테일을 훌륭하게 묘사했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "요구된 클로즈업을 무시하고 프레임이 너무 넓으며, 우측 근경에 프롬프트에 없는 피사체가 추가된 치명적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선이 화면 우측의 찰리를 향해 위로 향함.",
        "built_space": "새벽의 해안가 언덕, 배경에 바다와 픽업트럭이 배치됨.",
        "entities": "앰버(머리 부상 묘사 누락됨), 찰리, 그리고 우측 근경에 명시되지 않은 세 번째 기계 인물의 어깨 일부가 등장함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 세 번째 피사체(화면 우측 근경의 기계 어깨)가 추가됨."
        ],
        "physics": "앰버의 손이 얼굴 뺨에 닿아 눈물을 닦고 있으며, 캐릭터들이 서 있음."
       },
       {
        "label": "B",
        "direction": "앰버의 시선이 화면 상단에 위치한 찰리의 기계 팔 방향으로 위를 향함.",
        "built_space": "새벽의 해안가 언덕, 배경에 바다와 수풀이 배치됨.",
        "entities": "앰버(머리에 긁힌 부상과 핏자국 묘사됨, 눈물 흘림), 찰리의 기계 팔 일부.",
        "hard_violations": [],
        "physics": "찰리의 기계 팔이 앰버의 머리 위쪽을 자연스럽게 짚고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 프레이밍을 정확히 따랐으며 앰버의 머리 부상과 눈물 등 디테일을 훌륭하게 묘사했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "요구된 클로즈업을 무시하고 프레임이 너무 넓으며, 우측 근경에 프롬프트에 없는 피사체가 추가된 치명적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선이 화면 우측의 찰리를 향해 위로 향함.",
        "built_space": "새벽의 해안가 언덕, 배경에 바다와 픽업트럭이 배치됨.",
        "entities": "앰버(머리 부상 묘사 누락됨), 찰리, 그리고 우측 근경에 명시되지 않은 세 번째 기계 인물의 어깨 일부가 등장함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 세 번째 피사체(화면 우측 근경의 기계 어깨)가 추가됨."
        ],
        "physics": "앰버의 손이 얼굴 뺨에 닿아 눈물을 닦고 있으며, 캐릭터들이 서 있음."
       },
       {
        "label": "B",
        "direction": "앰버의 시선이 화면 상단에 위치한 찰리의 기계 팔 방향으로 위를 향함.",
        "built_space": "새벽의 해안가 언덕, 배경에 바다와 수풀이 배치됨.",
        "entities": "앰버(머리에 긁힌 부상과 핏자국 묘사됨, 눈물 흘림), 찰리의 기계 팔 일부.",
        "hard_violations": [],
        "physics": "찰리의 기계 팔이 앰버의 머리 위쪽을 자연스럽게 짚고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "눈물 맺힌 굳은 얼굴을 크게 잡고 찰리 쪽을 올려다보게 하며 머리 부상까지 유지해, 핵심 클로즈업 지시를 더 충실히 구현한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "찰리와의 시선 교환과 눈물 닦기는 맞지만, 앰버의 얼굴 클로즈업을 넓은 상반신 구도로 바꾸고 드러난 이마의 기존 부상도 재현하지 않았다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 두 눈은 화면 왼쪽 위, 머리 위로 들어온 찰리의 팔이 이어지는 화면 밖 방향을 향한다. 찰리의 얼굴은 보이지 않아 정확한 눈맞춤까지 확인할 수는 없지만, 자신보다 높은 찰리를 올려다보는 방향은 성립한다.",
        "built_space": "앰버의 얼굴이 화면 중앙과 오른쪽을 크게 차지한다. 뒤에는 흙길, 바위, 낮은 관목과 작은 보라색 꽃이 있고 오른쪽 위로 바다가 보인다. 도로나 고정 시설은 이 좁은 구도에 포함되지 않는다. 해안 언덕의 재료와 식생은 참조와 부합하지만, 지면 배경은 요구한 부드러운 처리보다 조금 선명하다.",
        "entities": "금발의 약 10세 여자아이 한 명과 찰리의 샌드 베이지 장갑 팔·손 일부가 보인다. 앰버는 큰 눈과 어린 얼굴을 갖췄으나 참조보다 얼굴이 다소 길고 눈 색도 달라 보인다. 혼혈 배경은 외모만으로 확정할 수 없다. 눈가의 수분, 볼을 흐르는 눈물, 다문 입과 굳은 표정, 이마의 상처와 작은 드레싱이 보인다. 참조의 머리 위 고글은 없다. 의상 대부분과 찰리의 가슴 부품, 트럭, 꽃 스케치는 프레임 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 장갑 손은 앰버의 정수리에 닿고 팔은 화면 왼쪽 밖으로 이어져, 독립적으로 떠 있는 물체가 아니다. 앰버의 머리와 목은 아래쪽 몸통으로 자연스럽게 연결된다. 발과 지면 접촉은 클로즈업 밖이므로 판단 대상이 아니다. 눈물은 피부를 따라 아래로 흐르며, 눈물을 닦는 앰버의 손은 이 구도에 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "앰버는 화면 오른쪽 위에 있는 찰리의 얼굴을 올려다보고, 찰리는 고개를 숙여 앰버를 향한다. 시선의 대상이 명확하며 앰버의 손끝은 자신의 눈가를 향해 눈물을 닦는 위치에 놓여 있다.",
        "built_space": "앰버의 머리와 상반신이 왼쪽 아래에 있고 찰리의 머리·가슴과 큰 전경 팔 부분이 오른쪽을 차지한다. 뒤쪽 왼편에는 정차한 픽업트럭 한 대와 도로를 따라 이어진 가드레일이 보이며, 바다와 해안 능선, 바위와 꽃이 있는 언덕이 넓게 드러난다. 장소의 배치는 참조와 잘 맞지만, 얼굴 클로즈업과 낮게 깔린 부드러운 지면 배경 대신 인물 둘과 풍경을 함께 보여주는 넓은 구도다.",
        "entities": "금발의 어린 여자아이와 흰 마스크형 얼굴, 모자, 올리브색 코트, 샌드 베이지 장갑판을 지닌 찰리가 보인다. 앰버의 남색 상의와 갈색 작업복은 참조에 부합하며 눈가에 눈물도 있다. 다만 참조의 고글이 없고 드러난 이마에는 지속되어야 할 부상이 보이지 않는다. 찰리의 원형 가슴 부품은 보이지만 손상된 개방부와 내부 부품 노출은 뚜렷하지 않다. 꽃 스케치는 보이지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 들어 올린 팔은 어깨부터 팔꿈치와 손목까지 연결되며 손가락이 눈가에 닿아, 눈물을 닦는 동작이 물리적으로 성립한다. 찰리의 머리와 몸통도 연결되어 있다. 오른쪽 전경의 큰 소매와 장갑 부분은 심하게 잘려 관절 연결을 전부 확인할 수 없지만, 별도로 떠 있는 물체라고 단정할 근거는 없다. 두 인물의 발은 프레임 밖이다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "눈물 맺힌 굳은 얼굴을 크게 잡고 찰리 쪽을 올려다보게 하며 머리 부상까지 유지해, 핵심 클로즈업 지시를 더 충실히 구현한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "찰리와의 시선 교환과 눈물 닦기는 맞지만, 앰버의 얼굴 클로즈업을 넓은 상반신 구도로 바꾸고 드러난 이마의 기존 부상도 재현하지 않았다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 두 눈은 화면 왼쪽 위, 머리 위로 들어온 찰리의 팔이 이어지는 화면 밖 방향을 향한다. 찰리의 얼굴은 보이지 않아 정확한 눈맞춤까지 확인할 수는 없지만, 자신보다 높은 찰리를 올려다보는 방향은 성립한다.",
        "built_space": "앰버의 얼굴이 화면 중앙과 오른쪽을 크게 차지한다. 뒤에는 흙길, 바위, 낮은 관목과 작은 보라색 꽃이 있고 오른쪽 위로 바다가 보인다. 도로나 고정 시설은 이 좁은 구도에 포함되지 않는다. 해안 언덕의 재료와 식생은 참조와 부합하지만, 지면 배경은 요구한 부드러운 처리보다 조금 선명하다.",
        "entities": "금발의 약 10세 여자아이 한 명과 찰리의 샌드 베이지 장갑 팔·손 일부가 보인다. 앰버는 큰 눈과 어린 얼굴을 갖췄으나 참조보다 얼굴이 다소 길고 눈 색도 달라 보인다. 혼혈 배경은 외모만으로 확정할 수 없다. 눈가의 수분, 볼을 흐르는 눈물, 다문 입과 굳은 표정, 이마의 상처와 작은 드레싱이 보인다. 참조의 머리 위 고글은 없다. 의상 대부분과 찰리의 가슴 부품, 트럭, 꽃 스케치는 프레임 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 장갑 손은 앰버의 정수리에 닿고 팔은 화면 왼쪽 밖으로 이어져, 독립적으로 떠 있는 물체가 아니다. 앰버의 머리와 목은 아래쪽 몸통으로 자연스럽게 연결된다. 발과 지면 접촉은 클로즈업 밖이므로 판단 대상이 아니다. 눈물은 피부를 따라 아래로 흐르며, 눈물을 닦는 앰버의 손은 이 구도에 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "앰버는 화면 오른쪽 위에 있는 찰리의 얼굴을 올려다보고, 찰리는 고개를 숙여 앰버를 향한다. 시선의 대상이 명확하며 앰버의 손끝은 자신의 눈가를 향해 눈물을 닦는 위치에 놓여 있다.",
        "built_space": "앰버의 머리와 상반신이 왼쪽 아래에 있고 찰리의 머리·가슴과 큰 전경 팔 부분이 오른쪽을 차지한다. 뒤쪽 왼편에는 정차한 픽업트럭 한 대와 도로를 따라 이어진 가드레일이 보이며, 바다와 해안 능선, 바위와 꽃이 있는 언덕이 넓게 드러난다. 장소의 배치는 참조와 잘 맞지만, 얼굴 클로즈업과 낮게 깔린 부드러운 지면 배경 대신 인물 둘과 풍경을 함께 보여주는 넓은 구도다.",
        "entities": "금발의 어린 여자아이와 흰 마스크형 얼굴, 모자, 올리브색 코트, 샌드 베이지 장갑판을 지닌 찰리가 보인다. 앰버의 남색 상의와 갈색 작업복은 참조에 부합하며 눈가에 눈물도 있다. 다만 참조의 고글이 없고 드러난 이마에는 지속되어야 할 부상이 보이지 않는다. 찰리의 원형 가슴 부품은 보이지만 손상된 개방부와 내부 부품 노출은 뚜렷하지 않다. 꽃 스케치는 보이지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 들어 올린 팔은 어깨부터 팔꿈치와 손목까지 연결되며 손가락이 눈가에 닿아, 눈물을 닦는 동작이 물리적으로 성립한다. 찰리의 머리와 몸통도 연결되어 있다. 오른쪽 전경의 큰 소매와 장갑 부분은 심하게 잘려 관절 연결을 전부 확인할 수 없지만, 별도로 떠 있는 물체라고 단정할 근거는 없다. 두 인물의 발은 프레임 밖이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.911,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.661,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 세 번째 피사체(화면 우측 근경의 기계 어깨)가 추가됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 661
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 프레이밍을 정확히 따랐으며 앰버의 머리 부상과 눈물 등 디테일을 훌륭하게 묘사했습니다."
   },
   {
    "label": "A",
    "score": 661,
    "verdict_ko": "요구된 클로즈업을 무시하고 프레임이 너무 넓으며, 우측 근경에 프롬프트에 없는 피사체가 추가된 치명적 오류가 있습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 세 번째 피사체(화면 우측 근경의 기계 어깨)가 추가됨."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_coastal_search_hill_7092b9.png",
    "asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-8364-797f-a512-ab5bb77e16d7",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh7__bgfirst_bg.png",
   "bg_asset_id": "0cca7227-2e69-4db8-ab19-8acea477fee5",
   "bg_record_key": "S71sh7::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "coastal_search_hill",
   "groupbg_asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C06"
  ]
 },
 "S71sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:21:36.630560+00:00",
  "fingerprint": "69f6e79039d834494015829cc327bb4dd27206fe9c330386df1337165a74d162",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S71sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S71sh7_sel.png",
  "source_sha256": "c9cab391707599c136a296429842032cb8b56fc8a9beab9a275685c3e4936d5d",
  "file": "S71sh7_cine.png",
  "staged_sha256": "009481222db541f1c22be27b229bc551577b45c274ba1d4d21f7fede451fbdef",
  "latency_ms": 10488
 },
 "S71sh17::confined_fp_apt": {
  "applies": true,
  "reason_ko": "트럭 조종석(운전석 및 조수석) 내부에서 인물들이 어느 자리에 앉아 앞 유리창 밖을 바라보고 있는지 정확한 위치 배치가 중요하기 때문입니다.",
  "input_fingerprint": "824462859e254c68"
 },
 "S71sh17::signage": {
  "fp": "4210a21e1515e6aa",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::3f20774e773a": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_3f20774e773a.png",
  "place_text": "Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield.",
  "input_fingerprint": "a815d3660a9511db"
 },
 "S71sh17::confined_fp": {
  "reads": {
   "controls": "The steering wheel is attached to the driver's station on the left.",
   "mirrors": "No mirrors or reflective surfaces are depicted in the diagram.",
   "camera": "The camera is positioned in the bottom center behind the seats, pointing straight forward toward the windshield.",
   "occupants": "현우 occupies the driver seat on the left, and 앰버 occupies the passenger seat on the right."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned in the lower center of the truck cab, facing straight forward. In the near left foreground, 현우 occupies the driver seat, facing away from the camera toward the front. The steering wheel is located at his station on the left side of the dashboard. In the near right foreground, 앰버 occupies the passenger seat, also facing away from the camera. The center of the frame remains an unobstructed viewing gap between the two front seats. In the far center, the wide windshield spans the cab, revealing the outside violet field in the upper background. No mirrors or reflective surfaces are present in this view.",
  "fixed": false,
  "input_fingerprint": "d5debdc1a7c0e8d2"
 },
 "S71sh17": {
  "input_fingerprint": "f4304a648aea3567",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 트럭 앞 유리창 너머로 끝없이 펼쳐진 보라색 제비꽃 밭을 멍하니 바라보는 현우와 앰버의 뒷모습.\n\nLOCATION (lock): Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed windshield view of the violet field in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Truck windshield (The violet field is visible through it) — Seen from inside the cab, with the field beyond the glass rather than reflected on it; used as Creates an unobstructed central view between the occupants; Violet field (Extensively covered in blooming purple violets); used as Provides the shared distant object of attention and the visual payoff beyond the cab; Truck seats (Occupied by 현우 and 앰버) — Their rear portions border the central viewing gap; used as Establishes the interior depth and supports the two separated foreground figures.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight balances the cab and exterior sufficiently to retain both backs while allowing the purple violets to provide the scene's richer color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An extensive field of purple violets stretches ahead, with Ji Dong-hyun's research institute visible in the distance. Charlie retains his damaged bodywork, chest opening and B-200's chest component. 현우: He is at the wheel after being startled awake from drowsy driving, with his earlier injuries still present. 앰버: She is riding in the truck after wiping away her tears, pointing forward. Her earlier head injury has no stated treatment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the lower center of the truck cab, facing straight forward. In the near left foreground, 현우 occupies the driver seat, facing away from the camera toward the front. The steering wheel is located at his station on the left side of the dashboard. In the near right foreground, 앰버 occupies the passenger seat, also facing away from the camera. The center of the frame remains an unobstructed viewing gap between the two front seats. In the far center, the wide windshield spans the cab, revealing the outside violet field in the upper background. No mirrors or reflective surfaces are present in this view.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 트럭 앞 유리창 너머로 끝없이 펼쳐진 보라색 제비꽃 밭을 멍하니 바라보는 현우와 앰버의 뒷모습.\n\nLOCATION (lock): Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight balances the cab and exterior sufficiently to retain both backs while allowing the purple violets to provide the scene's richer color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An extensive field of purple violets stretches ahead, with Ji Dong-hyun's research institute visible in the distance. Charlie retains his damaged bodywork, chest opening and B-200's chest component. 현우: He is at the wheel after being startled awake from drowsy driving, with his earlier injuries still present. 앰버: She is riding in the truck after wiping away her tears, pointing forward. Her earlier head injury has no stated treatment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the lower center of the truck cab, facing straight forward. In the near left foreground, 현우 occupies the driver seat, facing away from the camera toward the front. The steering wheel is located at his station on the left side of the dashboard. In the near right foreground, 앰버 occupies the passenger seat, also facing away from the camera. The center of the frame remains an unobstructed viewing gap between the two front seats. In the far center, the wide windshield spans the cab, revealing the outside violet field in the upper background. No mirrors or reflective surfaces are present in this view.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 트럭 앞 유리창 너머로 끝없이 펼쳐진 보라색 제비꽃 밭을 멍하니 바라보는 현우와 앰버의 뒷모습.\n\nLOCATION (lock): Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight balances the cab and exterior sufficiently to retain both backs while allowing the purple violets to provide the scene's richer color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An extensive field of purple violets stretches ahead, with Ji Dong-hyun's research institute visible in the distance. Charlie retains his damaged bodywork, chest opening and B-200's chest component. 현우: He is at the wheel after being startled awake from drowsy driving, with his earlier injuries still present. 앰버: She is riding in the truck after wiping away her tears, pointing forward. Her earlier head injury has no stated treatment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "앰버",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "앰버",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 앰버 모두 트럭 앞유리창 너머의 보라색 제비꽃 밭을 향해 시선을 두고 있으며, 앰버는 왼팔을 뻗어 앞을 가리키고 있습니다.",
    "built_space": "트럭 운전석 내부에 현우와 앰버가 좌우 좌석에 앉아 있으며, 대시보드와 앞유리창 너머로 외부 풍경이 넓게 보입니다.",
    "entities": "현우는 헝클어진 검은 머리에 흙이 묻은 회색 셔츠를 입고 상처가 보이며, 앰버는 금발 머리에 파란색 셔츠와 갈색 멜빵바지를 입고 있습니다.",
    "hard_violations": [],
    "physics": "두 인물 모두 시트에 안정적으로 앉아 있으며, 현우의 오른손은 기어 스틱 위에 얹혀 있고 앰버의 뻗은 팔은 어깨에 의해 지탱됩니다."
   },
   {
    "label": "B",
    "direction": "현우와 앰버 모두 창밖 정면의 보라색 밭과 먼 건물을 향하고 있으나, 앰버는 아무것도 가리키지 않습니다.",
    "built_space": "트럭 내부로 운전석, 조수석, 금이 간 대시보드와 스티어링 휠이 배치되어 있으며, 앞유리를 통해 바깥 풍경이 보입니다.",
    "entities": "현우는 검은 머리와 회색 셔츠를 입고 상처가 보이며, 앰버는 금발 머리와 멜빵바지를 입고 있습니다. 멀리 연구소로 보이는 건물이 있습니다.",
    "hard_violations": [],
    "physics": "현우는 스티어링 휠을 오른손으로 쥐고 좌석에 앉아 있으며, 앰버는 양팔을 내린 채 시트에 앉아 체중을 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "앰버가 앞으로 가리키는 동작을 정확히 표현했으며, 프레임과 차량 내부의 배치가 프롬프트의 지시를 잘 따르고 있습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "새벽 시간대와 멀리 보이는 연구소 건물은 잘 묘사되었으나, 앰버가 손을 뻗어 가리키는 핵심 동작이 완전히 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버 모두 트럭 앞유리창 너머의 보라색 제비꽃 밭을 향해 시선을 두고 있으며, 앰버는 왼팔을 뻗어 앞을 가리키고 있습니다.",
        "built_space": "트럭 운전석 내부에 현우와 앰버가 좌우 좌석에 앉아 있으며, 대시보드와 앞유리창 너머로 외부 풍경이 넓게 보입니다.",
        "entities": "현우는 헝클어진 검은 머리에 흙이 묻은 회색 셔츠를 입고 상처가 보이며, 앰버는 금발 머리에 파란색 셔츠와 갈색 멜빵바지를 입고 있습니다.",
        "hard_violations": [],
        "physics": "두 인물 모두 시트에 안정적으로 앉아 있으며, 현우의 오른손은 기어 스틱 위에 얹혀 있고 앰버의 뻗은 팔은 어깨에 의해 지탱됩니다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버 모두 창밖 정면의 보라색 밭과 먼 건물을 향하고 있으나, 앰버는 아무것도 가리키지 않습니다.",
        "built_space": "트럭 내부로 운전석, 조수석, 금이 간 대시보드와 스티어링 휠이 배치되어 있으며, 앞유리를 통해 바깥 풍경이 보입니다.",
        "entities": "현우는 검은 머리와 회색 셔츠를 입고 상처가 보이며, 앰버는 금발 머리와 멜빵바지를 입고 있습니다. 멀리 연구소로 보이는 건물이 있습니다.",
        "hard_violations": [],
        "physics": "현우는 스티어링 휠을 오른손으로 쥐고 좌석에 앉아 있으며, 앰버는 양팔을 내린 채 시트에 앉아 체중을 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "앰버가 앞으로 가리키는 동작을 정확히 표현했으며, 프레임과 차량 내부의 배치가 프롬프트의 지시를 잘 따르고 있습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "새벽 시간대와 멀리 보이는 연구소 건물은 잘 묘사되었으나, 앰버가 손을 뻗어 가리키는 핵심 동작이 완전히 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버 모두 트럭 앞유리창 너머의 보라색 제비꽃 밭을 향해 시선을 두고 있으며, 앰버는 왼팔을 뻗어 앞을 가리키고 있습니다.",
        "built_space": "트럭 운전석 내부에 현우와 앰버가 좌우 좌석에 앉아 있으며, 대시보드와 앞유리창 너머로 외부 풍경이 넓게 보입니다.",
        "entities": "현우는 헝클어진 검은 머리에 흙이 묻은 회색 셔츠를 입고 상처가 보이며, 앰버는 금발 머리에 파란색 셔츠와 갈색 멜빵바지를 입고 있습니다.",
        "hard_violations": [],
        "physics": "두 인물 모두 시트에 안정적으로 앉아 있으며, 현우의 오른손은 기어 스틱 위에 얹혀 있고 앰버의 뻗은 팔은 어깨에 의해 지탱됩니다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버 모두 창밖 정면의 보라색 밭과 먼 건물을 향하고 있으나, 앰버는 아무것도 가리키지 않습니다.",
        "built_space": "트럭 내부로 운전석, 조수석, 금이 간 대시보드와 스티어링 휠이 배치되어 있으며, 앞유리를 통해 바깥 풍경이 보입니다.",
        "entities": "현우는 검은 머리와 회색 셔츠를 입고 상처가 보이며, 앰버는 금발 머리와 멜빵바지를 입고 있습니다. 멀리 연구소로 보이는 건물이 있습니다.",
        "hard_violations": [],
        "physics": "현우는 스티어링 휠을 오른손으로 쥐고 좌석에 앉아 있으며, 앰버는 양팔을 내린 채 시트에 앉아 체중을 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 사람의 뒷모습 사이로 새벽의 광활한 꽃밭과 먼 연구소가 보이는 와이드 구도가 가장 충실하지만, 앰버가 앞을 가리키는 동작은 빠졌다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "앰버의 전방 지시와 좌석 배치는 맞지만, 가까운 꽃들이 앞유리를 대부분 채워 끝없는 원경과 새벽의 인상이 약하고 먼 연구소도 확인되지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버 모두 카메라에 등을 보이고 앞유리 너머 보라색 꽃밭을 향해 머리를 두고 있다. 눈은 보이지 않아 정확한 시선은 확인할 수 없지만 몸과 머리의 방향은 공동의 관찰 대상에 맞는다. 현우의 오른손은 운전대를 잡고 있으며, 앰버의 보이는 팔은 내려가 있어 전방을 가리키라는 유지 상태는 구현되지 않았다.",
        "built_space": "운전실 뒤쪽 중앙에서 앞을 보는 구도다. 앞좌석 두 개에 현우는 왼쪽, 앰버는 오른쪽으로 각각 앉아 있고, 등받이 두 개가 중앙의 열린 시야를 둘러싼다. 앞유리 한 면, 왼쪽 운전대 하나, 대시보드 하나, 와이퍼 두 개, 양쪽 외부 거울 두 개, 위쪽 햇빛가리개 두 개가 보인다. 중앙 바닥에는 변속 레버와 주차브레이크 레버가 있다. 꽃밭은 유리에 반사된 것이 아니라 유리 밖으로 이어지며, 거울에 보이는 차체 일부도 부자연스럽지 않다.",
        "entities": "인물은 두 명뿐이다. 현우의 검은 머리, 젊은 남성 체격, 회색 셔츠와 올리브색 바지는 참고 이미지와 부합하며 목 뒤에 상처가 보인다. 앰버는 작은 아동 체격에 금발이며 남색 상의와 갈색 멜빵옷을 입었다. 두 사람의 얼굴이 가려져 정확한 나이와 민족적 외모는 검증할 수 없다. 앰버 참고 이미지의 머리 위 고글형 장비는 없고, 머리 상처나 처치는 확인되지 않는다. 낡은 트럭, 넓은 보라색 꽃밭, 지평선의 작은 연구소형 건물이 보이며 낮은 따뜻한 빛은 새벽에 부합한다. 개별 꽃의 종은 확정하기 어렵다. 찰리나 추가 인물은 없고, 읽을 수 있는 글자도 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 골반과 허벅지는 각자 좌석 방석에 지지되고 등은 등받이 앞에 놓여 있다. 현우의 오른손은 운전대 테두리를 실제로 쥐고 있으며 팔과 어깨의 연결도 자연스럽다. 앰버의 내려간 팔과 앉은 자세에도 부유하거나 지지되지 않는 부분은 없다. 대시보드와 레버는 차체에 고정되어 있고 꽃은 지면에서 자란다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버의 머리와 몸은 앞유리 밖 꽃밭을 향한다. 앰버는 왼팔을 뻗어 검지로 앞유리 중앙 부근의 꽃밭을 가리키므로 전방 지시 동작이 명확하다. 특정 먼 건물을 가리키는지는 확인되지 않는다. 현우의 보이는 오른손은 운전대가 아니라 중앙 변속 레버 쪽에 놓여 있다.",
        "built_space": "카메라는 두 앞좌석 뒤의 중앙에 있고 현우는 왼쪽 운전석, 앰버는 오른쪽 조수석에 앉아 있다. 등받이 두 개 사이로 앞유리 한 면과 대시보드가 보이며 운전대는 왼쪽에 하나 있다. 와이퍼 두 개, 햇빛가리개 두 개, 양쪽 외부 거울 두 개와 오른쪽 위의 원형 보조 거울 하나가 보인다. 중앙에는 변속 레버와 수납 트레이가 있다. 좌석과 조작부의 관계는 가능하지만 꽃밭이 앞유리 거의 전체를 차지해 지평선과 먼 공간의 여유가 A보다 적다. 꽃밭은 반사가 아니라 차 밖의 풍경으로 묘사된다.",
        "entities": "현우와 앰버 두 명만 등장한다. 현우의 헝클어진 검은 머리, 젊은 남성 체격, 해진 회색 셔츠와 올리브색 바지는 참고와 맞고 목덜미에 상처 흔적이 있다. 앰버의 금발, 아동 체격, 남색 반소매와 갈색 멜빵옷도 부합한다. 얼굴을 볼 수 없어 세부 정체성은 확인할 수 없으며, 머리 위 고글형 장비는 빠져 있다. 앰버의 머리 상처나 처치는 확인되지 않는다. 낡은 트럭과 빽빽한 보라색 꽃밭은 있지만 먼 연구소는 식별되지 않는다. 꽃송이가 크게 드러나며 정확한 제비꽃 종류인지는 불확실하다. 빛은 흐린 주간처럼 보여 새벽의 시간 단서가 약하다. 추가 인물이나 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 각자의 좌석 방석에 앉아 체중을 싣고 있다. 앰버의 지시하는 팔은 어깨에서 이어져 팔꿈치를 조금 굽힌 채 뻗어 있어 가능한 동작이다. 현우의 오른손은 중앙 변속 레버 상단에 놓여 있고 팔은 자연스럽게 연결된다. 보조 거울은 외부 지지대에 연결되어 있으며, 지지 없이 떠 있는 사람이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 사람의 뒷모습 사이로 새벽의 광활한 꽃밭과 먼 연구소가 보이는 와이드 구도가 가장 충실하지만, 앰버가 앞을 가리키는 동작은 빠졌다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "앰버의 전방 지시와 좌석 배치는 맞지만, 가까운 꽃들이 앞유리를 대부분 채워 끝없는 원경과 새벽의 인상이 약하고 먼 연구소도 확인되지 않는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 앰버 모두 카메라에 등을 보이고 앞유리 너머 보라색 꽃밭을 향해 머리를 두고 있다. 눈은 보이지 않아 정확한 시선은 확인할 수 없지만 몸과 머리의 방향은 공동의 관찰 대상에 맞는다. 현우의 오른손은 운전대를 잡고 있으며, 앰버의 보이는 팔은 내려가 있어 전방을 가리키라는 유지 상태는 구현되지 않았다.",
        "built_space": "운전실 뒤쪽 중앙에서 앞을 보는 구도다. 앞좌석 두 개에 현우는 왼쪽, 앰버는 오른쪽으로 각각 앉아 있고, 등받이 두 개가 중앙의 열린 시야를 둘러싼다. 앞유리 한 면, 왼쪽 운전대 하나, 대시보드 하나, 와이퍼 두 개, 양쪽 외부 거울 두 개, 위쪽 햇빛가리개 두 개가 보인다. 중앙 바닥에는 변속 레버와 주차브레이크 레버가 있다. 꽃밭은 유리에 반사된 것이 아니라 유리 밖으로 이어지며, 거울에 보이는 차체 일부도 부자연스럽지 않다.",
        "entities": "인물은 두 명뿐이다. 현우의 검은 머리, 젊은 남성 체격, 회색 셔츠와 올리브색 바지는 참고 이미지와 부합하며 목 뒤에 상처가 보인다. 앰버는 작은 아동 체격에 금발이며 남색 상의와 갈색 멜빵옷을 입었다. 두 사람의 얼굴이 가려져 정확한 나이와 민족적 외모는 검증할 수 없다. 앰버 참고 이미지의 머리 위 고글형 장비는 없고, 머리 상처나 처치는 확인되지 않는다. 낡은 트럭, 넓은 보라색 꽃밭, 지평선의 작은 연구소형 건물이 보이며 낮은 따뜻한 빛은 새벽에 부합한다. 개별 꽃의 종은 확정하기 어렵다. 찰리나 추가 인물은 없고, 읽을 수 있는 글자도 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 골반과 허벅지는 각자 좌석 방석에 지지되고 등은 등받이 앞에 놓여 있다. 현우의 오른손은 운전대 테두리를 실제로 쥐고 있으며 팔과 어깨의 연결도 자연스럽다. 앰버의 내려간 팔과 앉은 자세에도 부유하거나 지지되지 않는 부분은 없다. 대시보드와 레버는 차체에 고정되어 있고 꽃은 지면에서 자란다."
       },
       {
        "label": "A",
        "direction": "현우와 앰버의 머리와 몸은 앞유리 밖 꽃밭을 향한다. 앰버는 왼팔을 뻗어 검지로 앞유리 중앙 부근의 꽃밭을 가리키므로 전방 지시 동작이 명확하다. 특정 먼 건물을 가리키는지는 확인되지 않는다. 현우의 보이는 오른손은 운전대가 아니라 중앙 변속 레버 쪽에 놓여 있다.",
        "built_space": "카메라는 두 앞좌석 뒤의 중앙에 있고 현우는 왼쪽 운전석, 앰버는 오른쪽 조수석에 앉아 있다. 등받이 두 개 사이로 앞유리 한 면과 대시보드가 보이며 운전대는 왼쪽에 하나 있다. 와이퍼 두 개, 햇빛가리개 두 개, 양쪽 외부 거울 두 개와 오른쪽 위의 원형 보조 거울 하나가 보인다. 중앙에는 변속 레버와 수납 트레이가 있다. 좌석과 조작부의 관계는 가능하지만 꽃밭이 앞유리 거의 전체를 차지해 지평선과 먼 공간의 여유가 A보다 적다. 꽃밭은 반사가 아니라 차 밖의 풍경으로 묘사된다.",
        "entities": "현우와 앰버 두 명만 등장한다. 현우의 헝클어진 검은 머리, 젊은 남성 체격, 해진 회색 셔츠와 올리브색 바지는 참고와 맞고 목덜미에 상처 흔적이 있다. 앰버의 금발, 아동 체격, 남색 반소매와 갈색 멜빵옷도 부합한다. 얼굴을 볼 수 없어 세부 정체성은 확인할 수 없으며, 머리 위 고글형 장비는 빠져 있다. 앰버의 머리 상처나 처치는 확인되지 않는다. 낡은 트럭과 빽빽한 보라색 꽃밭은 있지만 먼 연구소는 식별되지 않는다. 꽃송이가 크게 드러나며 정확한 제비꽃 종류인지는 불확실하다. 빛은 흐린 주간처럼 보여 새벽의 시간 단서가 약하다. 추가 인물이나 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 각자의 좌석 방석에 앉아 체중을 싣고 있다. 앰버의 지시하는 팔은 어깨에서 이어져 팔꿈치를 조금 굽힌 채 뻗어 있어 가능한 동작이다. 현우의 오른손은 중앙 변속 레버 상단에 놓여 있고 팔은 자연스럽게 연결된다. 보조 거울은 외부 지지대에 연결되어 있으며, 지지 없이 떠 있는 사람이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.875,
    "B": 1.714
   },
   "adjusted": {
    "A": 1.875,
    "B": 1.714
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1875,
   "B": 1714
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1875,
    "verdict_ko": "앰버가 앞으로 가리키는 동작을 정확히 표현했으며, 프레임과 차량 내부의 배치가 프롬프트의 지시를 잘 따르고 있습니다."
   },
   {
    "label": "B",
    "score": 1714,
    "verdict_ko": "새벽 시간대와 멀리 보이는 연구소 건물은 잘 묘사되었으나, 앰버가 손을 뻗어 가리키는 핵심 동작이 완전히 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh17_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "앰버",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-882e-79cc-9ad9-a0f8e8d4dcd4",
  "confined_fp": {
   "base_key": "confinedfp::3f20774e773a",
   "apt_reason": "트럭 조종석(운전석 및 조수석) 내부에서 인물들이 어느 자리에 앉아 앞 유리창 밖을 바라보고 있는지 정확한 위치 배치가 중요하기 때문입니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S71sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:22:49.889732+00:00",
  "fingerprint": "966c3e4944ac8686dd0fe5411e1c191b22464ffb32058ef72864815d4be62524",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S71sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S71sh17_sel.png",
  "source_sha256": "3d8b86336381b309ce21c7134053385080f46b4c7c8d0b13983d3cd80b998710",
  "file": "S71sh17_cine.png",
  "staged_sha256": "6a81a909107a262673d2c920b9d772c7af6902c0326eec1d42c236788c0bc65e",
  "latency_ms": 10541
 },
 "S72sh38::signage": {
  "fp": "3342b1e6949800e8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S72sh38": {
  "input_fingerprint": "280ca32f5165de8e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 푸른빛에 휩싸인 채 공중으로 번쩍 치켜 올려진 윤성찬의 자동차 공중 찰나.\n\nLOCATION (lock): Above the outdoor vehicle pursuit route beside the coastal research facility's violet fields, where the pursuing car is lifted into the air. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Entire airborne car with visible space beneath in the upper-left of the frame, midground; Truck carrying 찰리 farther ahead in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 윤성찬's car (Airborne at the peak of its lift, before the crash) — Its left flank, rear quarter, and part of the underside are visible from the low camera; used as The complete vehicle and the clear space beneath it establish the physical reversal of the pursuit; 현우's truck (Ahead of the pursuing car with 찰리 aboard) — Seen obliquely from behind along the established pursuit direction; used as Provides the distant source of the action and preserves the vehicle-to-vehicle geography; Violet field (The pursuit is taking place among the violets); used as A lower strip of terrain supplies a stable reference for the car's elevation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established sunset ambience is interrupted by the blue light emitted from 찰리's chest and enveloping the airborne car, with enough tonal separation to read its outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The pursuit continues through the violet-field area at dusk, with the pursuing car lifted into the air before its rollover impact. Charlie's chest opening exposes the interior, and his chest ring is active despite the accumulated body damage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 푸른빛에 휩싸인 채 공중으로 번쩍 치켜 올려진 윤성찬의 자동차 공중 찰나.\n\nLOCATION (lock): Above the outdoor vehicle pursuit route beside the coastal research facility's violet fields, where the pursuing car is lifted into the air. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Entire airborne car with visible space beneath in the upper-left of the frame, midground; Truck carrying 찰리 farther ahead in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 윤성찬's car (Airborne at the peak of its lift, before the crash) — Its left flank, rear quarter, and part of the underside are visible from the low camera; used as The complete vehicle and the clear space beneath it establish the physical reversal of the pursuit; 현우's truck (Ahead of the pursuing car with 찰리 aboard) — Seen obliquely from behind along the established pursuit direction; used as Provides the distant source of the action and preserves the vehicle-to-vehicle geography; Violet field (The pursuit is taking place among the violets); used as A lower strip of terrain supplies a stable reference for the car's elevation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established sunset ambience is interrupted by the blue light emitted from 찰리's chest and enveloping the airborne car, with enough tonal separation to read its outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The pursuit continues through the violet-field area at dusk, with the pursuing car lifted into the air before its rollover impact. Charlie's chest opening exposes the interior, and his chest ring is active despite the accumulated body damage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 푸른빛에 휩싸인 채 공중으로 번쩍 치켜 올려진 윤성찬의 자동차 공중 찰나.\n\nLOCATION (lock): Above the outdoor vehicle pursuit route beside the coastal research facility's violet fields, where the pursuing car is lifted into the air. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Entire airborne car with visible space beneath in the upper-left of the frame, midground; Truck carrying 찰리 farther ahead in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 윤성찬's car (Airborne at the peak of its lift, before the crash) — Its left flank, rear quarter, and part of the underside are visible from the low camera; used as The complete vehicle and the clear space beneath it establish the physical reversal of the pursuit; 현우's truck (Ahead of the pursuing car with 찰리 aboard) — Seen obliquely from behind along the established pursuit direction; used as Provides the distant source of the action and preserves the vehicle-to-vehicle geography; Violet field (The pursuit is taking place among the violets); used as A lower strip of terrain supplies a stable reference for the car's elevation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established sunset ambience is interrupted by the blue light emitted from 찰리's chest and enveloping the airborne car, with enough tonal separation to read its outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The pursuit continues through the violet-field area at dusk, with the pursuing car lifted into the air before its rollover impact. Charlie's chest opening exposes the interior, and his chest ring is active despite the accumulated body damage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "푸른빛의 광선이 찰리의 가슴이 아닌 오른손에서 뿜어져 나와 자동차를 향함.",
    "built_space": "레퍼런스와 동일한 도로. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
    "entities": "공중에 띄워진 자동차 1대, 박스 형태의 트럭, 트럭 짐칸에 앉아있는 찰리.",
    "hard_violations": [],
    "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 짐칸에 앉아 지탱됨."
   },
   {
    "label": "B",
    "direction": "푸른빛의 광선이 찰리의 가슴 중앙에서 뿜어져 나와 자동차 전면부를 향함.",
    "built_space": "레퍼런스와 동일한 해질녘 도로와 우측 들판. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
    "entities": "공중에 띄워진 자동차 1대, 소형 평판 트럭, 짐칸에 앉은 찰리(가슴 부위 링 활성화).",
    "hard_violations": [],
    "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 화물칸에 앉아 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 가슴에서 빛이 뿜어지는 설정을 정확히 연출하고, 요구된 넓은 샷의 구도와 피사체 배치를 충실히 따름."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "푸른빛이 가슴이 아닌 손에서 발사되어 '가슴에서 뿜어진 푸른빛'이라는 핵심 묘사를 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "푸른빛의 광선이 찰리의 가슴이 아닌 오른손에서 뿜어져 나와 자동차를 향함.",
        "built_space": "레퍼런스와 동일한 도로. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 박스 형태의 트럭, 트럭 짐칸에 앉아있는 찰리.",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 짐칸에 앉아 지탱됨."
       },
       {
        "label": "B",
        "direction": "푸른빛의 광선이 찰리의 가슴 중앙에서 뿜어져 나와 자동차 전면부를 향함.",
        "built_space": "레퍼런스와 동일한 해질녘 도로와 우측 들판. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 소형 평판 트럭, 짐칸에 앉은 찰리(가슴 부위 링 활성화).",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 화물칸에 앉아 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 가슴에서 빛이 뿜어지는 설정을 정확히 연출하고, 요구된 넓은 샷의 구도와 피사체 배치를 충실히 따름."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "푸른빛이 가슴이 아닌 손에서 발사되어 '가슴에서 뿜어진 푸른빛'이라는 핵심 묘사를 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "푸른빛의 광선이 찰리의 가슴이 아닌 오른손에서 뿜어져 나와 자동차를 향함.",
        "built_space": "레퍼런스와 동일한 도로. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 박스 형태의 트럭, 트럭 짐칸에 앉아있는 찰리.",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 짐칸에 앉아 지탱됨."
       },
       {
        "label": "B",
        "direction": "푸른빛의 광선이 찰리의 가슴 중앙에서 뿜어져 나와 자동차 전면부를 향함.",
        "built_space": "레퍼런스와 동일한 해질녘 도로와 우측 들판. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 소형 평판 트럭, 짐칸에 앉은 찰리(가슴 부위 링 활성화).",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 화물칸에 앉아 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "좌상단 공중 차량과 우측 후방 트럭의 와이드 배치, 가슴에서 차량으로 이어지는 청색광이 정확하지만 차량은 요청한 왼쪽 대신 오른쪽 측면을 드러낸다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "차량이 중경보다 전경에서 지나치게 크게 보이고, 청색광의 출발점도 가슴보다 뻗은 손으로 읽혀 핵심 구도와 작용 관계가 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 차량 모두 화면 오른쪽 안쪽으로 굽는 도로를 따라 전진하는 방향이며 트럭이 앞선다. 찰리는 트럭 뒤쪽에서 추격 차량을 향해 몸과 얼굴을 돌리고 있다. 청색 광선은 찰리의 가슴 고리에서 승용차 앞부분으로 정확히 연결된다. 차량의 뒤와 오른쪽 측면, 하부가 보여 요청된 왼쪽 측면과는 다르다.",
        "built_space": "굽은 포장도로 하나, 오른쪽의 열린 콘크리트 배수로 하나, 도로 양옆 꽃밭, 우측 전주열과 굽이의 가드레일이 참조 장소와 대응한다. 승용차 전체는 좌상단에 있고 아래로 도로와 빈 공간이 보인다. 소형 적재함 트럭 한 대는 우측의 더 먼 곳에 있으며, 찰리는 적재함 안에 자리한다. 배경을 과도하게 확대하지 않았다.",
        "entities": "손상된 검은 승용차 한 대, 소형 트럭 한 대, 찰리로 읽히는 성인 남성형 인물 한 명이 보인다. 찰리의 손상된 몸통과 청색 가슴 고리는 식별되지만 얼굴의 연령·민족적 특징 및 가슴 내부의 세부는 거리상 확정하기 어렵다. 윤성찬과 현우는 창 안에서 식별되지 않으며 다른 인물은 추가되지 않았다. 노을과 보라·분홍 꽃밭은 유지된다. 차체의 작은 배지는 있으나 글자는 명확히 판독되지 않는다.",
        "hard_violations": [],
        "physics": "승용차의 바퀴는 모두 도로에서 떨어져 있고 하부와 지면 사이가 분명하다. 가슴에서 차량까지 연결된 청색광이 요청된 들어 올림의 작용원으로 제시되므로 원인 없는 부유는 아니다. 차량은 아직 충돌하지 않은 상태다. 트럭은 바퀴로 도로에 지지되고 찰리의 하체는 적재함 안에 있어 적재함의 지지를 받는 배치다. 다만 청색광은 차 전체를 감싸기보다 앞부분에 집중된다."
       },
       {
        "label": "B",
        "direction": "트럭은 굽은 도로의 오른쪽 안쪽으로 앞서가고 승용차도 같은 방향을 향한다. 찰리는 뒤의 승용차를 바라보며 한 팔을 뻗는다. 청색광은 승용차에 닿지만 트럭 쪽 끝이 가슴 고리보다 뻗은 손에 연결되어 보인다. 승용차는 뒤와 오른쪽 측면, 넓은 하부를 드러내므로 요청된 왼쪽 측면은 아니다.",
        "built_space": "포장도로 하나와 우측 콘크리트 배수로 하나, 양옆 꽃밭, 전주열과 먼 가드레일이 참조 장소를 유지한다. 찰리는 뒤가 열린 상자형 트럭의 적재함 문턱에 앉아 있다. 승용차 전체와 그 아래 빈 공간은 확보되지만 차가 화면 왼쪽 대부분을 차지하며 중앙 높이까지 내려와, 지정된 좌상단 중경보다 큰 전경 피사체로 읽힌다.",
        "entities": "낡고 손상된 회색 승용차 한 대, 상자형 화물 트럭 한 대, 성인 남성 한 명이 보인다. 남성은 찰리 역할로 읽히며 손상된 상의와 몸통 중앙의 청색 고리가 있지만 노출된 가슴 내부는 분명하지 않다. 얼굴의 정확한 연령과 민족적 특징은 확정하기 어렵다. 윤성찬과 현우는 식별되지 않고 불필요한 추가 인물도 없다. 일몰과 꽃밭은 유지되며 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "승용차는 네 바퀴가 지면에서 떨어져 있고 아래에 그림자와 빈 공간이 있다. 차량을 감싸며 트럭 쪽으로 연결되는 청색광이 들어 올림의 원인으로 제시되어 단순한 무근거 부유는 아니다. 다만 그 힘이 가슴에서 나온다는 연결은 약하다. 찰리의 엉덩이는 적재함 문턱에 지지되고 다리는 아래로 늘어져 있으며, 트럭은 바퀴로 도로에 지지된다. 충돌 전 공중 순간으로는 읽힌다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "좌상단 공중 차량과 우측 후방 트럭의 와이드 배치, 가슴에서 차량으로 이어지는 청색광이 정확하지만 차량은 요청한 왼쪽 대신 오른쪽 측면을 드러낸다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "차량이 중경보다 전경에서 지나치게 크게 보이고, 청색광의 출발점도 가슴보다 뻗은 손으로 읽혀 핵심 구도와 작용 관계가 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 차량 모두 화면 오른쪽 안쪽으로 굽는 도로를 따라 전진하는 방향이며 트럭이 앞선다. 찰리는 트럭 뒤쪽에서 추격 차량을 향해 몸과 얼굴을 돌리고 있다. 청색 광선은 찰리의 가슴 고리에서 승용차 앞부분으로 정확히 연결된다. 차량의 뒤와 오른쪽 측면, 하부가 보여 요청된 왼쪽 측면과는 다르다.",
        "built_space": "굽은 포장도로 하나, 오른쪽의 열린 콘크리트 배수로 하나, 도로 양옆 꽃밭, 우측 전주열과 굽이의 가드레일이 참조 장소와 대응한다. 승용차 전체는 좌상단에 있고 아래로 도로와 빈 공간이 보인다. 소형 적재함 트럭 한 대는 우측의 더 먼 곳에 있으며, 찰리는 적재함 안에 자리한다. 배경을 과도하게 확대하지 않았다.",
        "entities": "손상된 검은 승용차 한 대, 소형 트럭 한 대, 찰리로 읽히는 성인 남성형 인물 한 명이 보인다. 찰리의 손상된 몸통과 청색 가슴 고리는 식별되지만 얼굴의 연령·민족적 특징 및 가슴 내부의 세부는 거리상 확정하기 어렵다. 윤성찬과 현우는 창 안에서 식별되지 않으며 다른 인물은 추가되지 않았다. 노을과 보라·분홍 꽃밭은 유지된다. 차체의 작은 배지는 있으나 글자는 명확히 판독되지 않는다.",
        "hard_violations": [],
        "physics": "승용차의 바퀴는 모두 도로에서 떨어져 있고 하부와 지면 사이가 분명하다. 가슴에서 차량까지 연결된 청색광이 요청된 들어 올림의 작용원으로 제시되므로 원인 없는 부유는 아니다. 차량은 아직 충돌하지 않은 상태다. 트럭은 바퀴로 도로에 지지되고 찰리의 하체는 적재함 안에 있어 적재함의 지지를 받는 배치다. 다만 청색광은 차 전체를 감싸기보다 앞부분에 집중된다."
       },
       {
        "label": "A",
        "direction": "트럭은 굽은 도로의 오른쪽 안쪽으로 앞서가고 승용차도 같은 방향을 향한다. 찰리는 뒤의 승용차를 바라보며 한 팔을 뻗는다. 청색광은 승용차에 닿지만 트럭 쪽 끝이 가슴 고리보다 뻗은 손에 연결되어 보인다. 승용차는 뒤와 오른쪽 측면, 넓은 하부를 드러내므로 요청된 왼쪽 측면은 아니다.",
        "built_space": "포장도로 하나와 우측 콘크리트 배수로 하나, 양옆 꽃밭, 전주열과 먼 가드레일이 참조 장소를 유지한다. 찰리는 뒤가 열린 상자형 트럭의 적재함 문턱에 앉아 있다. 승용차 전체와 그 아래 빈 공간은 확보되지만 차가 화면 왼쪽 대부분을 차지하며 중앙 높이까지 내려와, 지정된 좌상단 중경보다 큰 전경 피사체로 읽힌다.",
        "entities": "낡고 손상된 회색 승용차 한 대, 상자형 화물 트럭 한 대, 성인 남성 한 명이 보인다. 남성은 찰리 역할로 읽히며 손상된 상의와 몸통 중앙의 청색 고리가 있지만 노출된 가슴 내부는 분명하지 않다. 얼굴의 정확한 연령과 민족적 특징은 확정하기 어렵다. 윤성찬과 현우는 식별되지 않고 불필요한 추가 인물도 없다. 일몰과 꽃밭은 유지되며 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "승용차는 네 바퀴가 지면에서 떨어져 있고 아래에 그림자와 빈 공간이 있다. 차량을 감싸며 트럭 쪽으로 연결되는 청색광이 들어 올림의 원인으로 제시되어 단순한 무근거 부유는 아니다. 다만 그 힘이 가슴에서 나온다는 연결은 약하다. 찰리의 엉덩이는 적재함 문턱에 지지되고 다리는 아래로 늘어져 있으며, 트럭은 바퀴로 도로에 지지된다. 충돌 전 공중 순간으로는 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.196,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.196,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1196
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "찰리의 가슴에서 빛이 뿜어지는 설정을 정확히 연출하고, 요구된 넓은 샷의 구도와 피사체 배치를 충실히 따름."
   },
   {
    "label": "A",
    "score": 1196,
    "verdict_ko": "푸른빛이 가슴이 아닌 손에서 발사되어 '가슴에서 뿜어진 푸른빛'이라는 핵심 묘사를 위반함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L248B01.png",
    "asset_id": "53477491-7f2b-45f1-b840-dd1a1a2068e0",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-8b70-7446-874f-5c44ec6ca416",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S72sh38::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:26:07.329029+00:00",
  "fingerprint": "73a5bdbdda2b6c847ba545cab0825b19cb8d0804d5f828843011a321b67bc164",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S72sh38_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S72sh38_sel.png",
  "source_sha256": "991565214fa5ad6197b6af5060bd7bbba568028ed58cae5c8770aefbc9ba4df7",
  "file": "S72sh38_cine.png",
  "staged_sha256": "ae1dfa8b9e4bf8e2e34041e0c8cf154246a159313400e7b5250bdcbe4e6a4b32",
  "latency_ms": 11542
 },
 "S72sh43::signage": {
  "fp": "3e6305d7e277d74c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::d2a0a165e0f87beb": {
  "subjects": [],
  "subject_text": "해남 지동현의 연구소 1층\n낡은 철문 안으로 펼쳐지는 비어 있는 연구소 1층. 먼지가 쌓인 바닥과 위층으로 이어지는 나선형 계단이 눈에 띈다.",
  "identity": "canonical",
  "scope_id": "L248",
  "scope_role": "location_interior",
  "scope_sha": "7488a540c22abf73"
 },
 "S72sh43::bgfirst_bg": {
  "input_fingerprint": "4cf06b55bf705dc2",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43__bgfirst_bg.png",
  "asset_id": "d805328b-fd8e-4d2d-b2d9-3aebab763862",
  "input_asset_ids": [
   "5cff0d94-8701-405a-b9fc-ab23b5fb7cba",
   "20b9995f-79f3-4c42-9a56-159746fe0f8d"
  ]
 },
 "S72sh43": {
  "input_fingerprint": "d63db2cac3955077",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape truck is speeding along the road at night. Charlie's damaged chest remains open, and his diagnostic view displays a 'CAUTION' warning when checking the injury. 앰버: She has a severe, actively bleeding wound in her side and is groaning in pain. The earlier injury to the back of her head also remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape truck is speeding along the road at night. Charlie's damaged chest remains open, and his diagnostic view displays a 'CAUTION' warning when checking the injury. 앰버: She has a severe, actively bleeding wound in her side and is groaning in pain. The earlier injury to the back of her head also remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape truck is speeding along the road at night. Charlie's damaged chest remains open, and his diagnostic view displays a 'CAUTION' warning when checking the injury. 앰버: She has a severe, actively bleeding wound in her side and is groaning in pain. The earlier injury to the back of her head also remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43__bgfirst_bg.png",
     "asset_id": "d805328b-fd8e-4d2d-b2d9-3aebab763862",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S72sh43.png",
     "asset_id": "5cff0d94-8701-405a-b9fc-ab23b5fb7cba",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1453331>",
     "asset_id": "78e0f2d2-91e2-4e2e-ba1f-9758f1be07f9",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L248B03.png",
     "asset_id": "20b9995f-79f3-4c42-9a56-159746fe0f8d",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1453331>",
     "asset_id": "78e0f2d2-91e2-4e2e-ba1f-9758f1be07f9",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "눈을 감고 고개를 약간 숙인 채 정면 아래를 향함.",
    "built_space": "트럭 조수석 내부. 창밖의 야경 불빛이 선명하여 차량이 정지해 있는 것처럼 보임.",
    "entities": "머리에 방독면을 쓴 앰버. 피 묻은 손이 옆구리가 아닌 복부 앞쪽을 덮고 있음.",
    "hard_violations": [],
    "physics": "좌석에 앉아 체중을 지탱하고 있으나 오른손의 형태가 다소 뭉개져 있음."
   },
   {
    "label": "B",
    "direction": "고통스럽게 눈을 감고 고개를 아래로 숙이고 있음.",
    "built_space": "트럭 조수석 내부. 창문 너머로 해질녘 풍경이 모션 블러와 함께 보여 주행 중임을 나타냄.",
    "entities": "금발의 앰버. 피 묻은 손으로 우측 옆구리를 강하게 움켜쥐고 있음.",
    "hard_violations": [],
    "physics": "좌석에 앉아 엉덩이와 등으로 체중을 자연스럽게 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시문대로 옆구리를 움켜쥔 포즈를 취하고 있으며, 창밖의 모션 블러를 통해 고속으로 주행 중인 트럭의 상황을 잘 표현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "상처를 움켜쥔 위치가 옆구리가 아닌 복부 중앙이며, 창밖 풍경이 정지되어 있어 달리는 트럭이라는 설정에 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "고통스럽게 눈을 감고 고개를 아래로 숙이고 있음.",
        "built_space": "트럭 조수석 내부. 창문 너머로 해질녘 풍경이 모션 블러와 함께 보여 주행 중임을 나타냄.",
        "entities": "금발의 앰버. 피 묻은 손으로 우측 옆구리를 강하게 움켜쥐고 있음.",
        "hard_violations": [],
        "physics": "좌석에 앉아 엉덩이와 등으로 체중을 자연스럽게 지탱하고 있음."
       },
       {
        "label": "A",
        "direction": "눈을 감고 고개를 약간 숙인 채 정면 아래를 향함.",
        "built_space": "트럭 조수석 내부. 창밖의 야경 불빛이 선명하여 차량이 정지해 있는 것처럼 보임.",
        "entities": "머리에 방독면을 쓴 앰버. 피 묻은 손이 옆구리가 아닌 복부 앞쪽을 덮고 있음.",
        "hard_violations": [],
        "physics": "좌석에 앉아 체중을 지탱하고 있으나 오른손의 형태가 다소 뭉개져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시문대로 옆구리를 움켜쥔 포즈를 취하고 있으며, 창밖의 모션 블러를 통해 고속으로 주행 중인 트럭의 상황을 잘 표현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "상처를 움켜쥔 위치가 옆구리가 아닌 복부 중앙이며, 창밖 풍경이 정지되어 있어 달리는 트럭이라는 설정에 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "고통스럽게 눈을 감고 고개를 아래로 숙이고 있음.",
        "built_space": "트럭 조수석 내부. 창문 너머로 해질녘 풍경이 모션 블러와 함께 보여 주행 중임을 나타냄.",
        "entities": "금발의 앰버. 피 묻은 손으로 우측 옆구리를 강하게 움켜쥐고 있음.",
        "hard_violations": [],
        "physics": "좌석에 앉아 엉덩이와 등으로 체중을 자연스럽게 지탱하고 있음."
       },
       {
        "label": "A",
        "direction": "눈을 감고 고개를 약간 숙인 채 정면 아래를 향함.",
        "built_space": "트럭 조수석 내부. 창밖의 야경 불빛이 선명하여 차량이 정지해 있는 것처럼 보임.",
        "entities": "머리에 방독면을 쓴 앰버. 피 묻은 손이 옆구리가 아닌 복부 앞쪽을 덮고 있음.",
        "hard_violations": [],
        "physics": "좌석에 앉아 체중을 지탱하고 있으나 오른손의 형태가 다소 뭉개져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "감은 눈과 피 묻은 손의 옆구리 압박은 맞지만, 무릎까지 넓힌 구도와 넓은 실내 배경이 웅크린 상체 중심의 지정 프레이밍에서 벗어나며 머리의 방독면도 빠졌다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "상체를 더 크게 담아 감은 눈, 고통스러운 웅크림, 피 묻은 손의 압박을 선명하게 연결하고 참고 의상도 잘 유지하지만, 실내 배경은 요구된 좁은 띠보다 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "눈을 꽉 감고 얼굴을 아래로 숙였으며 카메라를 바라보지 않는다. 오른손은 자신의 오른쪽 옆구리 출혈 부위를 직접 누르고, 왼손은 왼쪽 허벅지 앞에 내려와 있다. 창밖 풍경은 가로로 흐려져 이동 중인 차량을 암시하지만 진행 방향은 확정할 수 없다.",
        "built_space": "회색 천 벤치 한 개의 등받이와 좌판, 오른쪽 창문 한 개, 낡은 금속 벽판이 보인다. 아이는 벤치에 엉덩이를 놓고 등받이를 뒤로 둔 채 앞으로 기울어 있어 좌석 방향과 맞는다. 참고 장소의 재료와 주요 배치는 일치하며 중복 설비나 불가능한 반사는 없다. 다만 벤치와 창문이 화면에 넓게 드러나 상체 뒤에 비스듬한 실내 띠만 남기라는 요구와 차이가 크다.",
        "entities": "금발의 어린 여자아이 한 명만 보이고, 체격과 얼굴은 참고 인물과 대체로 부합한다. 한국계 백인 혼혈 여부는 외형만으로 확정할 수 없다. 남색 반팔, 갈색 작업 멜빵바지, 허리띠와 공구는 참고와 맞지만, 머리가 보이는 구도인데 참고의 방독면은 없다. 얼굴의 상처와 옆구리·손의 젖은 피가 보이며 후두부 상처는 머리카락에 가려 확인할 수 없다. 추가 인물이나 판독 가능한 문구는 보이지 않는다. 창밖 주황색 하늘은 일몰 조건에 맞는다.",
        "hard_violations": [],
        "physics": "체중은 벤치 좌판에 놓인 골반과 허벅지가 받친다. 몸통을 앞으로 굽히고 오른팔을 접어 옆구리를 압박하는 자세는 가능하다. 내려온 왼팔은 어깨와 팔꿈치에 자연스럽게 연결되고, 공구는 허리띠에 걸리거나 좌판에 닿아 있다. 피는 손과 옷 표면을 따라 번지고 아래로 흘러 있으며, 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "두 눈을 강하게 감고 고개를 아래로 떨어뜨렸다. 오른팔을 배 앞쪽으로 가로질러 피 묻은 오른손으로 자신의 왼쪽 옆구리에 가까운 복부를 움켜누른다. 왼팔은 무릎 쪽으로 내려가며 손은 화면 밖이다. 얼굴의 찡그림과 손의 압박이 같은 고통 동작으로 연결된다. 창밖에서 차량의 진행 방향이나 속도는 뚜렷하게 읽히지 않는다.",
        "built_space": "회색 천 벤치 한 개, 오른쪽 창문 한 개, 그 아래의 낡은 금속 패널과 뒤쪽 벽판이 보인다. 아이는 등받이 앞 좌판에 앉아 상체를 앞으로 접고 있으며 좌석 사용 방향이 자연스럽다. 참고 장소의 창틀과 천·금속 재료가 잘 유지되고 설비 중복이나 불가능한 반사는 없다. A보다 상체 비중이 크지만 왼쪽 등받이와 오른쪽 창문은 여전히 넓어서 배경을 좁은 띠로 제한한 조건에는 완전히 맞지 않는다.",
        "entities": "금발의 어린 여자아이 한 명이며 둥근 얼굴과 어린 체격이 참고 인물에 가깝다. 혼혈 배경은 외형만으로 확정할 수 없고 큰 눈의 형태는 감겨 있어 판단하기 어렵다. 머리에 올린 방독면, 남색 반팔, 갈색 멜빵 작업복이 참고와 부합한다. 얼굴의 상처, 피에 젖은 손, 옆구리 아래와 무릎 쪽 옷의 혈흔이 보인다. 상처 자체는 손과 옷에 가려지고 후두부도 보이지 않는다. 추가 인물이나 읽을 수 있는 글자는 없으며 창밖에는 일몰 끝무렵의 어두운 하늘이 보인다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 벤치 좌판에 지지되고 몸통은 그 위에서 앞으로 굽혀져 있다. 어깨를 움츠리고 팔꿈치를 접어 복부 옆을 압박하는 동작이 해부학적으로 가능하다. 방독면은 머리와 둘레의 끈에 지지된다. 손의 피는 피부 굴곡을 따르고 옷의 혈흔은 천에 스며든 형태이며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "감은 눈과 피 묻은 손의 옆구리 압박은 맞지만, 무릎까지 넓힌 구도와 넓은 실내 배경이 웅크린 상체 중심의 지정 프레이밍에서 벗어나며 머리의 방독면도 빠졌다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "상체를 더 크게 담아 감은 눈, 고통스러운 웅크림, 피 묻은 손의 압박을 선명하게 연결하고 참고 의상도 잘 유지하지만, 실내 배경은 요구된 좁은 띠보다 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "눈을 꽉 감고 얼굴을 아래로 숙였으며 카메라를 바라보지 않는다. 오른손은 자신의 오른쪽 옆구리 출혈 부위를 직접 누르고, 왼손은 왼쪽 허벅지 앞에 내려와 있다. 창밖 풍경은 가로로 흐려져 이동 중인 차량을 암시하지만 진행 방향은 확정할 수 없다.",
        "built_space": "회색 천 벤치 한 개의 등받이와 좌판, 오른쪽 창문 한 개, 낡은 금속 벽판이 보인다. 아이는 벤치에 엉덩이를 놓고 등받이를 뒤로 둔 채 앞으로 기울어 있어 좌석 방향과 맞는다. 참고 장소의 재료와 주요 배치는 일치하며 중복 설비나 불가능한 반사는 없다. 다만 벤치와 창문이 화면에 넓게 드러나 상체 뒤에 비스듬한 실내 띠만 남기라는 요구와 차이가 크다.",
        "entities": "금발의 어린 여자아이 한 명만 보이고, 체격과 얼굴은 참고 인물과 대체로 부합한다. 한국계 백인 혼혈 여부는 외형만으로 확정할 수 없다. 남색 반팔, 갈색 작업 멜빵바지, 허리띠와 공구는 참고와 맞지만, 머리가 보이는 구도인데 참고의 방독면은 없다. 얼굴의 상처와 옆구리·손의 젖은 피가 보이며 후두부 상처는 머리카락에 가려 확인할 수 없다. 추가 인물이나 판독 가능한 문구는 보이지 않는다. 창밖 주황색 하늘은 일몰 조건에 맞는다.",
        "hard_violations": [],
        "physics": "체중은 벤치 좌판에 놓인 골반과 허벅지가 받친다. 몸통을 앞으로 굽히고 오른팔을 접어 옆구리를 압박하는 자세는 가능하다. 내려온 왼팔은 어깨와 팔꿈치에 자연스럽게 연결되고, 공구는 허리띠에 걸리거나 좌판에 닿아 있다. 피는 손과 옷 표면을 따라 번지고 아래로 흘러 있으며, 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "두 눈을 강하게 감고 고개를 아래로 떨어뜨렸다. 오른팔을 배 앞쪽으로 가로질러 피 묻은 오른손으로 자신의 왼쪽 옆구리에 가까운 복부를 움켜누른다. 왼팔은 무릎 쪽으로 내려가며 손은 화면 밖이다. 얼굴의 찡그림과 손의 압박이 같은 고통 동작으로 연결된다. 창밖에서 차량의 진행 방향이나 속도는 뚜렷하게 읽히지 않는다.",
        "built_space": "회색 천 벤치 한 개, 오른쪽 창문 한 개, 그 아래의 낡은 금속 패널과 뒤쪽 벽판이 보인다. 아이는 등받이 앞 좌판에 앉아 상체를 앞으로 접고 있으며 좌석 사용 방향이 자연스럽다. 참고 장소의 창틀과 천·금속 재료가 잘 유지되고 설비 중복이나 불가능한 반사는 없다. A보다 상체 비중이 크지만 왼쪽 등받이와 오른쪽 창문은 여전히 넓어서 배경을 좁은 띠로 제한한 조건에는 완전히 맞지 않는다.",
        "entities": "금발의 어린 여자아이 한 명이며 둥근 얼굴과 어린 체격이 참고 인물에 가깝다. 혼혈 배경은 외형만으로 확정할 수 없고 큰 눈의 형태는 감겨 있어 판단하기 어렵다. 머리에 올린 방독면, 남색 반팔, 갈색 멜빵 작업복이 참고와 부합한다. 얼굴의 상처, 피에 젖은 손, 옆구리 아래와 무릎 쪽 옷의 혈흔이 보인다. 상처 자체는 손과 옷에 가려지고 후두부도 보이지 않는다. 추가 인물이나 읽을 수 있는 글자는 없으며 창밖에는 일몰 끝무렵의 어두운 하늘이 보인다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 벤치 좌판에 지지되고 몸통은 그 위에서 앞으로 굽혀져 있다. 어깨를 움츠리고 팔꿈치를 접어 복부 옆을 압박하는 동작이 해부학적으로 가능하다. 방독면은 머리와 둘레의 끈에 지지된다. 손의 피는 피부 굴곡을 따르고 옷의 혈흔은 천에 스며든 형태이며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "지시문대로 옆구리를 움켜쥔 포즈를 취하고 있으며, 창밖의 모션 블러를 통해 고속으로 주행 중인 트럭의 상황을 잘 표현했습니다."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "상처를 움켜쥔 위치가 옆구리가 아닌 복부 중앙이며, 창밖 풍경이 정지되어 있어 달리는 트럭이라는 설정에 어긋납니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L248B03.png",
    "asset_id": "20b9995f-79f3-4c42-9a56-159746fe0f8d",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453331>",
    "asset_id": "78e0f2d2-91e2-4e2e-ba1f-9758f1be07f9",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-8d18-79ef-8f94-9935a869ff7e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43__bgfirst_bg.png",
   "bg_asset_id": "d805328b-fd8e-4d2d-b2d9-3aebab763862",
   "bg_record_key": "S72sh43::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S72sh58::signage": {
  "fp": "8220f4478eefb777",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S72sh58": {
  "input_fingerprint": "0a51c5fdff880781",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문 너머 눈부신 불빛 사이로 다정하게 미소를 띤 신부의 환한 얼굴.\n\nLOCATION (lock): Outside the stalled truck's open doorway on a rain-soaked road at night, amid the rescuers' bright flashlight beams. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck doorway (Open after the truck has stopped) — Viewed diagonally from the interior toward 신부 and 라울 outside; used as A narrow edge frames the recognition and preserves the inside-to-outside relationship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Intense flashlight glare surrounds the doorway while controlled exposure preserves 신부's bright, gently smiling face and the darker foreground shoulder.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped on the road with its fuel exhausted, and its door is now open in the rain. Charlie remains inside with his accumulated body damage and exposed chest opening. 현우: He is inside the stalled truck, crying, with his arms held tightly in an embracing position. His battered, dirty appearance persists. 라울: He has returned to the stopped truck after running away to seek help. 신부: He stands outside the open truck door, still wearing his clerical collar. Powerful flashlight beams shine into the truck through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문 너머 눈부신 불빛 사이로 다정하게 미소를 띤 신부의 환한 얼굴.\n\nLOCATION (lock): Outside the stalled truck's open doorway on a rain-soaked road at night, amid the rescuers' bright flashlight beams. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck doorway (Open after the truck has stopped) — Viewed diagonally from the interior toward 신부 and 라울 outside; used as A narrow edge frames the recognition and preserves the inside-to-outside relationship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Intense flashlight glare surrounds the doorway while controlled exposure preserves 신부's bright, gently smiling face and the darker foreground shoulder.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped on the road with its fuel exhausted, and its door is now open in the rain. Charlie remains inside with his accumulated body damage and exposed chest opening. 현우: He is inside the stalled truck, crying, with his arms held tightly in an embracing position. His battered, dirty appearance persists. 라울: He has returned to the stopped truck after running away to seek help. 신부: He stands outside the open truck door, still wearing his clerical collar. Powerful flashlight beams shine into the truck through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문 너머 눈부신 불빛 사이로 다정하게 미소를 띤 신부의 환한 얼굴.\n\nLOCATION (lock): Outside the stalled truck's open doorway on a rain-soaked road at night, amid the rescuers' bright flashlight beams. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck doorway (Open after the truck has stopped) — Viewed diagonally from the interior toward 신부 and 라울 outside; used as A narrow edge frames the recognition and preserves the inside-to-outside relationship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Intense flashlight glare surrounds the doorway while controlled exposure preserves 신부's bright, gently smiling face and the darker foreground shoulder.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped on the road with its fuel exhausted, and its door is now open in the rain. Charlie remains inside with his accumulated body damage and exposed chest opening. 현우: He is inside the stalled truck, crying, with his arms held tightly in an embracing position. His battered, dirty appearance persists. 라울: He has returned to the stopped truck after running away to seek help. 신부: He stands outside the open truck door, still wearing his clerical collar. Powerful flashlight beams shine into the truck through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 트럭 내부에서 왼쪽 바깥을 향하고 있으며, 신부가 트럭 밖에서 내부의 인물을 바라보며 미소 짓고 있음.",
    "built_space": "카메라가 트럭 내부에서 바깥을 봅니다. 그러나 신부는 닫혀 있거나 반쯤 열린 창문 너머에 서 있는데, 화면 우측에 트럭 문이 따로 열려 있는 것으로 묘사되어 '열린 차 문 너머'라는 단일 출입구 구조에 심각한 모순이 있습니다.",
    "entities": "신부의 외모와 복장은 레퍼런스와 일치합니다. 전경에는 현우로 추정되는 인물의 측면이 보입니다. 그러나 배경에 프롬프트에 없는 구급대원과 구경꾼 등 다수의 군중이 임의로 추가되었습니다.",
    "hard_violations": [
     "[gemini-pro] 지정되지 않은 발명된 인물들(배경의 군중) 추가",
     "[gemini-pro] 신부가 서 있는 곳과 별도로 화면 우측에 또 다른 문이 열려 있는 물리적/구조적 모순",
     "[gpt-high] 허용된 신부·현우·라울 외에 여러 성인 구조 인물을 실제 신체와 얼굴로 추가해, 다른 사람을 등장시키지 말라는 조건을 위반했다.",
     "[gpt-high] 장면에 명시되지 않은 소방차·구급차 형태의 구조 차량 여러 대를 배경의 주요 사물로 추가했다."
    ],
    "physics": "비가 자연스럽게 내리고 빛 번짐이 사실적이나, 프롬프트에 없는 인물들이 서 있는 구조적 오류가 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 트럭 내부에서 열린 문 바깥을 대각선으로 바라보며, 바깥에 선 신부와 라울이 안쪽을 바라보고 있음.",
    "built_space": "트럭의 열린 문틈과 외벽이 정확히 보이며, 프롬프트에서 요구한 내부와 외부의 관계를 좁은 가장자리로 분리해 내는 프레이밍을 정확히 구현함.",
    "entities": "신부(미소 짓는 노안, 사제복)와 라울(어린 흑인 혼혈 소년)이 레퍼런스와 완벽히 일치하게 바깥에 서 있습니다. 전경에는 스스로를 끌어안은 현우의 뒷모습이 렌즈에 걸려 어둡게 묘사되며, 불필요한 인물은 없습니다.",
    "hard_violations": [],
    "physics": "내리는 비와 강렬한 플래시 눈부심이 사실적으로 묘사되었으며, 전경 인물(현우)의 어깨에 얹힌 손은 팔을 교차해 스스로를 안고 있는 자연스러운 해부학적 자세를 보여줍니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "배경에 프롬프트가 금지한 다수의 인물을 임의로 추가하였으며, 신부가 서 있는 창문과 별개로 열린 문이 존재하는 구조적 모순이 발생하여 실격 사유가 됩니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 프레이밍(내부에서 바깥을 향한 대각선 근접 촬영)을 완벽히 따랐으며, 명시된 세 명의 인물만을 정확히 배치하고 다정한 신부의 미소와 플래시 불빛을 훌륭하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 트럭 내부에서 왼쪽 바깥을 향하고 있으며, 신부가 트럭 밖에서 내부의 인물을 바라보며 미소 짓고 있음.",
        "built_space": "카메라가 트럭 내부에서 바깥을 봅니다. 그러나 신부는 닫혀 있거나 반쯤 열린 창문 너머에 서 있는데, 화면 우측에 트럭 문이 따로 열려 있는 것으로 묘사되어 '열린 차 문 너머'라는 단일 출입구 구조에 심각한 모순이 있습니다.",
        "entities": "신부의 외모와 복장은 레퍼런스와 일치합니다. 전경에는 현우로 추정되는 인물의 측면이 보입니다. 그러나 배경에 프롬프트에 없는 구급대원과 구경꾼 등 다수의 군중이 임의로 추가되었습니다.",
        "hard_violations": [
         "지정되지 않은 발명된 인물들(배경의 군중) 추가",
         "신부가 서 있는 곳과 별도로 화면 우측에 또 다른 문이 열려 있는 물리적/구조적 모순"
        ],
        "physics": "비가 자연스럽게 내리고 빛 번짐이 사실적이나, 프롬프트에 없는 인물들이 서 있는 구조적 오류가 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 트럭 내부에서 열린 문 바깥을 대각선으로 바라보며, 바깥에 선 신부와 라울이 안쪽을 바라보고 있음.",
        "built_space": "트럭의 열린 문틈과 외벽이 정확히 보이며, 프롬프트에서 요구한 내부와 외부의 관계를 좁은 가장자리로 분리해 내는 프레이밍을 정확히 구현함.",
        "entities": "신부(미소 짓는 노안, 사제복)와 라울(어린 흑인 혼혈 소년)이 레퍼런스와 완벽히 일치하게 바깥에 서 있습니다. 전경에는 스스로를 끌어안은 현우의 뒷모습이 렌즈에 걸려 어둡게 묘사되며, 불필요한 인물은 없습니다.",
        "hard_violations": [],
        "physics": "내리는 비와 강렬한 플래시 눈부심이 사실적으로 묘사되었으며, 전경 인물(현우)의 어깨에 얹힌 손은 팔을 교차해 스스로를 안고 있는 자연스러운 해부학적 자세를 보여줍니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "배경에 프롬프트가 금지한 다수의 인물을 임의로 추가하였으며, 신부가 서 있는 창문과 별개로 열린 문이 존재하는 구조적 모순이 발생하여 실격 사유가 됩니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 프레이밍(내부에서 바깥을 향한 대각선 근접 촬영)을 완벽히 따랐으며, 명시된 세 명의 인물만을 정확히 배치하고 다정한 신부의 미소와 플래시 불빛을 훌륭하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 트럭 내부에서 왼쪽 바깥을 향하고 있으며, 신부가 트럭 밖에서 내부의 인물을 바라보며 미소 짓고 있음.",
        "built_space": "카메라가 트럭 내부에서 바깥을 봅니다. 그러나 신부는 닫혀 있거나 반쯤 열린 창문 너머에 서 있는데, 화면 우측에 트럭 문이 따로 열려 있는 것으로 묘사되어 '열린 차 문 너머'라는 단일 출입구 구조에 심각한 모순이 있습니다.",
        "entities": "신부의 외모와 복장은 레퍼런스와 일치합니다. 전경에는 현우로 추정되는 인물의 측면이 보입니다. 그러나 배경에 프롬프트에 없는 구급대원과 구경꾼 등 다수의 군중이 임의로 추가되었습니다.",
        "hard_violations": [
         "지정되지 않은 발명된 인물들(배경의 군중) 추가",
         "신부가 서 있는 곳과 별도로 화면 우측에 또 다른 문이 열려 있는 물리적/구조적 모순"
        ],
        "physics": "비가 자연스럽게 내리고 빛 번짐이 사실적이나, 프롬프트에 없는 인물들이 서 있는 구조적 오류가 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 트럭 내부에서 열린 문 바깥을 대각선으로 바라보며, 바깥에 선 신부와 라울이 안쪽을 바라보고 있음.",
        "built_space": "트럭의 열린 문틈과 외벽이 정확히 보이며, 프롬프트에서 요구한 내부와 외부의 관계를 좁은 가장자리로 분리해 내는 프레이밍을 정확히 구현함.",
        "entities": "신부(미소 짓는 노안, 사제복)와 라울(어린 흑인 혼혈 소년)이 레퍼런스와 완벽히 일치하게 바깥에 서 있습니다. 전경에는 스스로를 끌어안은 현우의 뒷모습이 렌즈에 걸려 어둡게 묘사되며, 불필요한 인물은 없습니다.",
        "hard_violations": [],
        "physics": "내리는 비와 강렬한 플래시 눈부심이 사실적으로 묘사되었으며, 전경 인물(현우)의 어깨에 얹힌 손은 팔을 교차해 스스로를 안고 있는 자연스러운 해부학적 자세를 보여줍니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "허용된 인물만으로 차 안팎의 시선과 신부의 다정한 미소를 구현했지만, 신부 얼굴의 클로즈업 대신 넓은 어깨너머 구도가 되었다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "명시되지 않은 다수의 구조 인물과 차량을 추가한 중대 위반이 있으며, 넓어진 문·도로 구도가 신부 얼굴 클로즈업과 좁은 문틀 조건도 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 오른쪽 전경의 차 안 인물을 바라보며 웃고, 라울도 같은 실내 방향을 본다. 전경 인물은 뒤통수와 어깨를 보이며 두 사람을 향한다. 신부의 가슴 앞 손전등은 열린 출입구를 통해 카메라가 있는 실내로 빛을 보내므로 요구된 조사 방향과 맞는다.",
        "built_space": "왼쪽에 젖고 낡은 금속 문짝 하나와 경첩·잠금장치가 보이고, 그 옆으로 열린 출입구 하나가 형성된다. 신부와 라울은 문밖에, 어두운 전경 인물은 안쪽에 있어 안에서 밖을 보는 관계가 성립한다. 왼쪽 표면의 번진 빛 반사는 손전등 위치와 모순되지 않는다. 다만 문짝과 전경 어깨가 화면 대부분을 차지하고 신부는 상반신까지 보여, 좁은 문틀 가장자리를 둔 얼굴 클로즈업이 아니다.",
        "entities": "신부는 주름진 얼굴과 짧은 회흑색 머리의 한국인 노년 남성으로 보이며, 참조와 유사한 얼굴 및 검은 성직복·흰 칼라를 갖췄다. 라울은 갈색 피부와 뒤로 정돈한 머리의 어린 남자아이로 보이지만 꽁지머리 끝은 가려져 확인할 수 없다. 전경 인물의 검은 헝클어진 머리는 현우와 양립하나 얼굴·눈물·상처는 보이지 않는다. 찰리와 포옹하는 팔은 구도 밖이다. 손전등, 비, 젖은 도로가 보이고 옅은 저녁 하늘은 있으나 참조의 주황색 일몰은 약하다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "신부와 라울은 문밖에서 정상적으로 서 있는 상체 자세이며, 발은 화면 밖이라 접지점은 확인되지 않는다. 신부의 굽힌 팔이 손전등 위치로 이어지고 손 부근은 강한 눈부심에 가려져 있다. 전경 인물의 어깨와 목은 자연스럽게 연결된다. 빗물은 문 표면에 맺히고 빗줄기는 아래로 떨어지며, 근거 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "신부는 왼쪽 전경의 현우 쪽으로 고개를 기울여 미소 짓는다. 도로 중앙의 아이와 뒤쪽 사람들도 대체로 트럭 출입구를 향한다. 신부 왼쪽과 출입구 오른쪽의 강한 빛은 실내 카메라 방향으로 들어와 손전등의 조사 방향 자체는 맞는다.",
        "built_space": "왼쪽 중앙에는 창과 외판이 있는 문짝, 오른쪽에는 창과 내장 손잡이가 있는 별도 문짝이 보여 두 문짝 사이로 도로를 바라보는 구도다. 신부의 몸은 왼쪽 문짝에 상당 부분 가려지고, 아이와 여러 사람은 그 뒤 도로에 배치된다. 지정된 좁은 문틀보다 두 문짝과 도로가 훨씬 큰 비중을 차지하며, 신부 얼굴은 작다. 기존 트럭 내부의 고정 설비 연속성을 확인할 부분은 제한적이다.",
        "entities": "신부의 나이 든 한국인 남성 외모, 회흑색 머리, 검은 성직복과 흰 칼라는 참조에 대체로 부합한다. 왼쪽 현우는 검은 머리와 남색 상의를 보이지만 흐려진 옆얼굴만으로 정확한 얼굴·상처·눈물은 확인하기 어렵다. 중앙의 남색 상의 아이는 라울 역할로 읽히나 얼굴과 꽁지머리는 불분명하다. 그 외에도 여러 성인과 배경 구조 차량 최소 세 대가 추가되어 있다. 비와 젖은 노면은 구현됐지만 하늘은 주황색 일몰보다 푸른 야간에 가깝다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "허용된 신부·현우·라울 외에 여러 성인 구조 인물을 실제 신체와 얼굴로 추가해, 다른 사람을 등장시키지 말라는 조건을 위반했다.",
         "장면에 명시되지 않은 소방차·구급차 형태의 구조 차량 여러 대를 배경의 주요 사물로 추가했다."
        ],
        "physics": "도로 위 아이와 뒤쪽 인물들은 발이 노면에 닿아 있거나 다른 사람에게 가려진 정상적인 기립 자세다. 신부는 문 옆에서 상체를 기울인 자세로, 공중에 뜬 모습은 아니다. 오른쪽 손전등은 사람의 손·팔 부근에서 빛나고 왼쪽 광원은 눈부심 때문에 파지 상태가 판독되지 않는다. 젖은 노면의 광원 반사는 자연스러우며, 명백히 무지지 상태인 물체는 확인되지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "허용된 인물만으로 차 안팎의 시선과 신부의 다정한 미소를 구현했지만, 신부 얼굴의 클로즈업 대신 넓은 어깨너머 구도가 되었다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "명시되지 않은 다수의 구조 인물과 차량을 추가한 중대 위반이 있으며, 넓어진 문·도로 구도가 신부 얼굴 클로즈업과 좁은 문틀 조건도 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "신부는 오른쪽 전경의 차 안 인물을 바라보며 웃고, 라울도 같은 실내 방향을 본다. 전경 인물은 뒤통수와 어깨를 보이며 두 사람을 향한다. 신부의 가슴 앞 손전등은 열린 출입구를 통해 카메라가 있는 실내로 빛을 보내므로 요구된 조사 방향과 맞는다.",
        "built_space": "왼쪽에 젖고 낡은 금속 문짝 하나와 경첩·잠금장치가 보이고, 그 옆으로 열린 출입구 하나가 형성된다. 신부와 라울은 문밖에, 어두운 전경 인물은 안쪽에 있어 안에서 밖을 보는 관계가 성립한다. 왼쪽 표면의 번진 빛 반사는 손전등 위치와 모순되지 않는다. 다만 문짝과 전경 어깨가 화면 대부분을 차지하고 신부는 상반신까지 보여, 좁은 문틀 가장자리를 둔 얼굴 클로즈업이 아니다.",
        "entities": "신부는 주름진 얼굴과 짧은 회흑색 머리의 한국인 노년 남성으로 보이며, 참조와 유사한 얼굴 및 검은 성직복·흰 칼라를 갖췄다. 라울은 갈색 피부와 뒤로 정돈한 머리의 어린 남자아이로 보이지만 꽁지머리 끝은 가려져 확인할 수 없다. 전경 인물의 검은 헝클어진 머리는 현우와 양립하나 얼굴·눈물·상처는 보이지 않는다. 찰리와 포옹하는 팔은 구도 밖이다. 손전등, 비, 젖은 도로가 보이고 옅은 저녁 하늘은 있으나 참조의 주황색 일몰은 약하다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "신부와 라울은 문밖에서 정상적으로 서 있는 상체 자세이며, 발은 화면 밖이라 접지점은 확인되지 않는다. 신부의 굽힌 팔이 손전등 위치로 이어지고 손 부근은 강한 눈부심에 가려져 있다. 전경 인물의 어깨와 목은 자연스럽게 연결된다. 빗물은 문 표면에 맺히고 빗줄기는 아래로 떨어지며, 근거 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "신부는 왼쪽 전경의 현우 쪽으로 고개를 기울여 미소 짓는다. 도로 중앙의 아이와 뒤쪽 사람들도 대체로 트럭 출입구를 향한다. 신부 왼쪽과 출입구 오른쪽의 강한 빛은 실내 카메라 방향으로 들어와 손전등의 조사 방향 자체는 맞는다.",
        "built_space": "왼쪽 중앙에는 창과 외판이 있는 문짝, 오른쪽에는 창과 내장 손잡이가 있는 별도 문짝이 보여 두 문짝 사이로 도로를 바라보는 구도다. 신부의 몸은 왼쪽 문짝에 상당 부분 가려지고, 아이와 여러 사람은 그 뒤 도로에 배치된다. 지정된 좁은 문틀보다 두 문짝과 도로가 훨씬 큰 비중을 차지하며, 신부 얼굴은 작다. 기존 트럭 내부의 고정 설비 연속성을 확인할 부분은 제한적이다.",
        "entities": "신부의 나이 든 한국인 남성 외모, 회흑색 머리, 검은 성직복과 흰 칼라는 참조에 대체로 부합한다. 왼쪽 현우는 검은 머리와 남색 상의를 보이지만 흐려진 옆얼굴만으로 정확한 얼굴·상처·눈물은 확인하기 어렵다. 중앙의 남색 상의 아이는 라울 역할로 읽히나 얼굴과 꽁지머리는 불분명하다. 그 외에도 여러 성인과 배경 구조 차량 최소 세 대가 추가되어 있다. 비와 젖은 노면은 구현됐지만 하늘은 주황색 일몰보다 푸른 야간에 가깝다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "허용된 신부·현우·라울 외에 여러 성인 구조 인물을 실제 신체와 얼굴로 추가해, 다른 사람을 등장시키지 말라는 조건을 위반했다.",
         "장면에 명시되지 않은 소방차·구급차 형태의 구조 차량 여러 대를 배경의 주요 사물로 추가했다."
        ],
        "physics": "도로 위 아이와 뒤쪽 인물들은 발이 노면에 닿아 있거나 다른 사람에게 가려진 정상적인 기립 자세다. 신부는 문 옆에서 상체를 기울인 자세로, 공중에 뜬 모습은 아니다. 오른쪽 손전등은 사람의 손·팔 부근에서 빛나고 왼쪽 광원은 눈부심 때문에 파지 상태가 판독되지 않는다. 젖은 노면의 광원 반사는 자연스러우며, 명백히 무지지 상태인 물체는 확인되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.667,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.417,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정되지 않은 발명된 인물들(배경의 군중) 추가",
     "[gemini-pro] 신부가 서 있는 곳과 별도로 화면 우측에 또 다른 문이 열려 있는 물리적/구조적 모순",
     "[gpt-high] 허용된 신부·현우·라울 외에 여러 성인 구조 인물을 실제 신체와 얼굴로 추가해, 다른 사람을 등장시키지 말라는 조건을 위반했다.",
     "[gpt-high] 장면에 명시되지 않은 소방차·구급차 형태의 구조 차량 여러 대를 배경의 주요 사물로 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 417,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 417,
    "verdict_ko": "배경에 프롬프트가 금지한 다수의 인물을 임의로 추가하였으며, 신부가 서 있는 창문과 별개로 열린 문이 존재하는 구조적 모순이 발생하여 실격 사유가 됩니다.  ★위반: [gemini-pro] 지정되지 않은 발명된 인물들(배경의 군중) 추가 / [gemini-pro] 신부가 서 있는 곳과 별도로 화면 우측에 또 다른 문이 열려 있는 물리적/구조적 모순 / [gpt-high] 허용된 신부·현우·라울 외에 여러 성인 구조 인물을 실제 신체와 얼굴로 추가해, 다른 사람을 등장시키지 말라는 조건을 위반했다. / [gpt-high] 장면에 명시되지 않은 소방차·구급차 형태의 구조 차량 여러 대를 배경의 주요 사물로 추가했다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 프레이밍(내부에서 바깥을 향한 대각선 근접 촬영)을 완벽히 따랐으며, 명시된 세 명의 인물만을 정확히 배치하고 다정한 신부의 미소와 플래시 불빛을 훌륭하게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43_sel.png",
    "asset_id": "f7056775-d06b-41a2-912c-04af85826fa7",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-9065-701e-b4e6-cd0943b61be0",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S72sh43"
  },
  "staged_characters_added": [
   "C01",
   "C08"
  ]
 },
 "S72sh58::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:27:56.216400+00:00",
  "fingerprint": "b26cc63f4a5a2ff9991f6a9d074712508b49882d05b9c29131635500af1de0f6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S72sh58_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S72sh58_sel.png",
  "source_sha256": "28c8a17285ac80ddc8449df401ca2cab31eeba64d9c48e2fc2bff9451637b821",
  "file": "S72sh58_cine.png",
  "staged_sha256": "1b4a270d558433589e90667dbf7c700c472087302b175d31314b44b96bb44230",
  "latency_ms": 11689
 },
 "S73sh2::signage": {
  "fp": "60df091bc68a947f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9c2ab35875802e03": {
  "subjects": [],
  "subject_text": "목포 항구도시 보건소 병실\n응급치료용 침대가 놓인 소규모 보건소 병실. 단출한 실내에 침대와 주변 통로가 마련되어 있고 출입문이 가까이 있다.",
  "identity": "canonical",
  "scope_id": "L253",
  "scope_role": "location_interior",
  "scope_sha": "bd4d2d740e8fc0c7"
 },
 "S73sh2::bgfirst_bg": {
  "input_fingerprint": "5df15547258838c8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2__bgfirst_bg.png",
  "asset_id": "a58436f5-7eb4-4977-a36e-faddd6803726",
  "input_asset_ids": [
   "480435f8-c957-407f-9883-a0ac016f25eb",
   "469b5a0c-e506-448f-a857-aa6ec1553e01"
  ]
 },
 "S73sh2": {
  "input_fingerprint": "6a346505e9fd4f9f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting is the health center at night, with a bed in use. Charlie's previously damaged bodywork and exposed chest opening remain unrepaired. 앰버: She is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment for her side injury; no specific dressing arrangement is established. 현우: He remains beside the bed, worried, with dirty skin, shabby clothing and accumulated injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting is the health center at night, with a bed in use. Charlie's previously damaged bodywork and exposed chest opening remain unrepaired. 앰버: She is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment for her side injury; no specific dressing arrangement is established. 현우: He remains beside the bed, worried, with dirty skin, shabby clothing and accumulated injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting is the health center at night, with a bed in use. Charlie's previously damaged bodywork and exposed chest opening remain unrepaired. 앰버: She is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment for her side injury; no specific dressing arrangement is established. 현우: He remains beside the bed, worried, with dirty skin, shabby clothing and accumulated injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2__bgfirst_bg.png",
     "asset_id": "a58436f5-7eb4-4977-a36e-faddd6803726",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S73sh2.png",
     "asset_id": "480435f8-c957-407f-9883-a0ac016f25eb",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L253B02.png",
     "asset_id": "469b5a0c-e506-448f-a857-aa6ec1553e01",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 침대에 누워 있는 앰버를 정확히 향해 아래로 향함.",
    "built_space": "기준 사진과 동일하게 왼쪽에 옷장과 침대, 정면에 창문이 배치된 병실 구조를 충실히 따름.",
    "entities": "현우의 얼굴과 체형은 일치하나 셔츠 안에 티셔츠가 추가됨. 앰버는 금발과 이목구비 묘사가 일치함.",
    "hard_violations": [],
    "physics": "앰버는 침대 매트리스에 체중을 싣고 안정적으로 누워 있으며, 현우는 바닥을 딛고 서 있음."
   },
   {
    "label": "B",
    "direction": "현우가 침대에 누운 앰버를 내려다보고 있음.",
    "built_space": "기준 사진과 반대로 침대 헤드와 옷장이 오른쪽에 위치하여 원본 공간의 구조를 심각하게 왜곡함.",
    "entities": "현우의 얼굴과 회색 셔츠 의상이 기준과 완벽히 일치함. 앰버 역시 일치함.",
    "hard_violations": [],
    "physics": "앰버는 침대에 정상적으로 누워 있고, 현우는 중력에 맞게 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "화면 좌측 하단을 가로지르는 침대 모서리의 사선 배치라는 프레임 지시사항과 기준 공간의 구조를 매우 정확하게 구현하여 우수합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 의상 구현은 훌륭하나, 침대를 수평으로 배치하여 프레임 지시를 어겼으며 기준 공간의 좌우 배치(옷장과 침대 위치)를 임의로 뒤바꾸어 감점 요소가 큽니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 침대에 누워 있는 앰버를 정확히 향해 아래로 향함.",
        "built_space": "기준 사진과 동일하게 왼쪽에 옷장과 침대, 정면에 창문이 배치된 병실 구조를 충실히 따름.",
        "entities": "현우의 얼굴과 체형은 일치하나 셔츠 안에 티셔츠가 추가됨. 앰버는 금발과 이목구비 묘사가 일치함.",
        "hard_violations": [],
        "physics": "앰버는 침대 매트리스에 체중을 싣고 안정적으로 누워 있으며, 현우는 바닥을 딛고 서 있음."
       },
       {
        "label": "B",
        "direction": "현우가 침대에 누운 앰버를 내려다보고 있음.",
        "built_space": "기준 사진과 반대로 침대 헤드와 옷장이 오른쪽에 위치하여 원본 공간의 구조를 심각하게 왜곡함.",
        "entities": "현우의 얼굴과 회색 셔츠 의상이 기준과 완벽히 일치함. 앰버 역시 일치함.",
        "hard_violations": [],
        "physics": "앰버는 침대에 정상적으로 누워 있고, 현우는 중력에 맞게 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "화면 좌측 하단을 가로지르는 침대 모서리의 사선 배치라는 프레임 지시사항과 기준 공간의 구조를 매우 정확하게 구현하여 우수합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 의상 구현은 훌륭하나, 침대를 수평으로 배치하여 프레임 지시를 어겼으며 기준 공간의 좌우 배치(옷장과 침대 위치)를 임의로 뒤바꾸어 감점 요소가 큽니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 침대에 누워 있는 앰버를 정확히 향해 아래로 향함.",
        "built_space": "기준 사진과 동일하게 왼쪽에 옷장과 침대, 정면에 창문이 배치된 병실 구조를 충실히 따름.",
        "entities": "현우의 얼굴과 체형은 일치하나 셔츠 안에 티셔츠가 추가됨. 앰버는 금발과 이목구비 묘사가 일치함.",
        "hard_violations": [],
        "physics": "앰버는 침대 매트리스에 체중을 싣고 안정적으로 누워 있으며, 현우는 바닥을 딛고 서 있음."
       },
       {
        "label": "B",
        "direction": "현우가 침대에 누운 앰버를 내려다보고 있음.",
        "built_space": "기준 사진과 반대로 침대 헤드와 옷장이 오른쪽에 위치하여 원본 공간의 구조를 심각하게 왜곡함.",
        "entities": "현우의 얼굴과 회색 셔츠 의상이 기준과 완벽히 일치함. 앰버 역시 일치함.",
        "hard_violations": [],
        "physics": "앰버는 침대에 정상적으로 누워 있고, 현우는 중력에 맞게 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "앰버를 걱정스럽게 내려다보는 행동은 맞지만, 침대가 하단 전체를 가로지르는 넓은 구도와 달라진 병실 설비 배치가 지정된 상체 중심 촬영 및 장소와 어긋난다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우의 상체 중심 미디엄 숏, 왼쪽 아래 대각선 침대와 앰버를 향한 시선, 참조 병실 배치를 충실히 구현했으나 겹쳐 입은 옷과 어깨끈은 인물 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 화면 오른쪽 아래의 앰버 쪽으로 내리고 있다. 앰버는 눈을 감고 누워 있으며 누구를 바라보지 않는다. 무기나 방향을 따져야 할 휴대 물체는 없다.",
        "built_space": "침대 한 개가 전경 하단 전체를 거의 가로지르고, 현우는 그 건너편에 서 있다. 오른쪽 벽에 의료용 벽 패널 한 개와 조명 한 개, 높은 수납장 한 개가 있고, 중앙 뒤에는 창 하나, 왼쪽에는 문 하나와 금속 수납장 및 카트가 보인다. 참조의 침대 머리 위 패널, 그 오른쪽 이중 창과 협탁이라는 배치와 상당히 다르다. 창에는 사람처럼 보이는 희미한 반사가 있으나, 이 화면만으로 별도 인물이나 광학적으로 불가능한 반사라고 단정하기는 어렵다.",
        "entities": "보이는 실물 인물은 현우와 앰버 두 명이다. 현우의 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 회색 단추 셔츠와 올리브색 바지는 참조와 대체로 맞으며 얼굴과 옷에 오염이 보인다. 앰버는 금발의 어린 여자아이로, 둥근 얼굴과 남색 상의가 참조에 부합한다. 잠든 상태이므로 눈 크기는 확인할 수 없다. 흰 베개와 파란 담요가 있는 의료용 침대가 보이고, 창밖은 밤이다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앰버의 머리와 목은 베개에, 누운 몸은 매트리스에 받쳐져 있다. 담요는 몸과 침대 위에 내려앉아 있고 공중에 든 신체 부위는 없다. 현우의 발은 프레임 밖이지만 몸통과 아래로 이어지는 다리는 정상적인 입식 자세로 보이며 부유 징후는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 화면 왼쪽 아래 침대에 누운 앰버를 향한다. 상체도 아이 쪽으로 조금 기울어 있어 걱정하며 내려다보는 관계가 명확하다. 앰버는 눈을 감고 잠들어 있다. 방향성 있는 휴대 물체는 없다.",
        "built_space": "침대 한 개가 왼쪽 아래 전경에서 대각선으로 뻗고, 현우는 침대 오른쪽 곁에 배치되어 상체가 크게 보인다. 왼쪽 수납장 한 개, 침대 머리 위 의료용 패널과 조명 각 한 개, 수액대와 펌프 각 한 개, 창 아래 협탁 한 개, 뒤쪽 이중 창과 커튼이 보인다. 이 설비들의 상대적 위치와 벽·바닥 재질은 장소 참조에 가깝다. 설비의 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "현우와 앰버만 보인다. 현우는 앳된 동아시아계 남성으로 검은 헝클어진 머리, 마른 체격, 얼굴의 상처와 오염이 요구에 맞는다. 다만 참조의 잠근 회색 셔츠 대신 회색 셔츠를 열어 안쪽 티셔츠를 드러내고, 후드처럼 보이는 겹옷과 어깨끈을 더했다. 앰버는 참조와 유사한 금발, 둥근 얼굴의 어린 여자아이이며 남색 상의 일부가 보인다. 감긴 눈 때문에 큰 눈의 일치 여부는 판단할 수 없다. 의료용 침대, 흰 베개, 파란 담요가 있고 창밖은 밤이며 배경 표기는 읽을 수 없다.",
        "hard_violations": [],
        "physics": "앰버의 머리는 베개에 충분히 놓여 있고 몸은 매트리스에 받쳐져 있다. 보이는 신체에 힘을 주어 들어 올린 부분이 없으며 담요도 몸의 굴곡을 따라 자연스럽게 덮여 있다. 현우는 침대 곁에서 상체를 조금 숙인 자세로 보인다. 하체와 발은 잘렸지만 공중에 떠 있거나 지지가 불가능하다고 볼 근거는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "앰버를 걱정스럽게 내려다보는 행동은 맞지만, 침대가 하단 전체를 가로지르는 넓은 구도와 달라진 병실 설비 배치가 지정된 상체 중심 촬영 및 장소와 어긋난다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우의 상체 중심 미디엄 숏, 왼쪽 아래 대각선 침대와 앰버를 향한 시선, 참조 병실 배치를 충실히 구현했으나 겹쳐 입은 옷과 어깨끈은 인물 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 화면 오른쪽 아래의 앰버 쪽으로 내리고 있다. 앰버는 눈을 감고 누워 있으며 누구를 바라보지 않는다. 무기나 방향을 따져야 할 휴대 물체는 없다.",
        "built_space": "침대 한 개가 전경 하단 전체를 거의 가로지르고, 현우는 그 건너편에 서 있다. 오른쪽 벽에 의료용 벽 패널 한 개와 조명 한 개, 높은 수납장 한 개가 있고, 중앙 뒤에는 창 하나, 왼쪽에는 문 하나와 금속 수납장 및 카트가 보인다. 참조의 침대 머리 위 패널, 그 오른쪽 이중 창과 협탁이라는 배치와 상당히 다르다. 창에는 사람처럼 보이는 희미한 반사가 있으나, 이 화면만으로 별도 인물이나 광학적으로 불가능한 반사라고 단정하기는 어렵다.",
        "entities": "보이는 실물 인물은 현우와 앰버 두 명이다. 현우의 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 회색 단추 셔츠와 올리브색 바지는 참조와 대체로 맞으며 얼굴과 옷에 오염이 보인다. 앰버는 금발의 어린 여자아이로, 둥근 얼굴과 남색 상의가 참조에 부합한다. 잠든 상태이므로 눈 크기는 확인할 수 없다. 흰 베개와 파란 담요가 있는 의료용 침대가 보이고, 창밖은 밤이다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앰버의 머리와 목은 베개에, 누운 몸은 매트리스에 받쳐져 있다. 담요는 몸과 침대 위에 내려앉아 있고 공중에 든 신체 부위는 없다. 현우의 발은 프레임 밖이지만 몸통과 아래로 이어지는 다리는 정상적인 입식 자세로 보이며 부유 징후는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 왼쪽 아래 침대에 누운 앰버를 향한다. 상체도 아이 쪽으로 조금 기울어 있어 걱정하며 내려다보는 관계가 명확하다. 앰버는 눈을 감고 잠들어 있다. 방향성 있는 휴대 물체는 없다.",
        "built_space": "침대 한 개가 왼쪽 아래 전경에서 대각선으로 뻗고, 현우는 침대 오른쪽 곁에 배치되어 상체가 크게 보인다. 왼쪽 수납장 한 개, 침대 머리 위 의료용 패널과 조명 각 한 개, 수액대와 펌프 각 한 개, 창 아래 협탁 한 개, 뒤쪽 이중 창과 커튼이 보인다. 이 설비들의 상대적 위치와 벽·바닥 재질은 장소 참조에 가깝다. 설비의 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "현우와 앰버만 보인다. 현우는 앳된 동아시아계 남성으로 검은 헝클어진 머리, 마른 체격, 얼굴의 상처와 오염이 요구에 맞는다. 다만 참조의 잠근 회색 셔츠 대신 회색 셔츠를 열어 안쪽 티셔츠를 드러내고, 후드처럼 보이는 겹옷과 어깨끈을 더했다. 앰버는 참조와 유사한 금발, 둥근 얼굴의 어린 여자아이이며 남색 상의 일부가 보인다. 감긴 눈 때문에 큰 눈의 일치 여부는 판단할 수 없다. 의료용 침대, 흰 베개, 파란 담요가 있고 창밖은 밤이며 배경 표기는 읽을 수 없다.",
        "hard_violations": [],
        "physics": "앰버의 머리는 베개에 충분히 놓여 있고 몸은 매트리스에 받쳐져 있다. 보이는 신체에 힘을 주어 들어 올린 부분이 없으며 담요도 몸의 굴곡을 따라 자연스럽게 덮여 있다. 현우는 침대 곁에서 상체를 조금 숙인 자세로 보인다. 하체와 발은 잘렸지만 공중에 떠 있거나 지지가 불가능하다고 볼 근거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.196
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.196
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1196
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "화면 좌측 하단을 가로지르는 침대 모서리의 사선 배치라는 프레임 지시사항과 기준 공간의 구조를 매우 정확하게 구현하여 우수합니다."
   },
   {
    "label": "B",
    "score": 1196,
    "verdict_ko": "현우의 의상 구현은 훌륭하나, 침대를 수평으로 배치하여 프레임 지시를 어겼으며 기준 공간의 좌우 배치(옷장과 침대 위치)를 임의로 뒤바꾸어 감점 요소가 큽니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L253B02.png",
    "asset_id": "469b5a0c-e506-448f-a857-aa6ec1553e01",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-9219-7bb2-8356-c0ab5a58726a",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2__bgfirst_bg.png",
   "bg_asset_id": "a58436f5-7eb4-4977-a36e-faddd6803726",
   "bg_record_key": "S73sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C03"
  ]
 },
 "S73sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:29:09.389429+00:00",
  "fingerprint": "9578aea3c4522e9aa7a8a19559c623b7e6fefbfe8fb28c14031572b7a8d97dc2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S73sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S73sh2_sel.png",
  "source_sha256": "9ed19fea33105c57efd913551cbfd8d48631bd69a3eaa4f5958637c24334352b",
  "file": "S73sh2_cine.png",
  "staged_sha256": "c7ed0b97e1443b98d6d8a3c6ff7a5cbae617313f24200634785045bb402ccde5",
  "latency_ms": 10440
 },
 "S73sh5::signage": {
  "fp": "ae0027474e2b8acb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S73sh5": {
  "input_fingerprint": "6f1bb9badbb70cf5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 꼬질꼬질한 현우의 어깨를 양손으로 꽉 감싸 쥔 신부의 밀착된 자세.\n\nLOCATION (lock): At the bedside inside the harbor-town clinic's patient room, under nighttime clinic lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bedside edge (Beside the two men as 앰버 rests outside the crop) — A short oblique section appears along the lower edge; used as Retains continuity with the bedside while keeping both supporting hands unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the clinic's restrained ambient illumination continuous, preserving 현우's visibly grubby appearance and the gentle modeling of the supporting hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the patient's bed, surrounding clinic surfaces, and nighttime interior lighting. Exclude the roadside vehicle, rain, and rescuers' flashlight beams from the preceding rescue.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged bodywork remains unrepaired. 현우: He stays beside the bed, visibly dirty and battered, with shabby clothing. 신부: He stands close beside the bed with his hands extended in a consoling gesture, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 꼬질꼬질한 현우의 어깨를 양손으로 꽉 감싸 쥔 신부의 밀착된 자세.\n\nLOCATION (lock): At the bedside inside the harbor-town clinic's patient room, under nighttime clinic lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bedside edge (Beside the two men as 앰버 rests outside the crop) — A short oblique section appears along the lower edge; used as Retains continuity with the bedside while keeping both supporting hands unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the clinic's restrained ambient illumination continuous, preserving 현우's visibly grubby appearance and the gentle modeling of the supporting hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the patient's bed, surrounding clinic surfaces, and nighttime interior lighting. Exclude the roadside vehicle, rain, and rescuers' flashlight beams from the preceding rescue.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged bodywork remains unrepaired. 현우: He stays beside the bed, visibly dirty and battered, with shabby clothing. 신부: He stands close beside the bed with his hands extended in a consoling gesture, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 꼬질꼬질한 현우의 어깨를 양손으로 꽉 감싸 쥔 신부의 밀착된 자세.\n\nLOCATION (lock): At the bedside inside the harbor-town clinic's patient room, under nighttime clinic lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bedside edge (Beside the two men as 앰버 rests outside the crop) — A short oblique section appears along the lower edge; used as Retains continuity with the bedside while keeping both supporting hands unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the clinic's restrained ambient illumination continuous, preserving 현우's visibly grubby appearance and the gentle modeling of the supporting hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the patient's bed, surrounding clinic surfaces, and nighttime interior lighting. Exclude the roadside vehicle, rain, and rescuers' flashlight beams from the preceding rescue.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged bodywork remains unrepaired. 현우: He stays beside the bed, visibly dirty and battered, with shabby clothing. 신부: He stands close beside the bed with his hands extended in a consoling gesture, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부는 현우 너머를 응시하고 있으며, 현우는 고개를 숙여 시선을 아래로 향하고 있음.",
    "built_space": "병실 내부. 왼쪽 뒤편에 침대 머리맡이 보이나, 지시된 하단 가장자리 배치가 아니며 베개가 비어 있음.",
    "entities": "현우와 신부의 얼굴은 레퍼런스와 일치하나 신부의 로만 칼라가 보이지 않음. 이전 샷에서 고정된 상태로 누워있던 소녀가 누락됨.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 신체 구조 및 의상 렌더링 (신부의 몸은 검은 옷을 입고 있으나 현우를 감싸는 신부의 팔 소매는 현우의 회색 겉옷과 동일하게 융합됨)",
     "[gemini-pro] 고정된 인물 누락 (프레임 내에 침대 베개가 포함되었음에도 이전 샷에서 제자리에 있어야 할 누워있는 소녀가 사라짐)"
    ],
    "physics": "신부가 현우를 안고 있는 자세이나, 감싸고 있는 팔이 물리적으로 불가능하게 현우의 옷과 융합되어 있어 구조적 모순이 발생함."
   },
   {
    "label": "B",
    "direction": "신부는 고개를 숙인 현우의 옆모습을 바라보고 있으며, 현우는 시선을 아래로 향함.",
    "built_space": "병실 내부. 하단에 침대 난간이 크게 가로질러 배치되어 있고 침대가 화면의 왼쪽 절반을 차지하여 미디엄 샷 구도 지시를 벗어남.",
    "entities": "현우와 신부는 명시된 겉모습 및 의상(지저분한 현우, 로만 칼라를 한 신부)과 일치함. 이전 샷의 소녀가 침대에 온전히 누워 있음.",
    "hard_violations": [
     "[gpt-high] 명시적으로 화면 밖에 두어야 하는 앰버의 얼굴과 누운 몸을 화면 안에 노출했다."
    ],
    "physics": "신부의 왼손이 현우의 어깨 위에 안정적으로 놓여 있으며 두 사람의 자세가 주변 사물과 자연스럽게 맞닿아 있음. 지탱되지 않고 떠 있는 물체는 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "치명적인 물리적 오류는 없으나, '양손으로 꽉 감싸 쥔'이라는 핵심 행동을 한 손으로 축소했고 소녀를 프레임 밖으로 제외하라는 구도 지시를 완전히 위반하여 아쉬운 결과물입니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "양손으로 안는 자세를 시도했으나 신부의 팔 소매가 현우의 옷과 융합되는 심각한 신체 렌더링 오류가 발생했으며, 프레임 내 침대에서 이전 샷의 소녀를 누락시켜 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 현우 너머를 응시하고 있으며, 현우는 고개를 숙여 시선을 아래로 향하고 있음.",
        "built_space": "병실 내부. 왼쪽 뒤편에 침대 머리맡이 보이나, 지시된 하단 가장자리 배치가 아니며 베개가 비어 있음.",
        "entities": "현우와 신부의 얼굴은 레퍼런스와 일치하나 신부의 로만 칼라가 보이지 않음. 이전 샷에서 고정된 상태로 누워있던 소녀가 누락됨.",
        "hard_violations": [
         "물리적으로 불가능한 신체 구조 및 의상 렌더링 (신부의 몸은 검은 옷을 입고 있으나 현우를 감싸는 신부의 팔 소매는 현우의 회색 겉옷과 동일하게 융합됨)",
         "고정된 인물 누락 (프레임 내에 침대 베개가 포함되었음에도 이전 샷에서 제자리에 있어야 할 누워있는 소녀가 사라짐)"
        ],
        "physics": "신부가 현우를 안고 있는 자세이나, 감싸고 있는 팔이 물리적으로 불가능하게 현우의 옷과 융합되어 있어 구조적 모순이 발생함."
       },
       {
        "label": "B",
        "direction": "신부는 고개를 숙인 현우의 옆모습을 바라보고 있으며, 현우는 시선을 아래로 향함.",
        "built_space": "병실 내부. 하단에 침대 난간이 크게 가로질러 배치되어 있고 침대가 화면의 왼쪽 절반을 차지하여 미디엄 샷 구도 지시를 벗어남.",
        "entities": "현우와 신부는 명시된 겉모습 및 의상(지저분한 현우, 로만 칼라를 한 신부)과 일치함. 이전 샷의 소녀가 침대에 온전히 누워 있음.",
        "hard_violations": [],
        "physics": "신부의 왼손이 현우의 어깨 위에 안정적으로 놓여 있으며 두 사람의 자세가 주변 사물과 자연스럽게 맞닿아 있음. 지탱되지 않고 떠 있는 물체는 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "치명적인 물리적 오류는 없으나, '양손으로 꽉 감싸 쥔'이라는 핵심 행동을 한 손으로 축소했고 소녀를 프레임 밖으로 제외하라는 구도 지시를 완전히 위반하여 아쉬운 결과물입니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "양손으로 안는 자세를 시도했으나 신부의 팔 소매가 현우의 옷과 융합되는 심각한 신체 렌더링 오류가 발생했으며, 프레임 내 침대에서 이전 샷의 소녀를 누락시켜 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부는 현우 너머를 응시하고 있으며, 현우는 고개를 숙여 시선을 아래로 향하고 있음.",
        "built_space": "병실 내부. 왼쪽 뒤편에 침대 머리맡이 보이나, 지시된 하단 가장자리 배치가 아니며 베개가 비어 있음.",
        "entities": "현우와 신부의 얼굴은 레퍼런스와 일치하나 신부의 로만 칼라가 보이지 않음. 이전 샷에서 고정된 상태로 누워있던 소녀가 누락됨.",
        "hard_violations": [
         "물리적으로 불가능한 신체 구조 및 의상 렌더링 (신부의 몸은 검은 옷을 입고 있으나 현우를 감싸는 신부의 팔 소매는 현우의 회색 겉옷과 동일하게 융합됨)",
         "고정된 인물 누락 (프레임 내에 침대 베개가 포함되었음에도 이전 샷에서 제자리에 있어야 할 누워있는 소녀가 사라짐)"
        ],
        "physics": "신부가 현우를 안고 있는 자세이나, 감싸고 있는 팔이 물리적으로 불가능하게 현우의 옷과 융합되어 있어 구조적 모순이 발생함."
       },
       {
        "label": "B",
        "direction": "신부는 고개를 숙인 현우의 옆모습을 바라보고 있으며, 현우는 시선을 아래로 향함.",
        "built_space": "병실 내부. 하단에 침대 난간이 크게 가로질러 배치되어 있고 침대가 화면의 왼쪽 절반을 차지하여 미디엄 샷 구도 지시를 벗어남.",
        "entities": "현우와 신부는 명시된 겉모습 및 의상(지저분한 현우, 로만 칼라를 한 신부)과 일치함. 이전 샷의 소녀가 침대에 온전히 누워 있음.",
        "hard_violations": [],
        "physics": "신부의 왼손이 현우의 어깨 위에 안정적으로 놓여 있으며 두 사람의 자세가 주변 사물과 자연스럽게 맞닿아 있음. 지탱되지 않고 떠 있는 물체는 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "화면 밖에 있어야 하는 앰버를 노출한 중대 위반이 있으며, 넓게 잡힌 침대와 거대한 전경 난간이 두 남자의 밀착된 미디엄 숏을 대신한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "신부가 현우의 양쪽 어깨를 두 손으로 감싸는 밀착된 미디엄 숏과 인물·병실 연속성을 충실히 구현하지만, 침대 노출은 요구한 짧은 하단 가장자리보다 크다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 현우의 얼굴을 바라보고, 현우는 침대 쪽 아래로 시선을 내린다. 신부의 두 손은 화면 오른쪽에 보이는 현우의 같은 어깨 앞뒤에 모여 있어, 양쪽 어깨를 각각 감싸 쥔 동작보다는 한쪽 어깨를 붙드는 모습이다. 침대의 여성은 눈을 감고 있다.",
        "built_space": "왼쪽에 붙박이장 하나, 벽등 하나와 의료용 벽면 패널 하나, 수액대 하나와 수액 주머니 하나 및 펌프 하나, 뒤쪽 창문 하나, 침상 옆 서랍장 하나가 보인다. 침대 하나가 화면 왼쪽과 하단 대부분을 차지하고, 전경 난간이 하단을 크게 가로지른다. 두 남자는 침대 오른쪽에 붙어 있다. 병실의 재질과 야간 조명은 참조와 유사하지만, 침대의 짧은 사선 가장자리만 남기라는 구도와 다르다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 더러운 회색 겹옷이 이전 숏에 부합한다. 신부는 주름과 희끗한 머리의 동아시아계 노년 남성이며, 검은 성직복과 흰 칼라, 십자가 목걸이가 보인다. 그러나 침대에 금발 여성 앰버의 얼굴과 몸까지 드러나, 이번 숏에 허용된 두 사람 외의 인물이 추가된다. 판독 가능한 글자는 확인되지 않는다.",
        "hard_violations": [
         "명시적으로 화면 밖에 두어야 하는 앰버의 얼굴과 누운 몸을 화면 안에 노출했다."
        ],
        "physics": "신부의 두 손은 현우의 어깨 옷감에 실제로 닿고, 팔은 신부의 몸으로 자연스럽게 이어진다. 두 남자의 하체 접지는 프레임에 가려져 있으나 공중에 떠 있는 징후는 없다. 여성은 베개와 매트리스에 받쳐져 있고, 수액 주머니는 수액대에 매달려 있다."
       },
       {
        "label": "B",
        "direction": "신부는 현우 바로 뒤에서 얼굴을 가까이 대고 현우의 어깨 너머 아래쪽을 바라본다. 현우의 시선은 화면 오른쪽 아래로 비껴간다. 신부의 두 손은 각각 현우의 양쪽 어깨와 위팔을 향해 뻗어 실제로 감싸 쥐고 있어 핵심 동작의 대상과 방향이 분명하다.",
        "built_space": "왼쪽 배경에 벽등 하나와 의료용 벽면 패널 하나, 수액대 하나와 수액 주머니 하나 및 펌프 하나가 있고, 뒤쪽에 창문 하나와 침상 옆 서랍장 하나가 보인다. 침대 하나의 머리판·측면 난간·베개 일부가 왼쪽 하단에 남는다. 두 남자는 침대 옆에서 상체 중심의 미디엄 숏으로 크게 잡히며 두 손 모두 가리지 않는다. 병실 설비와 야간의 따뜻한 조명은 참조와 이어지지만, 침대는 요구한 짧은 가장자리보다 넓게 노출된다.",
        "entities": "보이는 사람은 현우와 신부 두 명뿐이다. 현우의 어린 동아시아계 남성 얼굴, 헝클어진 검은 머리, 마른 체격, 얼굴의 흙과 상처, 낡고 더러운 회색 겹옷이 참조와 잘 맞는다. 신부도 참조와 유사한 얼굴 및 희끗한 머리의 동아시아계 노년 남성이고 검은 옷을 입었다. 성직자 칼라와 목걸이는 현우에게 가려져 확인할 수 없다. 앰버와 차량은 화면에 없으며, 판독 가능한 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부의 양팔은 현우의 양옆으로 돌아 들어오고, 양손의 손바닥과 손가락이 각 어깨·위팔의 옷감에 닿는다. 현우가 어깨를 움츠린 자세는 뒤에서 가까이 감싸는 동작으로 가능한 자세다. 하체는 잘렸지만 두 사람의 몸통은 아래 프레임까지 이어져 부유하는 모습이 아니다. 베개는 침대 위에 놓이고, 수액 주머니와 펌프는 수액대에 지지되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "화면 밖에 있어야 하는 앰버를 노출한 중대 위반이 있으며, 넓게 잡힌 침대와 거대한 전경 난간이 두 남자의 밀착된 미디엄 숏을 대신한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "신부가 현우의 양쪽 어깨를 두 손으로 감싸는 밀착된 미디엄 숏과 인물·병실 연속성을 충실히 구현하지만, 침대 노출은 요구한 짧은 하단 가장자리보다 크다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부는 현우의 얼굴을 바라보고, 현우는 침대 쪽 아래로 시선을 내린다. 신부의 두 손은 화면 오른쪽에 보이는 현우의 같은 어깨 앞뒤에 모여 있어, 양쪽 어깨를 각각 감싸 쥔 동작보다는 한쪽 어깨를 붙드는 모습이다. 침대의 여성은 눈을 감고 있다.",
        "built_space": "왼쪽에 붙박이장 하나, 벽등 하나와 의료용 벽면 패널 하나, 수액대 하나와 수액 주머니 하나 및 펌프 하나, 뒤쪽 창문 하나, 침상 옆 서랍장 하나가 보인다. 침대 하나가 화면 왼쪽과 하단 대부분을 차지하고, 전경 난간이 하단을 크게 가로지른다. 두 남자는 침대 오른쪽에 붙어 있다. 병실의 재질과 야간 조명은 참조와 유사하지만, 침대의 짧은 사선 가장자리만 남기라는 구도와 다르다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 더러운 회색 겹옷이 이전 숏에 부합한다. 신부는 주름과 희끗한 머리의 동아시아계 노년 남성이며, 검은 성직복과 흰 칼라, 십자가 목걸이가 보인다. 그러나 침대에 금발 여성 앰버의 얼굴과 몸까지 드러나, 이번 숏에 허용된 두 사람 외의 인물이 추가된다. 판독 가능한 글자는 확인되지 않는다.",
        "hard_violations": [
         "명시적으로 화면 밖에 두어야 하는 앰버의 얼굴과 누운 몸을 화면 안에 노출했다."
        ],
        "physics": "신부의 두 손은 현우의 어깨 옷감에 실제로 닿고, 팔은 신부의 몸으로 자연스럽게 이어진다. 두 남자의 하체 접지는 프레임에 가려져 있으나 공중에 떠 있는 징후는 없다. 여성은 베개와 매트리스에 받쳐져 있고, 수액 주머니는 수액대에 매달려 있다."
       },
       {
        "label": "A",
        "direction": "신부는 현우 바로 뒤에서 얼굴을 가까이 대고 현우의 어깨 너머 아래쪽을 바라본다. 현우의 시선은 화면 오른쪽 아래로 비껴간다. 신부의 두 손은 각각 현우의 양쪽 어깨와 위팔을 향해 뻗어 실제로 감싸 쥐고 있어 핵심 동작의 대상과 방향이 분명하다.",
        "built_space": "왼쪽 배경에 벽등 하나와 의료용 벽면 패널 하나, 수액대 하나와 수액 주머니 하나 및 펌프 하나가 있고, 뒤쪽에 창문 하나와 침상 옆 서랍장 하나가 보인다. 침대 하나의 머리판·측면 난간·베개 일부가 왼쪽 하단에 남는다. 두 남자는 침대 옆에서 상체 중심의 미디엄 숏으로 크게 잡히며 두 손 모두 가리지 않는다. 병실 설비와 야간의 따뜻한 조명은 참조와 이어지지만, 침대는 요구한 짧은 가장자리보다 넓게 노출된다.",
        "entities": "보이는 사람은 현우와 신부 두 명뿐이다. 현우의 어린 동아시아계 남성 얼굴, 헝클어진 검은 머리, 마른 체격, 얼굴의 흙과 상처, 낡고 더러운 회색 겹옷이 참조와 잘 맞는다. 신부도 참조와 유사한 얼굴 및 희끗한 머리의 동아시아계 노년 남성이고 검은 옷을 입었다. 성직자 칼라와 목걸이는 현우에게 가려져 확인할 수 없다. 앰버와 차량은 화면에 없으며, 판독 가능한 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부의 양팔은 현우의 양옆으로 돌아 들어오고, 양손의 손바닥과 손가락이 각 어깨·위팔의 옷감에 닿는다. 현우가 어깨를 움츠린 자세는 뒤에서 가까이 감싸는 동작으로 가능한 자세다. 하체는 잘렸지만 두 사람의 몸통은 아래 프레임까지 이어져 부유하는 모습이 아니다. 베개는 침대 위에 놓이고, 수액 주머니와 펌프는 수액대에 지지되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 1.25
   },
   "adjusted": {
    "A": 0.75,
    "B": 1.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 신체 구조 및 의상 렌더링 (신부의 몸은 검은 옷을 입고 있으나 현우를 감싸는 신부의 팔 소매는 현우의 회색 겉옷과 동일하게 융합됨)",
     "[gemini-pro] 고정된 인물 누락 (프레임 내에 침대 베개가 포함되었음에도 이전 샷에서 제자리에 있어야 할 누워있는 소녀가 사라짐)"
    ],
    "B": [
     "[gpt-high] 명시적으로 화면 밖에 두어야 하는 앰버의 얼굴과 누운 몸을 화면 안에 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1000,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1000,
    "verdict_ko": "치명적인 물리적 오류는 없으나, '양손으로 꽉 감싸 쥔'이라는 핵심 행동을 한 손으로 축소했고 소녀를 프레임 밖으로 제외하라는 구도 지시를 완전히 위반하여 아쉬운 결과물입니다.  ★위반: [gpt-high] 명시적으로 화면 밖에 두어야 하는 앰버의 얼굴과 누운 몸을 화면 안에 노출했다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "양손으로 안는 자세를 시도했으나 신부의 팔 소매가 현우의 옷과 융합되는 심각한 신체 렌더링 오류가 발생했으며, 프레임 내 침대에서 이전 샷의 소녀를 누락시켜 실격입니다.  ★위반: [gemini-pro] 물리적으로 불가능한 신체 구조 및 의상 렌더링 (신부의 몸은 검은 옷을 입고 있으나 현우를 감싸는 신부의 팔 소매는 현우의 회색 겉옷과 동일하게 융합됨) / [gemini-pro] 고정된 인물 누락 (프레임 내에 침대 베개가 포함되었음에도 이전 샷에서 제자리에 있어야 할 누워있는 소녀가 사라짐)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2_sel.png",
    "asset_id": "ae7edd90-60df-4f9c-a673-266c884b31f9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-9577-7ad4-a8e9-0273d5a11bb7",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S73sh2"
  }
 },
 "S73sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:30:31.003347+00:00",
  "fingerprint": "9f2067bfa63016fb6c4de98086d67875e1f9e70574170b2bad942c28b3c107ae",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S73sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S73sh5_sel.png",
  "source_sha256": "f92e2a9bb78689b380022c242d14a811ad1008f48673056f5dfee0e510da0fb6",
  "file": "S73sh5_cine.png",
  "staged_sha256": "043bde79dfa49608d1673a70333ce39dc89730b6d38a52704ce1bd3ce51afa99",
  "latency_ms": 13787
 },
 "S73sh11::signage": {
  "fp": "77462a96f9fb24b2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S73sh11": {
  "input_fingerprint": "46ab5651f79b8fe4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서로를 양팔로 빈틈없이 끌어안은 현우와 쿠마의 밀착된 전신.\n\nLOCATION (lock): In the open floor area of the clinic patient room near the child's bed, under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Clinic entrance (The entrance through which 쿠마 has just arrived) — Seen obliquely behind the embracing pair without specifying the door's position; used as Preserves the direction of his arrival in the wider composition; Clinic floor (Visible beneath both men's feet); used as Makes their distinct weight distribution and complete head-to-foot embrace readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the same restrained clinic ambience and controlled contrast, allowing the closeness of the embrace rather than a lighting shift to convey relief.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic bed, surrounding room surfaces, and nighttime lighting. Exclude the stranded truck and outdoor rescue lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied at night. Charlie retains his accumulated body damage and exposed chest opening. 현우: He stands with both arms raised in an embrace, still dirty and visibly battered. 쿠마: He has entered the health center and stands with both arms raised in an embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서로를 양팔로 빈틈없이 끌어안은 현우와 쿠마의 밀착된 전신.\n\nLOCATION (lock): In the open floor area of the clinic patient room near the child's bed, under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Clinic entrance (The entrance through which 쿠마 has just arrived) — Seen obliquely behind the embracing pair without specifying the door's position; used as Preserves the direction of his arrival in the wider composition; Clinic floor (Visible beneath both men's feet); used as Makes their distinct weight distribution and complete head-to-foot embrace readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the same restrained clinic ambience and controlled contrast, allowing the closeness of the embrace rather than a lighting shift to convey relief.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic bed, surrounding room surfaces, and nighttime lighting. Exclude the stranded truck and outdoor rescue lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied at night. Charlie retains his accumulated body damage and exposed chest opening. 현우: He stands with both arms raised in an embrace, still dirty and visibly battered. 쿠마: He has entered the health center and stands with both arms raised in an embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서로를 양팔로 빈틈없이 끌어안은 현우와 쿠마의 밀착된 전신.\n\nLOCATION (lock): In the open floor area of the clinic patient room near the child's bed, under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Clinic entrance (The entrance through which 쿠마 has just arrived) — Seen obliquely behind the embracing pair without specifying the door's position; used as Preserves the direction of his arrival in the wider composition; Clinic floor (Visible beneath both men's feet); used as Makes their distinct weight distribution and complete head-to-foot embrace readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the same restrained clinic ambience and controlled contrast, allowing the closeness of the embrace rather than a lighting shift to convey relief.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic bed, surrounding room surfaces, and nighttime lighting. Exclude the stranded truck and outdoor rescue lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied at night. Charlie retains his accumulated body damage and exposed chest opening. 현우: He stands with both arms raised in an embrace, still dirty and visibly battered. 쿠마: He has entered the health center and stands with both arms raised in an embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 쿠마가 마주보며 밀착하여 서로를 안고 있습니다.",
    "built_space": "병실 내부에 세 개의 침대가 나열되어 있으며, 고정되어야 할 아이의 침대가 창문 아래 수직 방향으로 임의로 이동되어 참조 이미지의 공간 배치와 완전히 어긋납니다.",
    "entities": "현우(회색 셔츠, 녹색 바지)와 쿠마(가죽 재킷, 얼굴 일치)의 인상착의가 참조와 대체로 일치하며, 금발의 아이가 덮개 아래에 누워 있습니다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (바닥을 보면 현우의 오른발 바로 뒤편에 누구의 것도 아닌 여분의 부츠 앞코가 하나 더 그려져 총 5개의 발이 존재함)",
     "[gemini-pro] 고정된 세팅 위반 (이동 불가능한 아이의 침대 위치와 방향이 원본과 다르게 완전히 재배치됨)",
     "[gpt-high] 숏에 허용되지 않은 침대의 아이를 화면에 노출했습니다.",
     "[gpt-high] 잠긴 장소의 침대 외에 전경 양쪽으로 침대 두 개를 추가하여 고정 설비를 중복하고 병실 구성을 변경했습니다."
    ],
    "physics": "두 사람 모두 바닥에 발을 딛고 서 있으나, 현우의 다리 쪽에 추가적인 발이 그려져 물리적 형태가 붕괴되었습니다."
   },
   {
    "label": "B",
    "direction": "현우와 쿠마가 마주보며 빈틈없이 전신으로 끌어안고 있습니다.",
    "built_space": "아이가 누워있는 침대가 전경에 창문 벽면과 평행하게 배치되어 있고, 두 사람 뒤편으로 출입구가 보여 참조 이미지의 병실 공간 구조를 잘 따르고 있습니다.",
    "entities": "현우(회색 셔츠, 녹색 바지)와 쿠마(가죽 재킷, 참조와 일치하는 얼굴)의 모습이 명확하게 구현되었으며, 아이가 침대에 누워 자고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (쿠마의 오른팔이 현우의 허리를 감싸고 있음에도 불구하고, 현우의 등 왼쪽 상단에 회색 소매를 입고 시계를 찬 정체불명의 세 번째 손이 불가능한 각도에서 뻗어 나와 있음)",
     "[gpt-high] 숏에 허용된 현우와 쿠마 외에 침대에 누운 아이의 얼굴과 몸을 노출했습니다."
    ],
    "physics": "두 인물은 바닥에 딛고 서서 포옹의 무게를 지탱하고 있으나, 현우의 등 뒤에 위치한 손은 해부학적으로 두 사람 중 누구의 팔 구조로도 설명되지 않는 잉여 신체 부위입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "공간 배치와 인물 외형은 참조 이미지와 일치하나, 현우의 왼쪽 어깨 뒤에서 뻗어 나온 불가능한 각도의 회색 소매 손이 발견되어 해부학적 오류로 탈락입니다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "바닥에 주인을 알 수 없는 다섯 번째 신발이 그려지는 신체적 오류가 있으며, 병실 내 아이의 침대 위치와 방향이 임의로 변경되어 탈락입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 쿠마가 마주보며 밀착하여 서로를 안고 있습니다.",
        "built_space": "병실 내부에 세 개의 침대가 나열되어 있으며, 고정되어야 할 아이의 침대가 창문 아래 수직 방향으로 임의로 이동되어 참조 이미지의 공간 배치와 완전히 어긋납니다.",
        "entities": "현우(회색 셔츠, 녹색 바지)와 쿠마(가죽 재킷, 얼굴 일치)의 인상착의가 참조와 대체로 일치하며, 금발의 아이가 덮개 아래에 누워 있습니다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (바닥을 보면 현우의 오른발 바로 뒤편에 누구의 것도 아닌 여분의 부츠 앞코가 하나 더 그려져 총 5개의 발이 존재함)",
         "고정된 세팅 위반 (이동 불가능한 아이의 침대 위치와 방향이 원본과 다르게 완전히 재배치됨)"
        ],
        "physics": "두 사람 모두 바닥에 발을 딛고 서 있으나, 현우의 다리 쪽에 추가적인 발이 그려져 물리적 형태가 붕괴되었습니다."
       },
       {
        "label": "B",
        "direction": "현우와 쿠마가 마주보며 빈틈없이 전신으로 끌어안고 있습니다.",
        "built_space": "아이가 누워있는 침대가 전경에 창문 벽면과 평행하게 배치되어 있고, 두 사람 뒤편으로 출입구가 보여 참조 이미지의 병실 공간 구조를 잘 따르고 있습니다.",
        "entities": "현우(회색 셔츠, 녹색 바지)와 쿠마(가죽 재킷, 참조와 일치하는 얼굴)의 모습이 명확하게 구현되었으며, 아이가 침대에 누워 자고 있습니다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (쿠마의 오른팔이 현우의 허리를 감싸고 있음에도 불구하고, 현우의 등 왼쪽 상단에 회색 소매를 입고 시계를 찬 정체불명의 세 번째 손이 불가능한 각도에서 뻗어 나와 있음)"
        ],
        "physics": "두 인물은 바닥에 딛고 서서 포옹의 무게를 지탱하고 있으나, 현우의 등 뒤에 위치한 손은 해부학적으로 두 사람 중 누구의 팔 구조로도 설명되지 않는 잉여 신체 부위입니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "공간 배치와 인물 외형은 참조 이미지와 일치하나, 현우의 왼쪽 어깨 뒤에서 뻗어 나온 불가능한 각도의 회색 소매 손이 발견되어 해부학적 오류로 탈락입니다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "바닥에 주인을 알 수 없는 다섯 번째 신발이 그려지는 신체적 오류가 있으며, 병실 내 아이의 침대 위치와 방향이 임의로 변경되어 탈락입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 쿠마가 마주보며 밀착하여 서로를 안고 있습니다.",
        "built_space": "병실 내부에 세 개의 침대가 나열되어 있으며, 고정되어야 할 아이의 침대가 창문 아래 수직 방향으로 임의로 이동되어 참조 이미지의 공간 배치와 완전히 어긋납니다.",
        "entities": "현우(회색 셔츠, 녹색 바지)와 쿠마(가죽 재킷, 얼굴 일치)의 인상착의가 참조와 대체로 일치하며, 금발의 아이가 덮개 아래에 누워 있습니다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (바닥을 보면 현우의 오른발 바로 뒤편에 누구의 것도 아닌 여분의 부츠 앞코가 하나 더 그려져 총 5개의 발이 존재함)",
         "고정된 세팅 위반 (이동 불가능한 아이의 침대 위치와 방향이 원본과 다르게 완전히 재배치됨)"
        ],
        "physics": "두 사람 모두 바닥에 발을 딛고 서 있으나, 현우의 다리 쪽에 추가적인 발이 그려져 물리적 형태가 붕괴되었습니다."
       },
       {
        "label": "B",
        "direction": "현우와 쿠마가 마주보며 빈틈없이 전신으로 끌어안고 있습니다.",
        "built_space": "아이가 누워있는 침대가 전경에 창문 벽면과 평행하게 배치되어 있고, 두 사람 뒤편으로 출입구가 보여 참조 이미지의 병실 공간 구조를 잘 따르고 있습니다.",
        "entities": "현우(회색 셔츠, 녹색 바지)와 쿠마(가죽 재킷, 참조와 일치하는 얼굴)의 모습이 명확하게 구현되었으며, 아이가 침대에 누워 자고 있습니다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (쿠마의 오른팔이 현우의 허리를 감싸고 있음에도 불구하고, 현우의 등 왼쪽 상단에 회색 소매를 입고 시계를 찬 정체불명의 세 번째 손이 불가능한 각도에서 뻗어 나와 있음)"
        ],
        "physics": "두 인물은 바닥에 딛고 서서 포옹의 무게를 지탱하고 있으나, 현우의 등 뒤에 위치한 손은 해부학적으로 두 사람 중 누구의 팔 구조로도 설명되지 않는 잉여 신체 부위입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "밀착 포옹과 야간 병실의 연속성은 더 충실하지만, 화면에 나오면 안 되는 침대의 아이까지 보여 주어 탈락 사유가 있습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "두 사람의 전신과 발의 지지는 명확하지만, 금지된 아이를 노출하고 침대를 세 개로 늘려 장소 잠금을 크게 위반합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 등을 카메라 쪽으로 두고 얼굴을 쿠마의 어깨에 묻고 있으며, 쿠마는 눈을 내리깔고 현우의 머리와 어깨 쪽을 향합니다. 쿠마의 두 팔은 현우의 등 위아래를 감싸고, 현우의 보이는 팔은 쿠마의 몸통을 감쌉니다. 서로를 향한 밀착 포옹으로 읽히며 카메라를 응시하지 않습니다.",
        "built_space": "왼쪽 전경에 침대 한 개, 그 뒤에 협탁 한 개와 수액대 한 개, 벽 조명과 의료 설비 띠 한 줄, 야간 창문 한 구획이 보입니다. 두 사람은 침대 옆 빈 바닥에 서 있고, 오른쪽 뒤에는 열린 출입구 한 개가 비스듬히 보입니다. 침대 난간과 파란 담요, 협탁 주변 물품은 이전 장면에 비교적 가깝습니다. 다만 침대의 아이까지 화면 안에 포함했습니다.",
        "entities": "현우는 헝클어진 검은 머리, 오염된 회색 셔츠, 올리브색 카고 바지와 낡은 신발을 착용해 참고 외형과 대체로 맞지만 얼굴은 가려져 동일성을 세밀하게 확인할 수 없습니다. 쿠마는 짧은 검은 머리의 젊은 동아시아계 남성으로 보이며 갈색 가죽 재킷, 짙은 상의, 카고 바지, 부츠와 손목시계가 참고와 부합합니다. 혼혈 여부는 외관만으로 판별할 수 없습니다. 현우의 옷에는 오염이 있으나 뚜렷한 부상은 확인하기 어렵습니다. 두 남성 외에 금발 아이의 얼굴과 몸이 보이며, 이는 두 사람만 보여 달라는 지시와 충돌합니다. 복도 안내판은 있으나 문구를 확실히 판독하기는 어렵습니다.",
        "hard_violations": [
         "숏에 허용된 현우와 쿠마 외에 침대에 누운 아이의 얼굴과 몸을 노출했습니다."
        ],
        "physics": "두 사람 모두 신발을 바닥에 대고 서 있으며, 벌어진 발과 서로 다른 다리 각도로 체중을 지탱합니다. 팔과 손은 상대의 어깨와 등에 접촉해 실제로 가능한 포옹을 이룹니다. 현우의 먼쪽 팔 일부는 몸 사이에 가려져 있지만 비정상적인 추가 팔다리는 보이지 않습니다. 아이는 베개와 매트리스에 지지되고 수액 용기는 수액대에 매달려 있으며, 지지 없이 떠 있는 신체나 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우는 쿠마 쪽으로 얼굴을 돌려 어깨에 밀착하고, 쿠마는 현우의 머리 옆을 향해 아래로 시선을 둡니다. 쿠마의 두 팔이 현우의 등을 감싸며 현우의 보이는 팔과 손도 쿠마의 등 쪽으로 돌아갑니다. 몸통 사이가 붙어 있어 상호 포옹이라는 행동 방향은 맞습니다.",
        "built_space": "아이의 침대 한 개가 왼쪽 뒤에 있고, 별도의 침대가 왼쪽 전경과 오른쪽 전경에 각각 한 개씩 있어 총 세 개가 보입니다. 출입구는 두 사람의 오른쪽 뒤에 한 개, 창문은 왼쪽 뒤에 한 구획, 수액대는 왼쪽에 한 개 있습니다. 두 사람은 침대들 사이 통로에 서며 머리부터 신발까지 바닥과 함께 보입니다. 그러나 침대의 목재 패널과 파란 연결부는 참고 침대와 다르고, 추가 침대 두 개로 병실의 구성이 바뀌었습니다.",
        "entities": "현우의 검은 헝클어진 머리, 회색 셔츠와 카고 바지, 쿠마의 짧은 검은 머리와 갈색 가죽 재킷 및 부츠는 참고에 대체로 맞습니다. 쿠마는 젊은 동아시아계 남성으로 보이며 현우의 얼굴은 가려져 나이와 얼굴 일치 여부를 자세히 확인할 수 없습니다. 현우의 의복 오염은 보이지만 심한 부상 표현은 뚜렷하지 않습니다. 왼쪽 뒤 침대에 금발 아이가 추가로 보입니다. 오른쪽에는 참고에서 확인되지 않은 수납장, 서류철과 소형 기기가 배치되어 있습니다.",
        "hard_violations": [
         "숏에 허용되지 않은 침대의 아이를 화면에 노출했습니다.",
         "잠긴 장소의 침대 외에 전경 양쪽으로 침대 두 개를 추가하여 고정 설비를 중복하고 병실 구성을 변경했습니다."
        ],
        "physics": "두 남성의 네 신발은 바닥에 접촉하며, 쿠마의 벌어진 발과 현우의 앞뒤로 어긋난 발이 각각 몸을 지지합니다. 서로 기댄 상체와 상대 등에 놓인 손은 서서 포옹하는 동작으로 가능합니다. 아이는 침대에 누워 지지되고, 침대들은 바닥에 닿은 바퀴와 프레임으로 받쳐집니다. 지지 없이 떠 있는 인물이나 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "밀착 포옹과 야간 병실의 연속성은 더 충실하지만, 화면에 나오면 안 되는 침대의 아이까지 보여 주어 탈락 사유가 있습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "두 사람의 전신과 발의 지지는 명확하지만, 금지된 아이를 노출하고 침대를 세 개로 늘려 장소 잠금을 크게 위반합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 등을 카메라 쪽으로 두고 얼굴을 쿠마의 어깨에 묻고 있으며, 쿠마는 눈을 내리깔고 현우의 머리와 어깨 쪽을 향합니다. 쿠마의 두 팔은 현우의 등 위아래를 감싸고, 현우의 보이는 팔은 쿠마의 몸통을 감쌉니다. 서로를 향한 밀착 포옹으로 읽히며 카메라를 응시하지 않습니다.",
        "built_space": "왼쪽 전경에 침대 한 개, 그 뒤에 협탁 한 개와 수액대 한 개, 벽 조명과 의료 설비 띠 한 줄, 야간 창문 한 구획이 보입니다. 두 사람은 침대 옆 빈 바닥에 서 있고, 오른쪽 뒤에는 열린 출입구 한 개가 비스듬히 보입니다. 침대 난간과 파란 담요, 협탁 주변 물품은 이전 장면에 비교적 가깝습니다. 다만 침대의 아이까지 화면 안에 포함했습니다.",
        "entities": "현우는 헝클어진 검은 머리, 오염된 회색 셔츠, 올리브색 카고 바지와 낡은 신발을 착용해 참고 외형과 대체로 맞지만 얼굴은 가려져 동일성을 세밀하게 확인할 수 없습니다. 쿠마는 짧은 검은 머리의 젊은 동아시아계 남성으로 보이며 갈색 가죽 재킷, 짙은 상의, 카고 바지, 부츠와 손목시계가 참고와 부합합니다. 혼혈 여부는 외관만으로 판별할 수 없습니다. 현우의 옷에는 오염이 있으나 뚜렷한 부상은 확인하기 어렵습니다. 두 남성 외에 금발 아이의 얼굴과 몸이 보이며, 이는 두 사람만 보여 달라는 지시와 충돌합니다. 복도 안내판은 있으나 문구를 확실히 판독하기는 어렵습니다.",
        "hard_violations": [
         "숏에 허용된 현우와 쿠마 외에 침대에 누운 아이의 얼굴과 몸을 노출했습니다."
        ],
        "physics": "두 사람 모두 신발을 바닥에 대고 서 있으며, 벌어진 발과 서로 다른 다리 각도로 체중을 지탱합니다. 팔과 손은 상대의 어깨와 등에 접촉해 실제로 가능한 포옹을 이룹니다. 현우의 먼쪽 팔 일부는 몸 사이에 가려져 있지만 비정상적인 추가 팔다리는 보이지 않습니다. 아이는 베개와 매트리스에 지지되고 수액 용기는 수액대에 매달려 있으며, 지지 없이 떠 있는 신체나 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "현우는 쿠마 쪽으로 얼굴을 돌려 어깨에 밀착하고, 쿠마는 현우의 머리 옆을 향해 아래로 시선을 둡니다. 쿠마의 두 팔이 현우의 등을 감싸며 현우의 보이는 팔과 손도 쿠마의 등 쪽으로 돌아갑니다. 몸통 사이가 붙어 있어 상호 포옹이라는 행동 방향은 맞습니다.",
        "built_space": "아이의 침대 한 개가 왼쪽 뒤에 있고, 별도의 침대가 왼쪽 전경과 오른쪽 전경에 각각 한 개씩 있어 총 세 개가 보입니다. 출입구는 두 사람의 오른쪽 뒤에 한 개, 창문은 왼쪽 뒤에 한 구획, 수액대는 왼쪽에 한 개 있습니다. 두 사람은 침대들 사이 통로에 서며 머리부터 신발까지 바닥과 함께 보입니다. 그러나 침대의 목재 패널과 파란 연결부는 참고 침대와 다르고, 추가 침대 두 개로 병실의 구성이 바뀌었습니다.",
        "entities": "현우의 검은 헝클어진 머리, 회색 셔츠와 카고 바지, 쿠마의 짧은 검은 머리와 갈색 가죽 재킷 및 부츠는 참고에 대체로 맞습니다. 쿠마는 젊은 동아시아계 남성으로 보이며 현우의 얼굴은 가려져 나이와 얼굴 일치 여부를 자세히 확인할 수 없습니다. 현우의 의복 오염은 보이지만 심한 부상 표현은 뚜렷하지 않습니다. 왼쪽 뒤 침대에 금발 아이가 추가로 보입니다. 오른쪽에는 참고에서 확인되지 않은 수납장, 서류철과 소형 기기가 배치되어 있습니다.",
        "hard_violations": [
         "숏에 허용되지 않은 침대의 아이를 화면에 노출했습니다.",
         "잠긴 장소의 침대 외에 전경 양쪽으로 침대 두 개를 추가하여 고정 설비를 중복하고 병실 구성을 변경했습니다."
        ],
        "physics": "두 남성의 네 신발은 바닥에 접촉하며, 쿠마의 벌어진 발과 현우의 앞뒤로 어긋난 발이 각각 몸을 지지합니다. 서로 기댄 상체와 상대 등에 놓인 손은 서서 포옹하는 동작으로 가능합니다. 아이는 침대에 누워 지지되고, 침대들은 바닥에 닿은 바퀴와 프레임으로 받쳐집니다. 지지 없이 떠 있는 인물이나 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (바닥을 보면 현우의 오른발 바로 뒤편에 누구의 것도 아닌 여분의 부츠 앞코가 하나 더 그려져 총 5개의 발이 존재함)",
     "[gemini-pro] 고정된 세팅 위반 (이동 불가능한 아이의 침대 위치와 방향이 원본과 다르게 완전히 재배치됨)",
     "[gpt-high] 숏에 허용되지 않은 침대의 아이를 화면에 노출했습니다.",
     "[gpt-high] 잠긴 장소의 침대 외에 전경 양쪽으로 침대 두 개를 추가하여 고정 설비를 중복하고 병실 구성을 변경했습니다."
    ],
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (쿠마의 오른팔이 현우의 허리를 감싸고 있음에도 불구하고, 현우의 등 왼쪽 상단에 회색 소매를 입고 시계를 찬 정체불명의 세 번째 손이 불가능한 각도에서 뻗어 나와 있음)",
     "[gpt-high] 숏에 허용된 현우와 쿠마 외에 침대에 누운 아이의 얼굴과 몸을 노출했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "공간 배치와 인물 외형은 참조 이미지와 일치하나, 현우의 왼쪽 어깨 뒤에서 뻗어 나온 불가능한 각도의 회색 소매 손이 발견되어 해부학적 오류로 탈락입니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (쿠마의 오른팔이 현우의 허리를 감싸고 있음에도 불구하고, 현우의 등 왼쪽 상단에 회색 소매를 입고 시계를 찬 정체불명의 세 번째 손이 불가능한 각도에서 뻗어 나와 있음) / [gpt-high] 숏에 허용된 현우와 쿠마 외에 침대에 누운 아이의 얼굴과 몸을 노출했습니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "바닥에 주인을 알 수 없는 다섯 번째 신발이 그려지는 신체적 오류가 있으며, 병실 내 아이의 침대 위치와 방향이 임의로 변경되어 탈락입니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (바닥을 보면 현우의 오른발 바로 뒤편에 누구의 것도 아닌 여분의 부츠 앞코가 하나 더 그려져 총 5개의 발이 존재함) / [gemini-pro] 고정된 세팅 위반 (이동 불가능한 아이의 침대 위치와 방향이 원본과 다르게 완전히 재배치됨) / [gpt-high] 숏에 허용되지 않은 침대의 아이를 화면에 노출했습니다. / [gpt-high] 잠긴 장소의 침대 외에 전경 양쪽으로 침대 두 개를 추가하여 고정 설비를 중복하고 병실 구성을 변경했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh5_sel.png",
    "asset_id": "0cacc8b0-a16c-414f-bc31-b52cbb9dfb35",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:806945>",
    "asset_id": "fe00e8c2-fc45-4464-b8b5-95ce8056561f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-9721-77d2-9f61-8e8165260160",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S73sh5"
  }
 },
 "S73sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:32:24.747464+00:00",
  "fingerprint": "859ecfc361ebc63e364361b4266f22892a3091d00942d92712ae8c67f532843c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S73sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S73sh11_sel.png",
  "source_sha256": "ed9f5dda1559bbc484cdeaa964c4c7fbca74d6ff31d6603d36b363b041121392",
  "file": "S73sh11_cine.png",
  "staged_sha256": "5f31f61ff99ba9fa863a7863d133a90f03263a9ecdf63dd6839a00056d281a33",
  "latency_ms": 11113
 },
 "S74sh8::signage": {
  "fp": "ee8fec592733164f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S74sh8": {
  "input_fingerprint": "ff6ad2b0ba579305",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 접시 위에 담긴 피부가 벗겨지고 흉측한 돌연변이 생선회 덩어리 클로즈업.\n\nLOCATION (lock): On a dining table at a street-side seafood eatery in the quiet harbor village at night. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Plate of mutant fish (Holding skinned, misshapen pieces of raw mutant fish) — The serving side is visible obliquely from above; used as Central focal detail, kept small enough to retain tabletop context; Restaurant tabletop (Supporting the served plate) — The top surface surrounds the plate; used as Scale reference and uncluttered framing margin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination and controlled contrast reveal the fish's malformed contours without theatrical color or added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A plate of sliced, radiation-mutated fish has been served at the harbor street restaurant at night.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 접시 위에 담긴 피부가 벗겨지고 흉측한 돌연변이 생선회 덩어리 클로즈업.\n\nLOCATION (lock): On a dining table at a street-side seafood eatery in the quiet harbor village at night. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Plate of mutant fish (Holding skinned, misshapen pieces of raw mutant fish) — The serving side is visible obliquely from above; used as Central focal detail, kept small enough to retain tabletop context; Restaurant tabletop (Supporting the served plate) — The top surface surrounds the plate; used as Scale reference and uncluttered framing margin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination and controlled contrast reveal the fish's malformed contours without theatrical color or added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A plate of sliced, radiation-mutated fish has been served at the harbor street restaurant at night.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 접시 위에 담긴 피부가 벗겨지고 흉측한 돌연변이 생선회 덩어리 클로즈업.\n\nLOCATION (lock): On a dining table at a street-side seafood eatery in the quiet harbor village at night. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Plate of mutant fish (Holding skinned, misshapen pieces of raw mutant fish) — The serving side is visible obliquely from above; used as Central focal detail, kept small enough to retain tabletop context; Restaurant tabletop (Supporting the served plate) — The top surface surrounds the plate; used as Scale reference and uncluttered framing margin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination and controlled contrast reveal the fish's malformed contours without theatrical color or added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A plate of sliced, radiation-mutated fish has been served at the harbor street restaurant at night.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "특정 대상을 향한 시선이나 방향성 없음.",
    "built_space": "야간 항구 식당. 금속 테이블 위에 접시가 위치하며, 배경에 포장마차와 항구의 흐릿한 야간 조명이 보임.",
    "entities": "생선회 접시(껍질과 비늘이 제거되지 않고 평범한 생선 조각처럼 보임), 젊은 남성의 손과 팔(샷 텍스트에 명시되지 않은 인물).",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
    ],
    "physics": "손과 접시 모두 테이블 표면에 물리적으로 자연스럽게 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "특정 대상을 향한 시선이나 방향성 없음.",
    "built_space": "야간 항구 식당. 나무 테이블 위에 접시와 간장 종지가 놓여 있고, 배경에 어선과 야외 식당 구조물이 배치됨.",
    "entities": "흉측한 돌연변이 생선회 덩어리(두껍고 기괴한 형태), 굵고 주름진 노인의 손(샷 텍스트에 없으며 20대 설정과 어긋남).",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
     "[gemini-pro] 화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
     "[gpt-high] 접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
    ],
    "physics": "손, 접시, 젓가락, 종지 등이 테이블 표면에 안정적으로 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 없는 인물의 손이 등장하여 지침을 위반했으며, 생선회에 비늘이 그대로 있어 '피부가 벗겨진' 돌연변이라는 묘사를 제대로 살리지 못했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "돌연변이 생선의 형태는 텍스트에 부합하나, 샷 텍스트에 없는 인물의 손이 추가되었고 20대 설정과 전혀 맞지 않는 노인의 손이 그려져 심각한 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 금속 테이블 위에 접시가 위치하며, 배경에 포장마차와 항구의 흐릿한 야간 조명이 보임.",
        "entities": "생선회 접시(껍질과 비늘이 제거되지 않고 평범한 생선 조각처럼 보임), 젊은 남성의 손과 팔(샷 텍스트에 명시되지 않은 인물).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨"
        ],
        "physics": "손과 접시 모두 테이블 표면에 물리적으로 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 나무 테이블 위에 접시와 간장 종지가 놓여 있고, 배경에 어선과 야외 식당 구조물이 배치됨.",
        "entities": "흉측한 돌연변이 생선회 덩어리(두껍고 기괴한 형태), 굵고 주름진 노인의 손(샷 텍스트에 없으며 20대 설정과 어긋남).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
         "화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임"
        ],
        "physics": "손, 접시, 젓가락, 종지 등이 테이블 표면에 안정적으로 지지되어 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 없는 인물의 손이 등장하여 지침을 위반했으며, 생선회에 비늘이 그대로 있어 '피부가 벗겨진' 돌연변이라는 묘사를 제대로 살리지 못했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "돌연변이 생선의 형태는 텍스트에 부합하나, 샷 텍스트에 없는 인물의 손이 추가되었고 20대 설정과 전혀 맞지 않는 노인의 손이 그려져 심각한 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 금속 테이블 위에 접시가 위치하며, 배경에 포장마차와 항구의 흐릿한 야간 조명이 보임.",
        "entities": "생선회 접시(껍질과 비늘이 제거되지 않고 평범한 생선 조각처럼 보임), 젊은 남성의 손과 팔(샷 텍스트에 명시되지 않은 인물).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨"
        ],
        "physics": "손과 접시 모두 테이블 표면에 물리적으로 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 나무 테이블 위에 접시와 간장 종지가 놓여 있고, 배경에 어선과 야외 식당 구조물이 배치됨.",
        "entities": "흉측한 돌연변이 생선회 덩어리(두껍고 기괴한 형태), 굵고 주름진 노인의 손(샷 텍스트에 없으며 20대 설정과 어긋남).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
         "화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임"
        ],
        "physics": "손, 접시, 젓가락, 종지 등이 테이블 표면에 안정적으로 지지되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "울퉁불퉁하고 흉측한 돌연변이 생선 덩어리는 더 충실하지만, 지시되지 않은 사람과 식기를 추가했고 껍질도 상당 부분 남아 있다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "접시와 식탁의 근접 구도는 맞지만, 불필요한 사람의 상반신이 크게 들어오며 생선도 피부를 벗긴 돌연변이보다 껍질 붙은 일반 토막에 가깝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "접시의 담는 면이 카메라 쪽으로 비스듬히 위에서 보인다. 생선 덩어리는 접시 중앙에 모여 있다. 왼쪽 사람은 얼굴이 잘려 시선 방향을 확인할 수 없으며, 손은 생선을 집지 않고 식탁에 놓여 있다.",
        "built_space": "앞쪽 나무 식탁 위에 큰 접시 하나, 뒤쪽 작은 그릇 하나, 오른쪽 가장자리에 잘린 작은 그릇 하나와 금속 식기 두 가닥이 보인다. 배경에는 별도 식탁과 차양 기둥, 조명, 판매대 및 정박한 배들이 보인다. 접시 주변 식탁 여백은 있지만, 요구된 단순한 식탁 중심 화면보다 배경과 부가물이 많다. 장소 참조 사진은 없어 특정 구조의 일치 여부는 판단할 수 없다.",
        "entities": "접시에는 젖은 생선 덩어리 여러 개가 있고, 혹처럼 부푼 형태가 돌연변이의 흉측함을 드러낸다. 다만 회갈색 외피가 넓게 남아 있어 피부가 벗겨진 생선회라는 조건에는 미달한다. 왼쪽에는 검은 소매와 손이 추가되어 있다. 얼굴이 없어 쿠마의 나이·혈통·머리카락은 확인할 수 없으며, 이 샷에는 애초에 사람이 요구되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
         "접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
        ],
        "physics": "생선 덩어리는 접시 바닥이나 다른 덩어리에 받쳐져 있고, 접시는 식탁에 놓여 있다. 손과 작은 그릇, 금속 식기에도 식탁의 지지가 보인다. 공중에 떠 있거나 지지 없이 매달린 물체는 없다."
       },
       {
        "label": "B",
        "direction": "접시의 담는 면을 비스듬히 위에서 내려다보며, 절단된 생선 살과 껍질 면이 함께 카메라에 드러난다. 뒤쪽 사람의 얼굴은 화면 밖이므로 시선은 확인할 수 없다. 손은 접시 뒤 가장자리 가까이에 놓여 있지만 생선을 집거나 가리키지 않는다.",
        "built_space": "금속 식탁 하나 위에 큰 접시 하나가 있으며 다른 전경 식기는 없다. 접시 주변으로 식탁 표면이 충분히 보인다. 그러나 왼쪽 위를 사람의 상반신과 팔이 크게 차지하고, 상단에는 항구의 포장면과 판매대, 배들이 넓게 들어와 접시만을 다루는 클로즈업의 집중을 흐린다.",
        "entities": "접시에는 두껍고 불규칙하게 썬 생선 토막들이 있다. 밝은 생살은 보이지만 여러 조각의 가장자리에 비늘 무늬가 있는 껍질이 뚜렷하게 남아 있고, 돌연변이 특유의 기형성은 약하다. 검은 겉옷을 입은 사람의 상반신과 한 손이 추가되었다. 얼굴과 머리카락이 보이지 않아 지정 인물의 신원 특성은 확인할 수 없다. 배경 표지는 흐려 읽을 수 없다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
        ],
        "physics": "생선 토막들은 접시와 서로에게 받쳐져 있고, 접시는 금속 식탁에 안정적으로 놓여 있다. 사람의 팔은 식탁에 기대어 있으며 손도 식탁 높이에 놓여 있다. 지지 없는 부유나 불가능한 동작은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "울퉁불퉁하고 흉측한 돌연변이 생선 덩어리는 더 충실하지만, 지시되지 않은 사람과 식기를 추가했고 껍질도 상당 부분 남아 있다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "접시와 식탁의 근접 구도는 맞지만, 불필요한 사람의 상반신이 크게 들어오며 생선도 피부를 벗긴 돌연변이보다 껍질 붙은 일반 토막에 가깝다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "접시의 담는 면이 카메라 쪽으로 비스듬히 위에서 보인다. 생선 덩어리는 접시 중앙에 모여 있다. 왼쪽 사람은 얼굴이 잘려 시선 방향을 확인할 수 없으며, 손은 생선을 집지 않고 식탁에 놓여 있다.",
        "built_space": "앞쪽 나무 식탁 위에 큰 접시 하나, 뒤쪽 작은 그릇 하나, 오른쪽 가장자리에 잘린 작은 그릇 하나와 금속 식기 두 가닥이 보인다. 배경에는 별도 식탁과 차양 기둥, 조명, 판매대 및 정박한 배들이 보인다. 접시 주변 식탁 여백은 있지만, 요구된 단순한 식탁 중심 화면보다 배경과 부가물이 많다. 장소 참조 사진은 없어 특정 구조의 일치 여부는 판단할 수 없다.",
        "entities": "접시에는 젖은 생선 덩어리 여러 개가 있고, 혹처럼 부푼 형태가 돌연변이의 흉측함을 드러낸다. 다만 회갈색 외피가 넓게 남아 있어 피부가 벗겨진 생선회라는 조건에는 미달한다. 왼쪽에는 검은 소매와 손이 추가되어 있다. 얼굴이 없어 쿠마의 나이·혈통·머리카락은 확인할 수 없으며, 이 샷에는 애초에 사람이 요구되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
         "접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
        ],
        "physics": "생선 덩어리는 접시 바닥이나 다른 덩어리에 받쳐져 있고, 접시는 식탁에 놓여 있다. 손과 작은 그릇, 금속 식기에도 식탁의 지지가 보인다. 공중에 떠 있거나 지지 없이 매달린 물체는 없다."
       },
       {
        "label": "A",
        "direction": "접시의 담는 면을 비스듬히 위에서 내려다보며, 절단된 생선 살과 껍질 면이 함께 카메라에 드러난다. 뒤쪽 사람의 얼굴은 화면 밖이므로 시선은 확인할 수 없다. 손은 접시 뒤 가장자리 가까이에 놓여 있지만 생선을 집거나 가리키지 않는다.",
        "built_space": "금속 식탁 하나 위에 큰 접시 하나가 있으며 다른 전경 식기는 없다. 접시 주변으로 식탁 표면이 충분히 보인다. 그러나 왼쪽 위를 사람의 상반신과 팔이 크게 차지하고, 상단에는 항구의 포장면과 판매대, 배들이 넓게 들어와 접시만을 다루는 클로즈업의 집중을 흐린다.",
        "entities": "접시에는 두껍고 불규칙하게 썬 생선 토막들이 있다. 밝은 생살은 보이지만 여러 조각의 가장자리에 비늘 무늬가 있는 껍질이 뚜렷하게 남아 있고, 돌연변이 특유의 기형성은 약하다. 검은 겉옷을 입은 사람의 상반신과 한 손이 추가되었다. 얼굴과 머리카락이 보이지 않아 지정 인물의 신원 특성은 확인할 수 없다. 배경 표지는 흐려 읽을 수 없다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
        ],
        "physics": "생선 토막들은 접시와 서로에게 받쳐져 있고, 접시는 금속 식탁에 안정적으로 놓여 있다. 사람의 팔은 식탁에 기대어 있으며 손도 식탁 높이에 놓여 있다. 지지 없는 부유나 불가능한 동작은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
    ],
    "B": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
     "[gemini-pro] 화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
     "[gpt-high] 접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1417
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "샷 텍스트에 없는 인물의 손이 등장하여 지침을 위반했으며, 생선회에 비늘이 그대로 있어 '피부가 벗겨진' 돌연변이라는 묘사를 제대로 살리지 못했습니다.  ★위반: [gemini-pro] 샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨 / [gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
   },
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "돌연변이 생선의 형태는 텍스트에 부합하나, 샷 텍스트에 없는 인물의 손이 추가되었고 20대 설정과 전혀 맞지 않는 노인의 손이 그려져 심각한 오류가 발생했습니다.  ★위반: [gemini-pro] 샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨 / [gemini-pro] 화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임 / [gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다. / [gpt-high] 접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-98ce-7219-86dd-a3fe5ac5a09b",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S74sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:36:35.786166+00:00",
  "fingerprint": "bfefad1326ed0bcdaa2aabafb75b5fc9cf5203d5936d904b9c7f08a1d0fdd3fb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S74sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S74sh8_sel.png",
  "source_sha256": "8f5948a3673b92b0b26c031177f9b8115aeb21647bdce98fbf6d5cbe853c9915",
  "file": "S74sh8_cine.png",
  "staged_sha256": "b3da1480d272cc65bc76b0bd9b419938b37c6d25f511dc5bad6c521b01ae42f3",
  "latency_ms": 9650
 },
 "S74sh9::signage": {
  "fp": "8bd30630902575c7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S74sh9": {
  "input_fingerprint": "4cbe4ce162a06f38",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 젓가락으로 생선회를 허공에 치켜든 채 태연한 표정으로 앞을 주시하는 쿠마의 상체.\n\nLOCATION (lock): At the outdoor table of the harbor village's street-side seafood eatery, near the nighttime fish trade. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Raised chopsticks and fish (A piece of raw fish is held aloft) — The chopsticks angle upward beside 쿠마's face rather than toward the lens; used as Small foreground-to-midground link between the plate and 쿠마's calm expression; Restaurant table (The served plate remains on the table) — Only the near tabletop and part of the plate are visible along the lower edge; used as Continuity anchor for the completed rise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient illumination and restrained contrast, giving equal readability to 쿠마's expression and the lifted food.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The plate of sliced, radiation-mutated fish remains served at the harbor street restaurant. 쿠마: He remains at the harbor street restaurant.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마 right now, so 쿠마's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 젓가락으로 생선회를 허공에 치켜든 채 태연한 표정으로 앞을 주시하는 쿠마의 상체.\n\nLOCATION (lock): At the outdoor table of the harbor village's street-side seafood eatery, near the nighttime fish trade. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Raised chopsticks and fish (A piece of raw fish is held aloft) — The chopsticks angle upward beside 쿠마's face rather than toward the lens; used as Small foreground-to-midground link between the plate and 쿠마's calm expression; Restaurant table (The served plate remains on the table) — Only the near tabletop and part of the plate are visible along the lower edge; used as Continuity anchor for the completed rise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient illumination and restrained contrast, giving equal readability to 쿠마's expression and the lifted food.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The plate of sliced, radiation-mutated fish remains served at the harbor street restaurant. 쿠마: He remains at the harbor street restaurant.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마 right now, so 쿠마's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 젓가락으로 생선회를 허공에 치켜든 채 태연한 표정으로 앞을 주시하는 쿠마의 상체.\n\nLOCATION (lock): At the outdoor table of the harbor village's street-side seafood eatery, near the nighttime fish trade. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Raised chopsticks and fish (A piece of raw fish is held aloft) — The chopsticks angle upward beside 쿠마's face rather than toward the lens; used as Small foreground-to-midground link between the plate and 쿠마's calm expression; Restaurant table (The served plate remains on the table) — Only the near tabletop and part of the plate are visible along the lower edge; used as Continuity anchor for the completed rise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient illumination and restrained contrast, giving equal readability to 쿠마's expression and the lifted food.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The plate of sliced, radiation-mutated fish remains served at the harbor street restaurant. 쿠마: He remains at the harbor street restaurant.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마 right now, so 쿠마's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "쿠마의 시선은 정면을 향함. 젓가락은 화면 좌측 상단을 향해 비스듬히 위로 들려 있음.",
    "built_space": "근경에 금속 테이블이 있고 원경에 야간 어시장이 보임. 프롬프트 지시대로 접시의 하단부 일부만 화면 아래쪽 테두리에 걸쳐 있음.",
    "entities": "쿠마의 얼굴과 체형은 레퍼런스와 일치하나, 갈색 가죽 재킷 대신 녹색 계열의 캔버스 재킷을 입고 있음. 테이블 위에 회 접시가 있고 젓가락과 회 한 점이 구현됨.",
    "hard_violations": [
     "[gpt-high] 오른쪽 매대 뒤에 쿠마 이외의 배경 인물을 추가했습니다."
    ],
    "physics": "손이 젓가락을 쥐고 있으며, 두 가닥으로 집는 대신 회가 위쪽 젓가락 끝에 꽂힌 형태로 지탱되고 있음."
   },
   {
    "label": "B",
    "direction": "쿠마의 시선은 정면을 향함. 젓가락은 가슴 앞에서 우측을 향해 완전히 수평으로 들려 있음.",
    "built_space": "근경에 금속 테이블이 있고 원경에 야간 어시장이 보임. 접시가 크롭되지 않고 테이블 위에 온전히 전체 형태를 드러내고 있음.",
    "entities": "쿠마의 인물 생김새 및 낡은 갈색 가죽 재킷 의상이 레퍼런스와 정확히 일치함. 회와 젓가락이 존재함.",
    "hard_violations": [
     "[gpt-high] 쿠마만 등장해야 하는 장면에 매대 직원과 통로 보행자 등 여러 배경 인물을 추가했습니다."
    ],
    "physics": "손이 젓가락을 자연스럽게 쥐고 있으며, 두 가닥으로 회를 단단히 집어 공중에 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도 지침(얼굴 옆으로 비스듬히 든 젓가락, 하단에 걸친 접시 크롭)을 가장 충실히 따랐으나, 재킷 재질이 레퍼런스와 다르고 젓가락으로 회를 집는 형태가 다소 어색합니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "의상과 물리적 묘사는 우수하나, 젓가락을 수평으로 들고 접시 전체가 화면에 들어와 프레임과 구도에 대한 최우선 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "쿠마의 시선은 정면을 향함. 젓가락은 화면 좌측 상단을 향해 비스듬히 위로 들려 있음.",
        "built_space": "근경에 금속 테이블이 있고 원경에 야간 어시장이 보임. 프롬프트 지시대로 접시의 하단부 일부만 화면 아래쪽 테두리에 걸쳐 있음.",
        "entities": "쿠마의 얼굴과 체형은 레퍼런스와 일치하나, 갈색 가죽 재킷 대신 녹색 계열의 캔버스 재킷을 입고 있음. 테이블 위에 회 접시가 있고 젓가락과 회 한 점이 구현됨.",
        "hard_violations": [],
        "physics": "손이 젓가락을 쥐고 있으며, 두 가닥으로 집는 대신 회가 위쪽 젓가락 끝에 꽂힌 형태로 지탱되고 있음."
       },
       {
        "label": "B",
        "direction": "쿠마의 시선은 정면을 향함. 젓가락은 가슴 앞에서 우측을 향해 완전히 수평으로 들려 있음.",
        "built_space": "근경에 금속 테이블이 있고 원경에 야간 어시장이 보임. 접시가 크롭되지 않고 테이블 위에 온전히 전체 형태를 드러내고 있음.",
        "entities": "쿠마의 인물 생김새 및 낡은 갈색 가죽 재킷 의상이 레퍼런스와 정확히 일치함. 회와 젓가락이 존재함.",
        "hard_violations": [],
        "physics": "손이 젓가락을 자연스럽게 쥐고 있으며, 두 가닥으로 회를 단단히 집어 공중에 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도 지침(얼굴 옆으로 비스듬히 든 젓가락, 하단에 걸친 접시 크롭)을 가장 충실히 따랐으나, 재킷 재질이 레퍼런스와 다르고 젓가락으로 회를 집는 형태가 다소 어색합니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "의상과 물리적 묘사는 우수하나, 젓가락을 수평으로 들고 접시 전체가 화면에 들어와 프레임과 구도에 대한 최우선 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "쿠마의 시선은 정면을 향함. 젓가락은 화면 좌측 상단을 향해 비스듬히 위로 들려 있음.",
        "built_space": "근경에 금속 테이블이 있고 원경에 야간 어시장이 보임. 프롬프트 지시대로 접시의 하단부 일부만 화면 아래쪽 테두리에 걸쳐 있음.",
        "entities": "쿠마의 얼굴과 체형은 레퍼런스와 일치하나, 갈색 가죽 재킷 대신 녹색 계열의 캔버스 재킷을 입고 있음. 테이블 위에 회 접시가 있고 젓가락과 회 한 점이 구현됨.",
        "hard_violations": [],
        "physics": "손이 젓가락을 쥐고 있으며, 두 가닥으로 집는 대신 회가 위쪽 젓가락 끝에 꽂힌 형태로 지탱되고 있음."
       },
       {
        "label": "B",
        "direction": "쿠마의 시선은 정면을 향함. 젓가락은 가슴 앞에서 우측을 향해 완전히 수평으로 들려 있음.",
        "built_space": "근경에 금속 테이블이 있고 원경에 야간 어시장이 보임. 접시가 크롭되지 않고 테이블 위에 온전히 전체 형태를 드러내고 있음.",
        "entities": "쿠마의 인물 생김새 및 낡은 갈색 가죽 재킷 의상이 레퍼런스와 정확히 일치함. 회와 젓가락이 존재함.",
        "hard_violations": [],
        "physics": "손이 젓가락을 자연스럽게 쥐고 있으며, 두 가닥으로 회를 단단히 집어 공중에 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "상체 중심 구도와 의상은 비교적 충실하지만, 젓가락이 얼굴 옆으로 올라가지 않고 가슴 앞에서 거의 수평이며 금지된 배경 인물이 여럿 추가되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 옆으로 비스듬히 올린 젓가락과 회, 태연한 전방 주시, 하단에 걸친 접시는 더 정확하지만 배경 인물 추가와 재킷 소재 변경이 남습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "쿠마는 회가 아니라 카메라 쪽 정면을 태연하게 바라봅니다. 오른손의 젓가락은 화면 왼쪽에서 중앙으로 거의 수평으로 뻗으며 끝이 약간 내려갑니다. 회는 얼굴 옆이 아닌 턱 아래 가슴 앞에 있어, 얼굴 옆으로 위를 향해야 하는 명시적 방향과 다릅니다.",
        "built_space": "전경에는 젖고 긁힌 금속 식탁 하나와 하단에 잘린 접시 하나가 있습니다. 쿠마는 식탁 뒤에 앉아 있습니다. 왼쪽에는 차양이 있는 매대와 크게 보이는 등 세 개, 오른쪽에는 별도 매대와 통들이 있으며 중앙으로 젖은 통로가 이어집니다. 이전 사진의 금속 식탁과 야간 수산시장 재질은 이어지지만, 항구의 배보다 새로 드러난 좌측 매대가 배경을 지배합니다. 식탁의 빛 반사는 가능한 방향입니다.",
        "entities": "주인공은 검은 짧은 머리의 젊은 동아시아계 남성으로, 참조의 얼굴과 낡은 짙은 가죽 재킷에 비교적 가깝습니다. 혼혈 여부는 외관만으로 확정할 수 없습니다. 젓가락 한 쌍, 들린 생선 살 한 조각, 두꺼운 껍질 붙은 회가 담긴 접시가 보입니다. 그러나 왼쪽 매대의 인물들과 중앙 통로의 보행자 등 쿠마 이외의 사람이 여러 명 있습니다. 명확히 판독되는 문구나 화면 오버레이는 보이지 않습니다.",
        "hard_violations": [
         "쿠마만 등장해야 하는 장면에 매대 직원과 통로 보행자 등 여러 배경 인물을 추가했습니다."
        ],
        "physics": "쿠마의 오른손이 젓가락을 잡고 있고 젓가락 끝이 회의 윗부분을 집어, 회가 아래로 늘어진 상태를 지탱합니다. 오른쪽 팔꿈치는 식탁에 닿아 팔을 받칩니다. 접시는 식탁 위에 놓여 있습니다. 하체와 좌석은 가려져 있지만 상체가 공중에 떠 있다고 볼 근거는 없습니다."
       },
       {
        "label": "B",
        "direction": "쿠마는 얼굴을 정면으로 두고 카메라 부근의 앞쪽을 침착하게 주시합니다. 오른손의 젓가락은 화면 왼쪽 아래에서 오른쪽 위로 올라가며, 끝에 집힌 회가 쿠마의 오른쪽 얼굴 옆에 위치합니다. 렌즈를 향해 내미는 방향이 아니라 얼굴 옆으로 들어 올린 방향이므로 지시와 잘 맞습니다.",
        "built_space": "전경에는 금속 식탁 하나와 아래쪽 가장자리에 일부만 보이는 접시 하나가 있습니다. 쿠마는 식탁 뒤에 앉아 있고, 왼쪽 뒤에는 붉은 의자와 쌓인 플라스틱 통들이 보입니다. 양쪽 매대 사이의 젖은 통로 너머로 배와 계류 시설이 작게 보여 이전 사진의 항구 맥락을 유지합니다. 가까운 식탁만 하단에 배치되어 요구한 구도에 가깝고, 식탁과 도로의 반사도 조명 위치에 비추어 가능합니다.",
        "entities": "검은 짧은 머리와 젊은 성인 남성의 얼굴은 쿠마 참조와 유사합니다. 다만 재킷은 참조의 갈색 낡은 가죽보다 녹색 계열의 직물 야전 재킷처럼 보이며 주머니와 여밈도 다릅니다. 젓가락 한 쌍과 길게 늘어진 껍질 붙은 회, 하단의 회 접시가 있습니다. 오른쪽 매대 뒤에는 머리와 상체가 드러난 배경 인물이 보여 쿠마 단독 등장 조건에 어긋납니다. 간판은 흐려 명확히 읽을 수 없습니다.",
        "hard_violations": [
         "오른쪽 매대 뒤에 쿠마 이외의 배경 인물을 추가했습니다."
        ],
        "physics": "오른손의 손가락이 젓가락을 쥐고, 젓가락 끝이 회의 상단을 집고 있습니다. 길쭉한 회의 아래 부분은 중력 방향으로 처져 있어 지지 관계가 자연스럽습니다. 팔꿈치는 식탁 위에 닿고 접시도 식탁에 받쳐져 있습니다. 보이는 신체나 소품에 지지 없는 부유는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "상체 중심 구도와 의상은 비교적 충실하지만, 젓가락이 얼굴 옆으로 올라가지 않고 가슴 앞에서 거의 수평이며 금지된 배경 인물이 여럿 추가되었습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "얼굴 옆으로 비스듬히 올린 젓가락과 회, 태연한 전방 주시, 하단에 걸친 접시는 더 정확하지만 배경 인물 추가와 재킷 소재 변경이 남습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "쿠마는 회가 아니라 카메라 쪽 정면을 태연하게 바라봅니다. 오른손의 젓가락은 화면 왼쪽에서 중앙으로 거의 수평으로 뻗으며 끝이 약간 내려갑니다. 회는 얼굴 옆이 아닌 턱 아래 가슴 앞에 있어, 얼굴 옆으로 위를 향해야 하는 명시적 방향과 다릅니다.",
        "built_space": "전경에는 젖고 긁힌 금속 식탁 하나와 하단에 잘린 접시 하나가 있습니다. 쿠마는 식탁 뒤에 앉아 있습니다. 왼쪽에는 차양이 있는 매대와 크게 보이는 등 세 개, 오른쪽에는 별도 매대와 통들이 있으며 중앙으로 젖은 통로가 이어집니다. 이전 사진의 금속 식탁과 야간 수산시장 재질은 이어지지만, 항구의 배보다 새로 드러난 좌측 매대가 배경을 지배합니다. 식탁의 빛 반사는 가능한 방향입니다.",
        "entities": "주인공은 검은 짧은 머리의 젊은 동아시아계 남성으로, 참조의 얼굴과 낡은 짙은 가죽 재킷에 비교적 가깝습니다. 혼혈 여부는 외관만으로 확정할 수 없습니다. 젓가락 한 쌍, 들린 생선 살 한 조각, 두꺼운 껍질 붙은 회가 담긴 접시가 보입니다. 그러나 왼쪽 매대의 인물들과 중앙 통로의 보행자 등 쿠마 이외의 사람이 여러 명 있습니다. 명확히 판독되는 문구나 화면 오버레이는 보이지 않습니다.",
        "hard_violations": [
         "쿠마만 등장해야 하는 장면에 매대 직원과 통로 보행자 등 여러 배경 인물을 추가했습니다."
        ],
        "physics": "쿠마의 오른손이 젓가락을 잡고 있고 젓가락 끝이 회의 윗부분을 집어, 회가 아래로 늘어진 상태를 지탱합니다. 오른쪽 팔꿈치는 식탁에 닿아 팔을 받칩니다. 접시는 식탁 위에 놓여 있습니다. 하체와 좌석은 가려져 있지만 상체가 공중에 떠 있다고 볼 근거는 없습니다."
       },
       {
        "label": "A",
        "direction": "쿠마는 얼굴을 정면으로 두고 카메라 부근의 앞쪽을 침착하게 주시합니다. 오른손의 젓가락은 화면 왼쪽 아래에서 오른쪽 위로 올라가며, 끝에 집힌 회가 쿠마의 오른쪽 얼굴 옆에 위치합니다. 렌즈를 향해 내미는 방향이 아니라 얼굴 옆으로 들어 올린 방향이므로 지시와 잘 맞습니다.",
        "built_space": "전경에는 금속 식탁 하나와 아래쪽 가장자리에 일부만 보이는 접시 하나가 있습니다. 쿠마는 식탁 뒤에 앉아 있고, 왼쪽 뒤에는 붉은 의자와 쌓인 플라스틱 통들이 보입니다. 양쪽 매대 사이의 젖은 통로 너머로 배와 계류 시설이 작게 보여 이전 사진의 항구 맥락을 유지합니다. 가까운 식탁만 하단에 배치되어 요구한 구도에 가깝고, 식탁과 도로의 반사도 조명 위치에 비추어 가능합니다.",
        "entities": "검은 짧은 머리와 젊은 성인 남성의 얼굴은 쿠마 참조와 유사합니다. 다만 재킷은 참조의 갈색 낡은 가죽보다 녹색 계열의 직물 야전 재킷처럼 보이며 주머니와 여밈도 다릅니다. 젓가락 한 쌍과 길게 늘어진 껍질 붙은 회, 하단의 회 접시가 있습니다. 오른쪽 매대 뒤에는 머리와 상체가 드러난 배경 인물이 보여 쿠마 단독 등장 조건에 어긋납니다. 간판은 흐려 명확히 읽을 수 없습니다.",
        "hard_violations": [
         "오른쪽 매대 뒤에 쿠마 이외의 배경 인물을 추가했습니다."
        ],
        "physics": "오른손의 손가락이 젓가락을 쥐고, 젓가락 끝이 회의 상단을 집고 있습니다. 길쭉한 회의 아래 부분은 중력 방향으로 처져 있어 지지 관계가 자연스럽습니다. 팔꿈치는 식탁 위에 닿고 접시도 식탁에 받쳐져 있습니다. 보이는 신체나 소품에 지지 없는 부유는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.214
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.964
   },
   "violations": {
    "B": [
     "[gpt-high] 쿠마만 등장해야 하는 장면에 매대 직원과 통로 보행자 등 여러 배경 인물을 추가했습니다."
    ],
    "A": [
     "[gpt-high] 오른쪽 매대 뒤에 쿠마 이외의 배경 인물을 추가했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 964
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "구도 지침(얼굴 옆으로 비스듬히 든 젓가락, 하단에 걸친 접시 크롭)을 가장 충실히 따랐으나, 재킷 재질이 레퍼런스와 다르고 젓가락으로 회를 집는 형태가 다소 어색합니다.  ★위반: [gpt-high] 오른쪽 매대 뒤에 쿠마 이외의 배경 인물을 추가했습니다."
   },
   {
    "label": "B",
    "score": 964,
    "verdict_ko": "의상과 물리적 묘사는 우수하나, 젓가락을 수평으로 들고 접시 전체가 화면에 들어와 프레임과 구도에 대한 최우선 지침을 위반했습니다.  ★위반: [gpt-high] 쿠마만 등장해야 하는 장면에 매대 직원과 통로 보행자 등 여러 배경 인물을 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S74sh8_sel.png",
    "asset_id": "4dddf26c-90fb-4c6c-810a-c95bb5bff52e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:806945>",
    "asset_id": "fe00e8c2-fc45-4464-b8b5-95ce8056561f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-9a6e-7b0b-9468-2976f005bfbb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S74sh8"
  }
 },
 "S74sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:33:36.794568+00:00",
  "fingerprint": "4655829804a7413ddb3b93a3ad74079cabc06c477ae7941f3c145d3af44a9221",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S74sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S74sh9_sel.png",
  "source_sha256": "99b1249103c8dbdab1b043170589596084d84f937d0f86a1b6a101d0fe2b797b",
  "file": "S74sh9_cine.png",
  "staged_sha256": "1913fee2cde6bdac68c3fea41cf23e304714369bda137aaa9462dd395ac3b661",
  "latency_ms": 8803
 },
 "S75sh5::signage": {
  "fp": "c7aec2cf7688f135",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::ae001ce363d68bb2": {
  "subjects": [
   {
    "subject_native": "목포공항 입구 및 간판",
    "search_terms_native": [
     "목포공항 입구",
     "목포공항 터미널",
     "폐쇄 목포공항 전경"
    ],
    "language_lock_native": "반드시 한국어로만 검색을 진행하고 다른 언어로 번역하거나 외래 검색어를 추가하지 마십시오.",
    "reason_ko": "과거 실재했던 목포공항 특유의 소형 지방공항 터미널 외형 및 진입로 간판 양식을 정확히 반영하기 위함입니다."
   }
  ],
  "subject_text": "목포공항 입구와 공터\n오래 방치된 비행장 입구와 넓은 공터. 갈라진 도로에 이끼가 번지고 공항 간판은 낡아 일부가 떨어져 있다.",
  "identity": "canonical",
  "scope_id": "L256",
  "scope_role": "location_exterior",
  "scope_sha": "90bb07f185654788"
 },
 "era_fail::4398bc8971d3c7a3": {
  "stage": "research",
  "subject": "목포공항 입구 및 간판",
  "terms": [
   "목포공항 입구",
   "목포공항 터미널",
   "폐쇄 목포공항 전경"
  ],
  "status": "no_usable",
  "queries": [
   [
    "목포공항 옛 터미널 입구 사진"
   ],
   [
    "\"목포공항\" \"입구\" \"사진\"",
    "\"목포공항\" \"폐쇄\" \"전경\""
   ]
  ],
  "candidate_urls": [
   "https://upload.wikimedia.org/wikipedia/commons/thumb/c/c0/Mokpo_airport_South_Korea_20071008.jpg/3840px-Mokpo_airport_South_Korea_20071008.jpg",
   "https://mblogthumb-phinf.pstatic.net/MjAyNTA0MDJfMTQ5/MDAxNzQzNjAzNjE5NTYz.eryxm0IqD45wHOcV0oO2zMGFdPzYcVXA-Dth-rbeImog.ncB5IC7yEqBXgQDOd-j5Mqvlsg8lJXxCH4jR_FLA7lcg.PNG/image.PNG?type=w800",
   "https://mblogthumb-phinf.pstatic.net/MjAyNTAxMTZfMTAx/MDAxNzM2OTk1ODk3MDI5.-6sUDK3RU8s2QrJOanEcDnjQmjTBtyLROJP1zfmrRK8g.bcQjHNSENj-2AFgUzyq_dNRB9GpInSbB6z1q8ZxZIKsg.JPEG/KakaoTalk_20250116_115102911_19.jpg?type=w800",
   "https://mblogthumb-phinf.pstatic.net/MjAyMjEwMzFfMjEz/MDAxNjY3MjIwMzE3MzUx.tPtudk0ULtDpAV6_lIEDmcEi_jlnfJesKHAvWY_bRocg.Xmn0fMAKrythwk-MLcWCFJbO3yq6n8AIBww52wH0nBUg.JPEG.prettyye02/SE-922abeac-b7b3-4055-8ddb-faee8c87f67a.jpg?type=w800"
  ],
  "coarse": {
   "eligible": [],
   "chosen_index": 0,
   "reason": "종류·보임·기준을 다 만족하는 후보가 없다 — 이 라운드에선 안 고른다",
   "single_judge": true,
   "rejected_judges": {}
  },
  "verdicts": [
   {
    "index": 1,
    "object_type_match": "yes",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 100
   },
   {
    "index": 2,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 30
   },
   {
    "index": 3,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 30
   },
   {
    "index": 4,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 30
   }
  ],
  "chosen_reason_ko": "",
  "attempts": 12
 },
 "S75sh5::bgfirst_bg": {
  "input_fingerprint": "5d27d9bcc99f7e15",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5__bgfirst_bg.png",
  "asset_id": "6dae19f5-dd83-4234-9035-688927798687",
  "input_asset_ids": [
   "78ed0e38-f357-47d7-96e6-bcaf44bbe506",
   "16795a55-0399-4e31-9bed-0e880d7fe7e2"
  ]
 },
 "S75sh5": {
  "input_fingerprint": "9290bf6c2b798061",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large hangar door is open, revealing numerous old, nonfunctional light aircraft. The abandoned airport has moss, broken paving and a deteriorated, partly fallen sign. 현우: He is inside the hangar, still dirty and battered from the journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large hangar door is open, revealing numerous old, nonfunctional light aircraft. The abandoned airport has moss, broken paving and a deteriorated, partly fallen sign. 현우: He is inside the hangar, still dirty and battered from the journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large hangar door is open, revealing numerous old, nonfunctional light aircraft. The abandoned airport has moss, broken paving and a deteriorated, partly fallen sign. 현우: He is inside the hangar, still dirty and battered from the journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5__bgfirst_bg.png",
     "asset_id": "6dae19f5-dd83-4234-9035-688927798687",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S75sh5.png",
     "asset_id": "78ed0e38-f357-47d7-96e6-bcaf44bbe506",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L256B02.png",
     "asset_id": "16795a55-0399-4e31-9bed-0e880d7fe7e2",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 화면 밖 위쪽을 향하고 있으며, 보이지 않는 타깃(경비행기 무더기)을 바라보는 듯함.",
    "built_space": "격납고 내부로, 배경의 경비행기들과 건물 구조가 레퍼런스 이미지의 구도 및 초점과 완전히 동일하게 배치되어 있음.",
    "entities": "현우의 얼굴과 헝클어진 머리 모양은 일치하나, 지정된 버튼다운 셔츠가 아닌 낡은 회색 티셔츠를 입고 있음. 배경의 경비행기 형태는 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "자연스럽게 서 있는 자세로 보이며, 특별히 물리적으로 어색한 부분은 없음."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 화면 밖 위쪽 공간을 향해 고정되어 있음.",
    "built_space": "격납고 내부로, 양옆에 경비행기가 배치되어 있으나 우측의 프로펠러가 매우 크게 자리 잡고 있으며, 뒷벽에 지시되지 않은 따뜻한 색감의 조명이 켜져 있음.",
    "entities": "현우의 얼굴, 머리 모양, 그리고 회색 셔츠와 카고 바지(무릎 부분) 등 의상이 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "무릎을 세우고 앉거나 웅크린 자세로 보이며, 바닥의 지지를 받고 있어 물리적 오류는 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "캐릭터의 외모와 의상(버튼다운 셔츠)을 정확히 재현했으나, 배경에 지시되지 않은 조명이 추가되고 우측 프로펠러가 과도하게 크게 부각되어 아쉬움이 남습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "장소 레퍼런스의 카메라 구도를 그대로 복사하지 말라는 지시와 배경을 아웃포커싱하라는 지시를 어겼으며, 캐릭터의 의상(티셔츠)도 잘못 묘사되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 밖 위쪽을 향하고 있으며, 보이지 않는 타깃(경비행기 무더기)을 바라보는 듯함.",
        "built_space": "격납고 내부로, 배경의 경비행기들과 건물 구조가 레퍼런스 이미지의 구도 및 초점과 완전히 동일하게 배치되어 있음.",
        "entities": "현우의 얼굴과 헝클어진 머리 모양은 일치하나, 지정된 버튼다운 셔츠가 아닌 낡은 회색 티셔츠를 입고 있음. 배경의 경비행기 형태는 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 자세로 보이며, 특별히 물리적으로 어색한 부분은 없음."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 밖 위쪽 공간을 향해 고정되어 있음.",
        "built_space": "격납고 내부로, 양옆에 경비행기가 배치되어 있으나 우측의 프로펠러가 매우 크게 자리 잡고 있으며, 뒷벽에 지시되지 않은 따뜻한 색감의 조명이 켜져 있음.",
        "entities": "현우의 얼굴, 머리 모양, 그리고 회색 셔츠와 카고 바지(무릎 부분) 등 의상이 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "무릎을 세우고 앉거나 웅크린 자세로 보이며, 바닥의 지지를 받고 있어 물리적 오류는 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "캐릭터의 외모와 의상(버튼다운 셔츠)을 정확히 재현했으나, 배경에 지시되지 않은 조명이 추가되고 우측 프로펠러가 과도하게 크게 부각되어 아쉬움이 남습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "장소 레퍼런스의 카메라 구도를 그대로 복사하지 말라는 지시와 배경을 아웃포커싱하라는 지시를 어겼으며, 캐릭터의 의상(티셔츠)도 잘못 묘사되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 밖 위쪽을 향하고 있으며, 보이지 않는 타깃(경비행기 무더기)을 바라보는 듯함.",
        "built_space": "격납고 내부로, 배경의 경비행기들과 건물 구조가 레퍼런스 이미지의 구도 및 초점과 완전히 동일하게 배치되어 있음.",
        "entities": "현우의 얼굴과 헝클어진 머리 모양은 일치하나, 지정된 버튼다운 셔츠가 아닌 낡은 회색 티셔츠를 입고 있음. 배경의 경비행기 형태는 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 자세로 보이며, 특별히 물리적으로 어색한 부분은 없음."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 밖 위쪽 공간을 향해 고정되어 있음.",
        "built_space": "격납고 내부로, 양옆에 경비행기가 배치되어 있으나 우측의 프로펠러가 매우 크게 자리 잡고 있으며, 뒷벽에 지시되지 않은 따뜻한 색감의 조명이 켜져 있음.",
        "entities": "현우의 얼굴, 머리 모양, 그리고 회색 셔츠와 카고 바지(무릎 부분) 등 의상이 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "무릎을 세우고 앉거나 웅크린 자세로 보이며, 바닥의 지지를 받고 있어 물리적 오류는 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "올려다보는 멍한 표정과 회색 셔츠는 충실하지만, 몸통과 무릎까지 들어오는 넓은 구도와 양옆의 거대한 기수가 얼굴 클로즈업·작은 배경 파편 지시에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "더 큰 얼굴 중심 구도와 화면 밖 위쪽을 향한 시선으로 우세하지만, 지나치게 선명하고 큰 항공기 배경과 참고의 셔츠를 바꾼 티셔츠는 불일치합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 턱을 들고 화면 위쪽 바깥을 바라봅니다. 시선은 양옆에 보이는 기수가 아니라 프레임 밖 높은 곳으로 향하므로, 바라보는 항공기를 화면 밖에 두라는 지시와 양립합니다. 벌어진 입과 느슨한 얼굴 근육이 멍한 반응을 표현합니다.",
        "built_space": "중앙 인물 양옆으로 항공기 기수 두 개가 크게 보이고, 뒤에는 겹친 날개 윤곽이 있습니다. 철제 벽체, 상부 창, 지붕 골조와 후방의 밝은 등 하나가 보입니다. 인물은 항공기 사이에 있으나, 양옆 기체가 작은 배경 조각이 아니라 화면을 지배하는 가까운 물체로 제시됩니다. 배경에는 흐림이 적용되어 있습니다. 열린 대형 출입문은 이 구도에서 확인할 수 없습니다.",
        "entities": "보이는 사람은 현우에 해당하는 앳된 동아시아계 남성 한 명뿐이며, 헝클어진 검은 머리와 얼굴 형태, 낡은 회색 단추 셔츠가 참고와 대체로 맞습니다. 얼굴과 목, 옷에 때와 먼지가 보입니다. 하단에는 올리브색 바지를 입은 무릎으로 보이는 부분이 있습니다. 오래된 프로펠러 경비행기의 먼지와 손상도 표현됩니다. 눈은 정상적인 사람의 눈이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨, 몸통은 자연스럽게 연결되어 있습니다. 하단의 굽힌 무릎은 낮게 앉거나 웅크린 자세와 양립하지만, 엉덩이와 발의 지지점은 화면 밖이라 확인할 수 없습니다. 공중에 떠 있다고 볼 근거는 없습니다. 항공기에는 하부 지주가 일부 보이고, 프로펠러와 날개는 기체에 연결되어 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 얼굴을 화면 왼쪽으로 돌리고 눈을 왼쪽 위로 향합니다. 보이는 왼쪽 항공기의 기수보다 높은 화면 밖 지점을 바라보므로, 응시 대상 항공기를 프레임 밖에 둔 설정에 맞습니다. 입을 다물고 얼굴의 긴장을 낮춘 정적인 반응입니다.",
        "built_space": "현우의 얼굴과 어깨가 화면 오른쪽을 크게 차지합니다. 좌우 가까운 기체 두 대와 그 뒤의 기체 두 대, 중앙 통로, 철제 트러스, 양측 높은 창열, 후면 벽과 작은 문 하나가 보여 참고 격납고의 구조와 잘 대응합니다. 다만 항공기와 건축 구조가 넓고 선명하게 드러나며, 특히 왼쪽 기수가 커서 흐릿한 부분 윤곽만 남기라는 지시에는 맞지 않습니다. 열린 대형 출입문은 화면에 없습니다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리, 마른 체격이 현우의 설정에 부합합니다. 얼굴과 목에는 때와 긁힌 흔적이 있습니다. 그러나 옷은 참고의 칼라 달린 단추 셔츠가 아니라 둥근 목의 회색 티셔츠입니다. 항공기들은 먼지 쌓인 낡은 프로펠러 경비행기로 식별됩니다. 정상적인 눈과 실제 피부 질감이 보이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "목이 어깨 위에서 머리를 자연스럽게 받치고 있으며, 고개를 들어 옆 위쪽을 보는 자세에 무리가 없습니다. 하체와 발은 클로즈업 밖이므로 지면 접촉을 확인할 수 없지만 부유 징후는 없습니다. 항공기들은 보이는 착륙장치와 바퀴로 바닥에 지지되며, 날개와 프로펠러도 기체에 연결되어 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "올려다보는 멍한 표정과 회색 셔츠는 충실하지만, 몸통과 무릎까지 들어오는 넓은 구도와 양옆의 거대한 기수가 얼굴 클로즈업·작은 배경 파편 지시에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "더 큰 얼굴 중심 구도와 화면 밖 위쪽을 향한 시선으로 우세하지만, 지나치게 선명하고 큰 항공기 배경과 참고의 셔츠를 바꾼 티셔츠는 불일치합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 턱을 들고 화면 위쪽 바깥을 바라봅니다. 시선은 양옆에 보이는 기수가 아니라 프레임 밖 높은 곳으로 향하므로, 바라보는 항공기를 화면 밖에 두라는 지시와 양립합니다. 벌어진 입과 느슨한 얼굴 근육이 멍한 반응을 표현합니다.",
        "built_space": "중앙 인물 양옆으로 항공기 기수 두 개가 크게 보이고, 뒤에는 겹친 날개 윤곽이 있습니다. 철제 벽체, 상부 창, 지붕 골조와 후방의 밝은 등 하나가 보입니다. 인물은 항공기 사이에 있으나, 양옆 기체가 작은 배경 조각이 아니라 화면을 지배하는 가까운 물체로 제시됩니다. 배경에는 흐림이 적용되어 있습니다. 열린 대형 출입문은 이 구도에서 확인할 수 없습니다.",
        "entities": "보이는 사람은 현우에 해당하는 앳된 동아시아계 남성 한 명뿐이며, 헝클어진 검은 머리와 얼굴 형태, 낡은 회색 단추 셔츠가 참고와 대체로 맞습니다. 얼굴과 목, 옷에 때와 먼지가 보입니다. 하단에는 올리브색 바지를 입은 무릎으로 보이는 부분이 있습니다. 오래된 프로펠러 경비행기의 먼지와 손상도 표현됩니다. 눈은 정상적인 사람의 눈이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨, 몸통은 자연스럽게 연결되어 있습니다. 하단의 굽힌 무릎은 낮게 앉거나 웅크린 자세와 양립하지만, 엉덩이와 발의 지지점은 화면 밖이라 확인할 수 없습니다. 공중에 떠 있다고 볼 근거는 없습니다. 항공기에는 하부 지주가 일부 보이고, 프로펠러와 날개는 기체에 연결되어 있습니다."
       },
       {
        "label": "A",
        "direction": "현우는 얼굴을 화면 왼쪽으로 돌리고 눈을 왼쪽 위로 향합니다. 보이는 왼쪽 항공기의 기수보다 높은 화면 밖 지점을 바라보므로, 응시 대상 항공기를 프레임 밖에 둔 설정에 맞습니다. 입을 다물고 얼굴의 긴장을 낮춘 정적인 반응입니다.",
        "built_space": "현우의 얼굴과 어깨가 화면 오른쪽을 크게 차지합니다. 좌우 가까운 기체 두 대와 그 뒤의 기체 두 대, 중앙 통로, 철제 트러스, 양측 높은 창열, 후면 벽과 작은 문 하나가 보여 참고 격납고의 구조와 잘 대응합니다. 다만 항공기와 건축 구조가 넓고 선명하게 드러나며, 특히 왼쪽 기수가 커서 흐릿한 부분 윤곽만 남기라는 지시에는 맞지 않습니다. 열린 대형 출입문은 화면에 없습니다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리, 마른 체격이 현우의 설정에 부합합니다. 얼굴과 목에는 때와 긁힌 흔적이 있습니다. 그러나 옷은 참고의 칼라 달린 단추 셔츠가 아니라 둥근 목의 회색 티셔츠입니다. 항공기들은 먼지 쌓인 낡은 프로펠러 경비행기로 식별됩니다. 정상적인 눈과 실제 피부 질감이 보이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "목이 어깨 위에서 머리를 자연스럽게 받치고 있으며, 고개를 들어 옆 위쪽을 보는 자세에 무리가 없습니다. 하체와 발은 클로즈업 밖이므로 지면 접촉을 확인할 수 없지만 부유 징후는 없습니다. 항공기들은 보이는 착륙장치와 바퀴로 바닥에 지지되며, 날개와 프로펠러도 기체에 연결되어 있습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.833
   },
   "adjusted": {
    "A": 1.667,
    "B": 1.833
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1833,
   "A": 1667
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1833,
    "verdict_ko": "캐릭터의 외모와 의상(버튼다운 셔츠)을 정확히 재현했으나, 배경에 지시되지 않은 조명이 추가되고 우측 프로펠러가 과도하게 크게 부각되어 아쉬움이 남습니다."
   },
   {
    "label": "A",
    "score": 1667,
    "verdict_ko": "장소 레퍼런스의 카메라 구도를 그대로 복사하지 말라는 지시와 배경을 아웃포커싱하라는 지시를 어겼으며, 캐릭터의 의상(티셔츠)도 잘못 묘사되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L256B02.png",
    "asset_id": "16795a55-0399-4e31-9bed-0e880d7fe7e2",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40e-9c14-7834-9af2-66448f98772e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5__bgfirst_bg.png",
   "bg_asset_id": "6dae19f5-dd83-4234-9035-688927798687",
   "bg_record_key": "S75sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S75sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:35:22.462171+00:00",
  "fingerprint": "d819496ee08dd0e600ec74644d93b3a9faa7966b87abff2be779ed531241c600",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S75sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S75sh5_sel.png",
  "source_sha256": "842a0f5460d6bb7a6ffefbd026453cadd8e20d4a658210b83197b43a049cadfa",
  "file": "S75sh5_cine.png",
  "staged_sha256": "f63dd0bacfd4edbfc71c7933ce990a576f1811d0363df11aaa11fd076572183e",
  "latency_ms": 10604
 },
 "S75sh10::signage": {
  "fp": "90b29363b8ba304d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S75sh10": {
  "input_fingerprint": "1c353cec7b94b8a5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우와 쿠마를 향해 뛰어오는 mid-stride 자세로, 한 발이 지면에서 떨어져 있고 겉옷 자락이 뒤로 휘날리는 신부의 전신.\n\nLOCATION (lock): At the open entrance of the abandoned airfield's aircraft-storage hangar at night, where the arriving priest reaches the others. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open warehouse entrance behind the approaching priest in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Warehouse entrance (Open after 현우 and 쿠마 entered) — Seen diagonally from inside, with the exterior approach visible through the opening; used as Shared spatial anchor connecting the foreground men and the approaching 신부; Warehouse floor (Visible beneath the running figure); used as Clear ground reference beneath the lifted foot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime ambient illumination with enough tonal separation to read the running figure against the entrance, without importing the outdoor firelight into the warehouse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The hangar remains open with the old light aircraft still nonfunctional, and the airport campfire remains lit. 현우: He remains at the hangar, dirty and visibly battered. 쿠마: He remains at the hangar. 신부: He has arrived at the airport in alarm, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우와 쿠마를 향해 뛰어오는 mid-stride 자세로, 한 발이 지면에서 떨어져 있고 겉옷 자락이 뒤로 휘날리는 신부의 전신.\n\nLOCATION (lock): At the open entrance of the abandoned airfield's aircraft-storage hangar at night, where the arriving priest reaches the others. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open warehouse entrance behind the approaching priest in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Warehouse entrance (Open after 현우 and 쿠마 entered) — Seen diagonally from inside, with the exterior approach visible through the opening; used as Shared spatial anchor connecting the foreground men and the approaching 신부; Warehouse floor (Visible beneath the running figure); used as Clear ground reference beneath the lifted foot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime ambient illumination with enough tonal separation to read the running figure against the entrance, without importing the outdoor firelight into the warehouse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The hangar remains open with the old light aircraft still nonfunctional, and the airport campfire remains lit. 현우: He remains at the hangar, dirty and visibly battered. 쿠마: He remains at the hangar. 신부: He has arrived at the airport in alarm, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우와 쿠마를 향해 뛰어오는 mid-stride 자세로, 한 발이 지면에서 떨어져 있고 겉옷 자락이 뒤로 휘날리는 신부의 전신.\n\nLOCATION (lock): At the open entrance of the abandoned airfield's aircraft-storage hangar at night, where the arriving priest reaches the others. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open warehouse entrance behind the approaching priest in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Warehouse entrance (Open after 현우 and 쿠마 entered) — Seen diagonally from inside, with the exterior approach visible through the opening; used as Shared spatial anchor connecting the foreground men and the approaching 신부; Warehouse floor (Visible beneath the running figure); used as Clear ground reference beneath the lifted foot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime ambient illumination with enough tonal separation to read the running figure against the entrance, without importing the outdoor firelight into the warehouse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The hangar remains open with the old light aircraft still nonfunctional, and the airport campfire remains lit. 현우: He remains at the hangar, dirty and visibly battered. 쿠마: He remains at the hangar. 신부: He has arrived at the airport in alarm, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부의 시선과 달리는 방향이 전경에 등을 보이고 서 있는 현우와 쿠마를 정확히 향하고 있습니다.",
    "built_space": "격납고 내부에서 대각선으로 밖을 내다보는 구도이며, 좌측에는 경비행기, 우측 중단에는 열린 출입구가 있어 지시된 공간 배치를 정확히 따릅니다.",
    "entities": "신부의 얼굴, 복장, 클러지 칼라와 휘날리는 겉옷이 지시와 일치합니다. 뒷모습으로 등장한 현우(더러운 셔츠, 헝클어진 머리)와 쿠마(파란색 계열 셔츠)의 인상착의도 레퍼런스에 부합합니다.",
    "hard_violations": [],
    "physics": "신부의 왼쪽 발이 바닥을 단단히 딛고 오른쪽 발은 공중에 떠 있어 달리는 포즈의 체중 이동이 자연스럽고, 옷자락 역시 이동 방향의 반대로 정상적으로 휘날립니다."
   },
   {
    "label": "B",
    "direction": "신부는 앞을 향해 달리고 있으나, 전경의 현우와 쿠마는 신부를 등진 채 렌즈(카메라) 쪽을 응시하고 있어 동선과 시선이 엇갈립니다.",
    "built_space": "격납고 중앙에서 정면을 바라보는 대칭 구도로 출입구가 중앙에 위치해 '대각선 뷰(diagonally)' 지시를 위반했으며, 화면 하단 전체를 출처 불명의 거대한 구조물이 가리고 있습니다.",
    "entities": "세 인물의 얼굴과 기본 복장이 레퍼런스와 잘 일치하며 신부 역시 겉옷을 입고 있습니다.",
    "hard_violations": [
     "[gemini-pro] invented objects (화면 하단을 크게 가로막고 있는 정체불명의 거대한 전경 구조물)"
    ],
    "physics": "신부의 왼쪽 발이 지면을 지탱하고 오른쪽 발이 들려 있어 달리는 동작 자체는 물리적으로 성립하며, 전경 인물들도 바닥에 안정적으로 위치해 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "대각선 앵글과 우측 출입구 배치, 현우와 쿠마를 향해 달려오는 신부의 동선과 겉옷의 휘날림까지 프롬프트의 세부 연출 지시를 매우 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물들이 카메라를 응시하여 '향해 뛰어오는' 동선의 개연성을 깨뜨렸고, 대각선 구도 지시를 무시한 채 정면 앵글과 불필요한 전경 물체를 배치해 지시에서 크게 벗어났습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선과 달리는 방향이 전경에 등을 보이고 서 있는 현우와 쿠마를 정확히 향하고 있습니다.",
        "built_space": "격납고 내부에서 대각선으로 밖을 내다보는 구도이며, 좌측에는 경비행기, 우측 중단에는 열린 출입구가 있어 지시된 공간 배치를 정확히 따릅니다.",
        "entities": "신부의 얼굴, 복장, 클러지 칼라와 휘날리는 겉옷이 지시와 일치합니다. 뒷모습으로 등장한 현우(더러운 셔츠, 헝클어진 머리)와 쿠마(파란색 계열 셔츠)의 인상착의도 레퍼런스에 부합합니다.",
        "hard_violations": [],
        "physics": "신부의 왼쪽 발이 바닥을 단단히 딛고 오른쪽 발은 공중에 떠 있어 달리는 포즈의 체중 이동이 자연스럽고, 옷자락 역시 이동 방향의 반대로 정상적으로 휘날립니다."
       },
       {
        "label": "B",
        "direction": "신부는 앞을 향해 달리고 있으나, 전경의 현우와 쿠마는 신부를 등진 채 렌즈(카메라) 쪽을 응시하고 있어 동선과 시선이 엇갈립니다.",
        "built_space": "격납고 중앙에서 정면을 바라보는 대칭 구도로 출입구가 중앙에 위치해 '대각선 뷰(diagonally)' 지시를 위반했으며, 화면 하단 전체를 출처 불명의 거대한 구조물이 가리고 있습니다.",
        "entities": "세 인물의 얼굴과 기본 복장이 레퍼런스와 잘 일치하며 신부 역시 겉옷을 입고 있습니다.",
        "hard_violations": [
         "invented objects (화면 하단을 크게 가로막고 있는 정체불명의 거대한 전경 구조물)"
        ],
        "physics": "신부의 왼쪽 발이 지면을 지탱하고 오른쪽 발이 들려 있어 달리는 동작 자체는 물리적으로 성립하며, 전경 인물들도 바닥에 안정적으로 위치해 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "대각선 앵글과 우측 출입구 배치, 현우와 쿠마를 향해 달려오는 신부의 동선과 겉옷의 휘날림까지 프롬프트의 세부 연출 지시를 매우 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물들이 카메라를 응시하여 '향해 뛰어오는' 동선의 개연성을 깨뜨렸고, 대각선 구도 지시를 무시한 채 정면 앵글과 불필요한 전경 물체를 배치해 지시에서 크게 벗어났습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선과 달리는 방향이 전경에 등을 보이고 서 있는 현우와 쿠마를 정확히 향하고 있습니다.",
        "built_space": "격납고 내부에서 대각선으로 밖을 내다보는 구도이며, 좌측에는 경비행기, 우측 중단에는 열린 출입구가 있어 지시된 공간 배치를 정확히 따릅니다.",
        "entities": "신부의 얼굴, 복장, 클러지 칼라와 휘날리는 겉옷이 지시와 일치합니다. 뒷모습으로 등장한 현우(더러운 셔츠, 헝클어진 머리)와 쿠마(파란색 계열 셔츠)의 인상착의도 레퍼런스에 부합합니다.",
        "hard_violations": [],
        "physics": "신부의 왼쪽 발이 바닥을 단단히 딛고 오른쪽 발은 공중에 떠 있어 달리는 포즈의 체중 이동이 자연스럽고, 옷자락 역시 이동 방향의 반대로 정상적으로 휘날립니다."
       },
       {
        "label": "B",
        "direction": "신부는 앞을 향해 달리고 있으나, 전경의 현우와 쿠마는 신부를 등진 채 렌즈(카메라) 쪽을 응시하고 있어 동선과 시선이 엇갈립니다.",
        "built_space": "격납고 중앙에서 정면을 바라보는 대칭 구도로 출입구가 중앙에 위치해 '대각선 뷰(diagonally)' 지시를 위반했으며, 화면 하단 전체를 출처 불명의 거대한 구조물이 가리고 있습니다.",
        "entities": "세 인물의 얼굴과 기본 복장이 레퍼런스와 잘 일치하며 신부 역시 겉옷을 입고 있습니다.",
        "hard_violations": [
         "invented objects (화면 하단을 크게 가로막고 있는 정체불명의 거대한 전경 구조물)"
        ],
        "physics": "신부의 왼쪽 발이 지면을 지탱하고 오른쪽 발이 들려 있어 달리는 동작 자체는 물리적으로 성립하며, 전경 인물들도 바닥에 안정적으로 위치해 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "신부의 전신과 입구 배치는 맞지만, 세 사람의 시선이 주로 카메라 쪽을 향해 현우·쿠마에게 달려오는 관계가 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "전경의 현우·쿠마를 향하는 신부의 시선과 달리기, 중간 오른쪽 전신, 뒤로 날리는 옷자락이 지시된 순간을 더 충실하게 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 입구에서 화면 앞쪽으로 달려오며 얼굴과 시선도 거의 카메라를 향한다. 왼쪽 전경의 현우와 쿠마를 향한 뚜렷한 대각선 접근은 약하다. 두 청년도 신부를 돌아보기보다는 카메라 쪽을 보고 있어 세 사람의 시선이 연결되지 않는다. 옷자락은 신부의 화면 오른쪽 뒤로 날린다.",
        "built_space": "금속 벽과 철골, 왼쪽 상단 창, 중앙 오른쪽의 열린 대형 출입구 한 곳과 양옆 문짝이 보인다. 내부 양쪽에 낡은 경비행기 동체가 하나씩 있고 전경 하단에는 날개로 보이는 면이 걸린다. 신부는 출입구 바로 안쪽 통로, 청년 둘은 왼쪽 전경에 있다. 내부에서 외부 접근로를 보는 공간 관계와 바닥 기준은 명확하지만 출입구를 보는 각도는 비교적 정면이다. 항공기의 줄무늬와 배치는 이전 사진과 완전히 같지는 않다.",
        "entities": "총 세 명이며 추가 인물은 없다. 신부는 짧은 회색 머리와 주름진 얼굴의 고령 동아시아계 남성으로, 참조와 대체로 맞고 흰 성직자 칼라와 검은 옷이 보인다. 목걸이 줄은 보이지만 십자가 형태는 손과 옷에 가려 확인하기 어렵다. 현우는 헝클어진 검은 머리, 앳되고 더러운 얼굴, 낡은 회색 셔츠로 이전 사진의 외형을 따른다. 쿠마는 짧은 검은 머리의 젊은 동아시아계 남성이며 남색 상의를 입었다. 혼혈 배경 자체는 외모만으로 확인할 수 없다. 비행기는 먼지와 손상이 있는 실물로 보이며, 모닥불은 프레임에 없어 상태를 판단할 수 없다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부의 앞쪽 부츠는 바닥 바로 위에 있고 뒤쪽 발은 무릎을 굽혀 들려 있다. 접지 순간보다는 달리기의 짧은 공중 국면에 가깝지만, 교차한 다리와 팔 동작 및 앞발 아래 착지 공간이 있어 근거 없는 부유로 보이지 않는다. 다만 한쪽 발의 바닥 접촉은 명확하지 않다. 옷자락은 몸에 연결된 채 뒤로 펼쳐져 있다. 두 청년의 하체는 일부 가려졌지만 지면에 서 있는 배치이며, 항공기는 보이는 바퀴와 착륙장치로 지지된다."
       },
       {
        "label": "B",
        "direction": "신부는 화면 중간 오른쪽에서 왼쪽 전경의 두 청년 쪽으로 접근하며 얼굴과 시선도 그들을 향한다. 현우와 쿠마는 등을 카메라 쪽에 두고 신부를 바라본다. 따라서 달려오는 사람과 도착 대상이 직접 연결된다. 긴 겉옷 자락은 진행 반대쪽인 화면 오른쪽 뒤로 날린다.",
        "built_space": "철골과 낡은 금속 벽, 왼쪽 상단 창, 오른쪽 배경의 큰 열린 출입구 한 곳이 보인다. 왼쪽 내부에는 낡은 경비행기 한 대와 바퀴가 보이고, 출입구 밖에는 다른 격납고와 어두운 항공기 형상이 보인다. 현우와 쿠마는 왼쪽 전경, 신부는 그들 앞의 빈 바닥 통로에 있어 접근 동선이 성립한다. 내부에서 입구를 비스듬히 보는 와이드 구도이며 신부의 머리부터 발까지 들어온다. 다만 바깥 격납고는 이전 사진에서 확인되지 않은 배경이고, 기존 장소의 세부 일치도는 제한적으로만 검증할 수 있다.",
        "entities": "보이는 인물은 신부·현우·쿠마에 해당하는 세 명뿐이다. 신부는 참조와 유사한 회색 머리와 고령의 동아시아계 남성 얼굴, 흰 성직자 칼라를 갖췄다. 참조의 짧은 상의보다 긴 검은 겉옷을 입었고 십자가는 확인되지 않는다. 현우의 뒷머리와 심하게 더러워진 회색 셔츠는 이전 사진과 대체로 이어지며 얼굴과 상처는 뒷모습이라 평가할 수 없다. 쿠마는 짧은 검은 머리의 젊은 남성으로 보이나 얼굴 대부분이 가려져 정확한 동일성은 확인하기 어렵다. 남색 상의는 참조와 색이 비슷하지만 긴소매다. 낡은 항공기는 보이고 모닥불은 화면 밖이다. 항공기 표면에 잘린 표식은 있으나 명확히 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "신부는 뒤쪽 무릎을 굽히고 앞발을 착지 방향으로 내민 달리기의 공중 국면이다. 두 부츠 모두 바닥과 약간 떨어져 보이지만 추진을 잇는 다리 동작과 바로 아래의 착지면이 명확해 불가능한 부유는 아니다. 한 발을 확실히 접지한 순간은 아니라는 작은 차이는 있다. 팔은 달리기에 맞춰 굽혀져 있고 겉옷은 어깨와 몸통에 연결되어 뒤로 펄럭인다. 두 청년은 바닥에 선 수직 자세이며 발은 하단 크롭 밖에 있다. 왼쪽 비행기는 보이는 착륙장치와 바퀴로 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "신부의 전신과 입구 배치는 맞지만, 세 사람의 시선이 주로 카메라 쪽을 향해 현우·쿠마에게 달려오는 관계가 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "전경의 현우·쿠마를 향하는 신부의 시선과 달리기, 중간 오른쪽 전신, 뒤로 날리는 옷자락이 지시된 순간을 더 충실하게 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부는 입구에서 화면 앞쪽으로 달려오며 얼굴과 시선도 거의 카메라를 향한다. 왼쪽 전경의 현우와 쿠마를 향한 뚜렷한 대각선 접근은 약하다. 두 청년도 신부를 돌아보기보다는 카메라 쪽을 보고 있어 세 사람의 시선이 연결되지 않는다. 옷자락은 신부의 화면 오른쪽 뒤로 날린다.",
        "built_space": "금속 벽과 철골, 왼쪽 상단 창, 중앙 오른쪽의 열린 대형 출입구 한 곳과 양옆 문짝이 보인다. 내부 양쪽에 낡은 경비행기 동체가 하나씩 있고 전경 하단에는 날개로 보이는 면이 걸린다. 신부는 출입구 바로 안쪽 통로, 청년 둘은 왼쪽 전경에 있다. 내부에서 외부 접근로를 보는 공간 관계와 바닥 기준은 명확하지만 출입구를 보는 각도는 비교적 정면이다. 항공기의 줄무늬와 배치는 이전 사진과 완전히 같지는 않다.",
        "entities": "총 세 명이며 추가 인물은 없다. 신부는 짧은 회색 머리와 주름진 얼굴의 고령 동아시아계 남성으로, 참조와 대체로 맞고 흰 성직자 칼라와 검은 옷이 보인다. 목걸이 줄은 보이지만 십자가 형태는 손과 옷에 가려 확인하기 어렵다. 현우는 헝클어진 검은 머리, 앳되고 더러운 얼굴, 낡은 회색 셔츠로 이전 사진의 외형을 따른다. 쿠마는 짧은 검은 머리의 젊은 동아시아계 남성이며 남색 상의를 입었다. 혼혈 배경 자체는 외모만으로 확인할 수 없다. 비행기는 먼지와 손상이 있는 실물로 보이며, 모닥불은 프레임에 없어 상태를 판단할 수 없다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신부의 앞쪽 부츠는 바닥 바로 위에 있고 뒤쪽 발은 무릎을 굽혀 들려 있다. 접지 순간보다는 달리기의 짧은 공중 국면에 가깝지만, 교차한 다리와 팔 동작 및 앞발 아래 착지 공간이 있어 근거 없는 부유로 보이지 않는다. 다만 한쪽 발의 바닥 접촉은 명확하지 않다. 옷자락은 몸에 연결된 채 뒤로 펼쳐져 있다. 두 청년의 하체는 일부 가려졌지만 지면에 서 있는 배치이며, 항공기는 보이는 바퀴와 착륙장치로 지지된다."
       },
       {
        "label": "A",
        "direction": "신부는 화면 중간 오른쪽에서 왼쪽 전경의 두 청년 쪽으로 접근하며 얼굴과 시선도 그들을 향한다. 현우와 쿠마는 등을 카메라 쪽에 두고 신부를 바라본다. 따라서 달려오는 사람과 도착 대상이 직접 연결된다. 긴 겉옷 자락은 진행 반대쪽인 화면 오른쪽 뒤로 날린다.",
        "built_space": "철골과 낡은 금속 벽, 왼쪽 상단 창, 오른쪽 배경의 큰 열린 출입구 한 곳이 보인다. 왼쪽 내부에는 낡은 경비행기 한 대와 바퀴가 보이고, 출입구 밖에는 다른 격납고와 어두운 항공기 형상이 보인다. 현우와 쿠마는 왼쪽 전경, 신부는 그들 앞의 빈 바닥 통로에 있어 접근 동선이 성립한다. 내부에서 입구를 비스듬히 보는 와이드 구도이며 신부의 머리부터 발까지 들어온다. 다만 바깥 격납고는 이전 사진에서 확인되지 않은 배경이고, 기존 장소의 세부 일치도는 제한적으로만 검증할 수 있다.",
        "entities": "보이는 인물은 신부·현우·쿠마에 해당하는 세 명뿐이다. 신부는 참조와 유사한 회색 머리와 고령의 동아시아계 남성 얼굴, 흰 성직자 칼라를 갖췄다. 참조의 짧은 상의보다 긴 검은 겉옷을 입었고 십자가는 확인되지 않는다. 현우의 뒷머리와 심하게 더러워진 회색 셔츠는 이전 사진과 대체로 이어지며 얼굴과 상처는 뒷모습이라 평가할 수 없다. 쿠마는 짧은 검은 머리의 젊은 남성으로 보이나 얼굴 대부분이 가려져 정확한 동일성은 확인하기 어렵다. 남색 상의는 참조와 색이 비슷하지만 긴소매다. 낡은 항공기는 보이고 모닥불은 화면 밖이다. 항공기 표면에 잘린 표식은 있으나 명확히 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "신부는 뒤쪽 무릎을 굽히고 앞발을 착지 방향으로 내민 달리기의 공중 국면이다. 두 부츠 모두 바닥과 약간 떨어져 보이지만 추진을 잇는 다리 동작과 바로 아래의 착지면이 명확해 불가능한 부유는 아니다. 한 발을 확실히 접지한 순간은 아니라는 작은 차이는 있다. 팔은 달리기에 맞춰 굽혀져 있고 겉옷은 어깨와 몸통에 연결되어 뒤로 펄럭인다. 두 청년은 바닥에 선 수직 자세이며 발은 하단 크롭 밖에 있다. 왼쪽 비행기는 보이는 착륙장치와 바퀴로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.194
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.944
   },
   "violations": {
    "B": [
     "[gemini-pro] invented objects (화면 하단을 크게 가로막고 있는 정체불명의 거대한 전경 구조물)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 944
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "대각선 앵글과 우측 출입구 배치, 현우와 쿠마를 향해 달려오는 신부의 동선과 겉옷의 휘날림까지 프롬프트의 세부 연출 지시를 매우 훌륭하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 944,
    "verdict_ko": "인물들이 카메라를 응시하여 '향해 뛰어오는' 동선의 개연성을 깨뜨렸고, 대각선 구도 지시를 무시한 채 정면 앵글과 불필요한 전경 물체를 배치해 지시에서 크게 벗어났습니다.  ★위반: [gemini-pro] invented objects (화면 하단을 크게 가로막고 있는 정체불명의 거대한 전경 구조물)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5_sel.png",
    "asset_id": "dc271dc3-8964-47de-87f3-4c3068e964c2",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:863986>",
    "asset_id": "87f8936b-1f27-4182-b09b-9f6277b7d71a",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1218434>",
    "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-c4b9-7c94-ba5a-b7781330f967",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S75sh5"
  },
  "lane_policy": "ab_select_bypass:prev",
  "staged_characters_added": [
   "C01",
   "C13"
  ]
 },
 "S75sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:36:44.944687+00:00",
  "fingerprint": "c5fde4d0e60718495cac888feba20dd5b8e577b8b68b4132245d1d3bd869592b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S75sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S75sh10_sel.png",
  "source_sha256": "6490a371bb74934f877435cdc10883df98bd0f67129e23a6560a66eeaf619a54",
  "file": "S75sh10_cine.png",
  "staged_sha256": "b92d9fabbfcd9df2d85d563b28c09e27190d41fd56495352477ddfb0f9d21d15",
  "latency_ms": 8848
 },
 "S76sh1::signage": {
  "fp": "8e9414617eb6cc5c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S76sh1": {
  "input_fingerprint": "fdb7c383c07c275b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보건소 침대에 누워 식은땀을 뻘뻘 흘리며 찡그린 앰버의 고통스러운 얼굴 클로즈업.\n\nLOCATION (lock): On the child's bed inside the harbor-town clinic, under nighttime treatment-room lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버) — A narrow part of the supporting surface is visible beneath her head and shoulder; used as Minimal physical context for the isolated face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet interior ambient illumination gently reveals 앰버's cold sweat and strained expression without stylized fever colors or perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the clinic bed, surrounding room surfaces, and the established nighttime lighting. Exclude aircraft and hangar equipment from the intervening airport scene.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed is still in use at night. Charlie retains his unrepaired dents, punctures and exposed chest opening. 앰버: She remains in bed after emergency treatment of her side wound, now visibly sweating with a high fever.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보건소 침대에 누워 식은땀을 뻘뻘 흘리며 찡그린 앰버의 고통스러운 얼굴 클로즈업.\n\nLOCATION (lock): On the child's bed inside the harbor-town clinic, under nighttime treatment-room lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버) — A narrow part of the supporting surface is visible beneath her head and shoulder; used as Minimal physical context for the isolated face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet interior ambient illumination gently reveals 앰버's cold sweat and strained expression without stylized fever colors or perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the clinic bed, surrounding room surfaces, and the established nighttime lighting. Exclude aircraft and hangar equipment from the intervening airport scene.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed is still in use at night. Charlie retains his unrepaired dents, punctures and exposed chest opening. 앰버: She remains in bed after emergency treatment of her side wound, now visibly sweating with a high fever.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보건소 침대에 누워 식은땀을 뻘뻘 흘리며 찡그린 앰버의 고통스러운 얼굴 클로즈업.\n\nLOCATION (lock): On the child's bed inside the harbor-town clinic, under nighttime treatment-room lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버) — A narrow part of the supporting surface is visible beneath her head and shoulder; used as Minimal physical context for the isolated face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet interior ambient illumination gently reveals 앰버's cold sweat and strained expression without stylized fever colors or perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the clinic bed, surrounding room surfaces, and the established nighttime lighting. Exclude aircraft and hangar equipment from the intervening airport scene.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed is still in use at night. Charlie retains his unrepaired dents, punctures and exposed chest opening. 앰버: She remains in bed after emergency treatment of her side wound, now visibly sweating with a high fever.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 눈을 질끈 감고 있어 특정한 곳을 응시하지 않음.",
    "built_space": "환자용 침대의 베개 위에 머리가 놓여 있으며, 배경에는 레퍼런스와 일치하는 야간의 창문과 병실 구조가 아웃포커싱되어 보임.",
    "entities": "금발의 10세 소녀 앰버가 땀에 흠뻑 젖은 채로 얼굴을 찡그리며 고통스러워하는 모습이 정확히 묘사됨. 프롬프트에 없는 다른 인물은 없음.",
    "hard_violations": [],
    "physics": "머리가 베개 위에 자연스럽게 눕혀져 있으며 중력의 영향을 정상적으로 받고 있음."
   },
   {
    "label": "B",
    "direction": "두 눈을 감고 있음.",
    "built_space": "베개와 침대, 그리고 배경의 병실 벽면 의료용 콘센트 및 수액 걸이가 레퍼런스의 구조와 일치하게 배치됨.",
    "entities": "앰버가 식은땀을 흘리고 있으나, 찡그리거나 고통스러운 표정보다는 비교적 평온하게 눈을 감고 있는 모습에 가까움.",
    "hard_violations": [],
    "physics": "머리와 목이 베개와 침대에 정상적으로 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트에서 요구한 '식은땀을 뻘뻘 흘리며 찡그린 고통스러운 얼굴'을 훌륭하게 묘사했으며, 카메라 구도와 조명 역시 레퍼런스의 분위기를 잘 유지하고 있습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "식은땀을 흘리는 모습은 표현되었으나, 표정이 비교적 평온하여 프롬프트의 '찡그린 고통스러운 얼굴'이라는 핵심 지시사항을 충분히 충족시키지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 질끈 감고 있어 특정한 곳을 응시하지 않음.",
        "built_space": "환자용 침대의 베개 위에 머리가 놓여 있으며, 배경에는 레퍼런스와 일치하는 야간의 창문과 병실 구조가 아웃포커싱되어 보임.",
        "entities": "금발의 10세 소녀 앰버가 땀에 흠뻑 젖은 채로 얼굴을 찡그리며 고통스러워하는 모습이 정확히 묘사됨. 프롬프트에 없는 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "머리가 베개 위에 자연스럽게 눕혀져 있으며 중력의 영향을 정상적으로 받고 있음."
       },
       {
        "label": "B",
        "direction": "두 눈을 감고 있음.",
        "built_space": "베개와 침대, 그리고 배경의 병실 벽면 의료용 콘센트 및 수액 걸이가 레퍼런스의 구조와 일치하게 배치됨.",
        "entities": "앰버가 식은땀을 흘리고 있으나, 찡그리거나 고통스러운 표정보다는 비교적 평온하게 눈을 감고 있는 모습에 가까움.",
        "hard_violations": [],
        "physics": "머리와 목이 베개와 침대에 정상적으로 지지되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트에서 요구한 '식은땀을 뻘뻘 흘리며 찡그린 고통스러운 얼굴'을 훌륭하게 묘사했으며, 카메라 구도와 조명 역시 레퍼런스의 분위기를 잘 유지하고 있습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "식은땀을 흘리는 모습은 표현되었으나, 표정이 비교적 평온하여 프롬프트의 '찡그린 고통스러운 얼굴'이라는 핵심 지시사항을 충분히 충족시키지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 질끈 감고 있어 특정한 곳을 응시하지 않음.",
        "built_space": "환자용 침대의 베개 위에 머리가 놓여 있으며, 배경에는 레퍼런스와 일치하는 야간의 창문과 병실 구조가 아웃포커싱되어 보임.",
        "entities": "금발의 10세 소녀 앰버가 땀에 흠뻑 젖은 채로 얼굴을 찡그리며 고통스러워하는 모습이 정확히 묘사됨. 프롬프트에 없는 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "머리가 베개 위에 자연스럽게 눕혀져 있으며 중력의 영향을 정상적으로 받고 있음."
       },
       {
        "label": "B",
        "direction": "두 눈을 감고 있음.",
        "built_space": "베개와 침대, 그리고 배경의 병실 벽면 의료용 콘센트 및 수액 걸이가 레퍼런스의 구조와 일치하게 배치됨.",
        "entities": "앰버가 식은땀을 흘리고 있으나, 찡그리거나 고통스러운 표정보다는 비교적 평온하게 눈을 감고 있는 모습에 가까움.",
        "hard_violations": [],
        "physics": "머리와 목이 베개와 침대에 정상적으로 지지되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴 클로즈업과 식은땀, 침대에 지지된 자세는 충실하지만, 찡그림이 약해 요구된 고통의 순간이 B보다 덜 분명하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 중심 클로즈업에서 깊게 모인 미간과 긴장한 입매, 풍부한 식은땀이 보이며 야간 병상의 지지 관계와 외형도 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 위쪽과 화면 오른쪽을 향하고 두 눈은 감겨 있어 응시 대상이 없다. 누워 쉬는 장면에 맞으며, 방향을 확인할 무기나 손에 든 물건은 없다.",
        "built_space": "흰 베개와 침대 지지면이 머리와 어깨 아래에 보인다. 뒤쪽에는 침대 머리판 하나, 벽 조명 하나, 의료용 벽 패널 하나, 흐릿한 장치 하나와 오른쪽 창 일부가 보인다. 기준 사진의 병실 재료와 따뜻한 야간 조명이 이어지며 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "금발의 둥근 얼굴을 가진 어린 여자아이 한 명으로, 기준 사진의 앰버와 외형이 대체로 일치한다. 혼혈 여부는 외형만으로 확정할 수 없고 큰 눈은 감겨 있어 확인되지 않는다. 짙은 남색 상의와 하단의 푸른 담요, 흰 침구가 유지된다. 이마·볼·목에 땀이 맺혀 있으나 미간과 입매의 고통 표현은 비교적 약하다. 다른 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 머리카락은 베개에 놓여 있고 목과 상체는 이어지는 침대에 받쳐져 있다. 들려 있는 팔다리나 떠 있는 물체는 없으며, 땀방울은 피부에 붙어 빛을 반사한다."
       },
       {
        "label": "B",
        "direction": "얼굴은 위쪽을 향한 채 카메라 쪽으로 비스듬히 보인다. 눈꺼풀은 거의 닫혀 특정 대상을 응시하지 않는다. 모인 눈썹과 굳은 입매가 통증에 찡그리는 동작을 분명하게 보여 준다.",
        "built_space": "머리 아래 흰 베개와 그 뒤쪽 흰 침구, 침대 머리판 하나 및 오른쪽 난간 일부가 보인다. 배경에는 병실 벽과 어두운 창, 흐릿한 병상 주변 용품이 남아 있다. 얼굴 위주의 촬영으로 나머지 설비가 제외된 구성이며, 기준 사진과 충돌하는 설비 중복이나 반사는 보이지 않는다.",
        "entities": "금발의 둥근 얼굴을 가진 어린 여자아이 한 명이며 기준 사진의 앰버와 연령대·머리색·얼굴 윤곽이 대체로 맞는다. 혼혈 여부는 외형만으로 단정할 수 없고 눈 크기는 거의 감긴 상태라 확인하기 어렵다. 남색 상의, 푸른 담요 가장자리와 흰 침구가 보인다. 이마와 볼, 목의 식은땀 및 강한 미간 주름이 요구된 발열과 고통을 드러낸다. 다른 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 머리 옆면이 베개에 안정적으로 놓이고, 머리카락도 베개와 어깨 위에 늘어져 있다. 목과 어깨는 침대에 지지되어 있으며 공중에 유지되는 신체 부위나 물체는 없다. 얼굴의 근육 긴장은 요구된 찡그림으로 가능한 범위다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴 클로즈업과 식은땀, 침대에 지지된 자세는 충실하지만, 찡그림이 약해 요구된 고통의 순간이 B보다 덜 분명하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 중심 클로즈업에서 깊게 모인 미간과 긴장한 입매, 풍부한 식은땀이 보이며 야간 병상의 지지 관계와 외형도 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 위쪽과 화면 오른쪽을 향하고 두 눈은 감겨 있어 응시 대상이 없다. 누워 쉬는 장면에 맞으며, 방향을 확인할 무기나 손에 든 물건은 없다.",
        "built_space": "흰 베개와 침대 지지면이 머리와 어깨 아래에 보인다. 뒤쪽에는 침대 머리판 하나, 벽 조명 하나, 의료용 벽 패널 하나, 흐릿한 장치 하나와 오른쪽 창 일부가 보인다. 기준 사진의 병실 재료와 따뜻한 야간 조명이 이어지며 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "금발의 둥근 얼굴을 가진 어린 여자아이 한 명으로, 기준 사진의 앰버와 외형이 대체로 일치한다. 혼혈 여부는 외형만으로 확정할 수 없고 큰 눈은 감겨 있어 확인되지 않는다. 짙은 남색 상의와 하단의 푸른 담요, 흰 침구가 유지된다. 이마·볼·목에 땀이 맺혀 있으나 미간과 입매의 고통 표현은 비교적 약하다. 다른 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 머리카락은 베개에 놓여 있고 목과 상체는 이어지는 침대에 받쳐져 있다. 들려 있는 팔다리나 떠 있는 물체는 없으며, 땀방울은 피부에 붙어 빛을 반사한다."
       },
       {
        "label": "A",
        "direction": "얼굴은 위쪽을 향한 채 카메라 쪽으로 비스듬히 보인다. 눈꺼풀은 거의 닫혀 특정 대상을 응시하지 않는다. 모인 눈썹과 굳은 입매가 통증에 찡그리는 동작을 분명하게 보여 준다.",
        "built_space": "머리 아래 흰 베개와 그 뒤쪽 흰 침구, 침대 머리판 하나 및 오른쪽 난간 일부가 보인다. 배경에는 병실 벽과 어두운 창, 흐릿한 병상 주변 용품이 남아 있다. 얼굴 위주의 촬영으로 나머지 설비가 제외된 구성이며, 기준 사진과 충돌하는 설비 중복이나 반사는 보이지 않는다.",
        "entities": "금발의 둥근 얼굴을 가진 어린 여자아이 한 명이며 기준 사진의 앰버와 연령대·머리색·얼굴 윤곽이 대체로 맞는다. 혼혈 여부는 외형만으로 단정할 수 없고 눈 크기는 거의 감긴 상태라 확인하기 어렵다. 남색 상의, 푸른 담요 가장자리와 흰 침구가 보인다. 이마와 볼, 목의 식은땀 및 강한 미간 주름이 요구된 발열과 고통을 드러낸다. 다른 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 머리 옆면이 베개에 안정적으로 놓이고, 머리카락도 베개와 어깨 위에 늘어져 있다. 목과 어깨는 침대에 지지되어 있으며 공중에 유지되는 신체 부위나 물체는 없다. 얼굴의 근육 긴장은 요구된 찡그림으로 가능한 범위다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.556
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.556
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1556
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트에서 요구한 '식은땀을 뻘뻘 흘리며 찡그린 고통스러운 얼굴'을 훌륭하게 묘사했으며, 카메라 구도와 조명 역시 레퍼런스의 분위기를 잘 유지하고 있습니다."
   },
   {
    "label": "B",
    "score": 1556,
    "verdict_ko": "식은땀을 흘리는 모습은 표현되었으나, 표정이 비교적 평온하여 프롬프트의 '찡그린 고통스러운 얼굴'이라는 핵심 지시사항을 충분히 충족시키지 못했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2_sel.png",
    "asset_id": "ae7edd90-60df-4f9c-a673-266c884b31f9",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-c674-7b61-b4f3-c110a2de49e9",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S73sh2"
  },
  "locked_char_refs_excluded": [
   "앰버(C03)"
  ]
 },
 "S76sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:37:36.143563+00:00",
  "fingerprint": "c58aea2bbf7754fddcd3e9a78cebe161f7f2b7488e402469e93e13dfcdff9282",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S76sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S76sh1_sel.png",
  "source_sha256": "54171ea58122b90f5abc20a2a3aea8645d21d609646269522bb24c7d19957f04",
  "file": "S76sh1_cine.png",
  "staged_sha256": "5705143aaea1f44e64afb58fc75ba2a2964fb5546b66128f76d47effe93bb229",
  "latency_ms": 10361
 },
 "S76sh6::signage": {
  "fp": "77311fab223af5ca",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S76sh6": {
  "input_fingerprint": "d80b56fd81906051",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 중년 의사의 옷소매를 양손으로 꽉 틀어쥔 현우의 절박한 상체.\n\nLOCATION (lock): Beside the sick child's bed inside the clinic patient room, under ordinary nighttime clinical lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Doctor's sleeves (Clenched in both of 현우's hands) — The near sleeve runs along the right edge while both grasped areas remain visible below 현우's face; used as Physical evidence of the plea, linked directly to the facial performance; Clinic bed edge (Beside the conversation) — A small partial edge remains in the lower background without revealing the patient; used as Continuity anchor for the bedside camera path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the clinic's neutral ambient illumination and controlled contrast, keeping the pleading face and tightly gripping hands equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic furnishings, bed, and nighttime interior lighting. Exclude the airport's aircraft, tools, and warehouse fixtures.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied. Charlie's previously damaged bodywork remains unrepaired. 현우: He remains in the health center, visibly dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 중년 의사의 옷소매를 양손으로 꽉 틀어쥔 현우의 절박한 상체.\n\nLOCATION (lock): Beside the sick child's bed inside the clinic patient room, under ordinary nighttime clinical lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Doctor's sleeves (Clenched in both of 현우's hands) — The near sleeve runs along the right edge while both grasped areas remain visible below 현우's face; used as Physical evidence of the plea, linked directly to the facial performance; Clinic bed edge (Beside the conversation) — A small partial edge remains in the lower background without revealing the patient; used as Continuity anchor for the bedside camera path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the clinic's neutral ambient illumination and controlled contrast, keeping the pleading face and tightly gripping hands equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic furnishings, bed, and nighttime interior lighting. Exclude the airport's aircraft, tools, and warehouse fixtures.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied. Charlie's previously damaged bodywork remains unrepaired. 현우: He remains in the health center, visibly dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 중년 의사의 옷소매를 양손으로 꽉 틀어쥔 현우의 절박한 상체.\n\nLOCATION (lock): Beside the sick child's bed inside the clinic patient room, under ordinary nighttime clinical lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Doctor's sleeves (Clenched in both of 현우's hands) — The near sleeve runs along the right edge while both grasped areas remain visible below 현우's face; used as Physical evidence of the plea, linked directly to the facial performance; Clinic bed edge (Beside the conversation) — A small partial edge remains in the lower background without revealing the patient; used as Continuity anchor for the bedside camera path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the clinic's neutral ambient illumination and controlled contrast, keeping the pleading face and tightly gripping hands equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic furnishings, bed, and nighttime interior lighting. Exclude the airport's aircraft, tools, and warehouse fixtures.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied. Charlie's previously damaged bodywork remains unrepaired. 현우: He remains in the health center, visibly dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 프레임 밖 우측 상단의 의사를 향함.",
    "built_space": "병실 내부. 배경에 병상 테두리와 문이 보이며 야간 병실 조명이 적용됨.",
    "entities": "현우(레퍼런스와 일치하는 얼굴, 헤어스타일, 회색 셔츠, 흙투성이 분장). 우측에 의사의 옷소매.",
    "hard_violations": [],
    "physics": "현우의 양손이 옷소매를 단단히 쥐고 있으며, 앞으로 쏠린 상체의 체중과 손의 접촉이 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 프레임 밖 우측 상단의 의사를 향함.",
    "built_space": "병실 내부. 병상 테두리와 벽면이 보임.",
    "entities": "현우(얼굴과 헤어스타일은 일치하나 상의를 탈의함). 우측에 의사의 옷소매.",
    "hard_violations": [
     "[gemini-pro] 캐릭터 레퍼런스의 의상 지침 위반 (셔츠 누락 및 상의 탈의)"
    ],
    "physics": "양손으로 옷소매를 쥐고 있으나 아래쪽 손의 해부학적 연결과 파지가 다소 뭉개짐."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스에 주어진 캐릭터의 외형과 의상을 정확히 반영했으며, 양손으로 옷소매를 쥐는 절박한 동작을 자연스럽게 연출했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캐릭터 레퍼런스를 통해 확정된 의상(회색 셔츠)을 무시하고 상의를 탈의한 상태로 묘사하여 지침에 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 프레임 밖 우측 상단의 의사를 향함.",
        "built_space": "병실 내부. 배경에 병상 테두리와 문이 보이며 야간 병실 조명이 적용됨.",
        "entities": "현우(레퍼런스와 일치하는 얼굴, 헤어스타일, 회색 셔츠, 흙투성이 분장). 우측에 의사의 옷소매.",
        "hard_violations": [],
        "physics": "현우의 양손이 옷소매를 단단히 쥐고 있으며, 앞으로 쏠린 상체의 체중과 손의 접촉이 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 프레임 밖 우측 상단의 의사를 향함.",
        "built_space": "병실 내부. 병상 테두리와 벽면이 보임.",
        "entities": "현우(얼굴과 헤어스타일은 일치하나 상의를 탈의함). 우측에 의사의 옷소매.",
        "hard_violations": [
         "캐릭터 레퍼런스의 의상 지침 위반 (셔츠 누락 및 상의 탈의)"
        ],
        "physics": "양손으로 옷소매를 쥐고 있으나 아래쪽 손의 해부학적 연결과 파지가 다소 뭉개짐."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스에 주어진 캐릭터의 외형과 의상을 정확히 반영했으며, 양손으로 옷소매를 쥐는 절박한 동작을 자연스럽게 연출했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캐릭터 레퍼런스를 통해 확정된 의상(회색 셔츠)을 무시하고 상의를 탈의한 상태로 묘사하여 지침에 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 프레임 밖 우측 상단의 의사를 향함.",
        "built_space": "병실 내부. 배경에 병상 테두리와 문이 보이며 야간 병실 조명이 적용됨.",
        "entities": "현우(레퍼런스와 일치하는 얼굴, 헤어스타일, 회색 셔츠, 흙투성이 분장). 우측에 의사의 옷소매.",
        "hard_violations": [],
        "physics": "현우의 양손이 옷소매를 단단히 쥐고 있으며, 앞으로 쏠린 상체의 체중과 손의 접촉이 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 프레임 밖 우측 상단의 의사를 향함.",
        "built_space": "병실 내부. 병상 테두리와 벽면이 보임.",
        "entities": "현우(얼굴과 헤어스타일은 일치하나 상의를 탈의함). 우측에 의사의 옷소매.",
        "hard_violations": [
         "캐릭터 레퍼런스의 의상 지침 위반 (셔츠 누락 및 상의 탈의)"
        ],
        "physics": "양손으로 옷소매를 쥐고 있으나 아래쪽 손의 해부학적 연결과 파지가 다소 뭉개짐."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "절박한 표정과 양손의 소매 움켜쥠은 명확하지만, 현우가 상의를 벗고 있어 참조의 회색 셔츠 차림을 크게 벗어난다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 상체 중심 미디엄 숏에서 의사를 올려다보는 시선, 얼굴 아래 양손의 소매 잡기, 회색 셔츠와 더러워진 상태를 함께 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 오른쪽 위, 화면 밖에 있는 의사의 얼굴 위치를 올려다본다. 양손은 오른쪽에서 내려오는 의사의 같은 소매를 얼굴 아래 서로 다른 지점에서 움켜쥐고 있어 애원의 대상과 행동 방향이 일치한다.",
        "built_space": "현우 뒤 하단에 병상 하나의 금속 난간 일부와 밝은색 침대 끝판이 보인다. 오른쪽에는 의사의 팔과 몸통 일부가 있고 환자는 노출되지 않는다. 밝은 벽과 침구는 참조의 병실 재질에 대체로 부합하나, 참조에서 보이지 않았던 침대 끝판의 정확한 동일성은 확인할 수 없다. 반사나 중복 병상은 없다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리와 얼굴의 윤곽이 참조에 대체로 맞는다. 얼굴과 맨몸에는 흙먼지와 긁힌 흔적이 뚜렷하지만, 참조의 회색 셔츠가 사라지고 상체가 벗겨져 있다. 의사는 밝은색 긴소매 옷과 손 일부만 보여 중년 여부와 얼굴 정체성은 확인할 수 없다. 아이와 찰리는 화면 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 두 손가락이 실제로 소매 천을 감싸고 있으며 잡힌 부분을 중심으로 천에 당김 주름이 생긴다. 의사의 소매는 오른쪽 몸통과 이어지고 손은 소매 끝 아래로 내려와 있어 떠 있는 옷이 아니다. 현우의 하체와 지지면은 잘려 있지만 상체 자세에 불가능한 부유나 관절 연결은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우의 눈은 오른쪽 위 의사의 화면 밖 얼굴을 향한다. 두 손은 얼굴 아래에서 의사의 소매 끝 부근을 양쪽으로 움켜쥐며, 팔과 상체도 그 의사 쪽으로 기울어 있다. 카메라를 바라보지 않는다.",
        "built_space": "하단 배경에 병상 하나의 흰 침구와 밝은색 끝판 일부가 보이고 환자는 가려져 있다. 뒤에는 목재 문 하나와 벽면 설비 패널 하나가 보인다. 현우는 병상 바로 옆, 의사는 오른쪽 전경에 있어 대화 위치가 자연스럽다. 참조의 밝은 벽·침구 및 따뜻한 기가 약간 도는 실내 조명과 대체로 이어지며, 설비 중복이나 불가능한 반사는 없다.",
        "entities": "현우의 앳된 동아시아계 남성 외모, 검은 머리, 마른 체격과 회색 단추 셔츠가 참조에 잘 맞는다. 얼굴과 손의 오염 및 상처도 유지된다. 의사는 회색 긴소매 겉옷을 입은 몸통·팔·손 일부만 보여 직업과 중년 나이를 외모만으로 확정할 수는 없다. 아이는 드러나지 않고 찰리도 프레임 밖이다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "현우의 양손이 의사 소매를 직접 감싸며 천이 손 사이에서 조여진다. 의사의 손은 소매 끝에서 자연스럽게 이어지고 팔은 오른쪽 몸통에 연결된다. 현우의 팔꿈치와 손목 굽힘은 소매를 붙잡아 당기는 동작으로 가능하다. 하체 지지는 프레임 밖이라 확인되지 않지만, 공중에 뜬 몸이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "절박한 표정과 양손의 소매 움켜쥠은 명확하지만, 현우가 상의를 벗고 있어 참조의 회색 셔츠 차림을 크게 벗어난다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우의 상체 중심 미디엄 숏에서 의사를 올려다보는 시선, 얼굴 아래 양손의 소매 잡기, 회색 셔츠와 더러워진 상태를 함께 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 오른쪽 위, 화면 밖에 있는 의사의 얼굴 위치를 올려다본다. 양손은 오른쪽에서 내려오는 의사의 같은 소매를 얼굴 아래 서로 다른 지점에서 움켜쥐고 있어 애원의 대상과 행동 방향이 일치한다.",
        "built_space": "현우 뒤 하단에 병상 하나의 금속 난간 일부와 밝은색 침대 끝판이 보인다. 오른쪽에는 의사의 팔과 몸통 일부가 있고 환자는 노출되지 않는다. 밝은 벽과 침구는 참조의 병실 재질에 대체로 부합하나, 참조에서 보이지 않았던 침대 끝판의 정확한 동일성은 확인할 수 없다. 반사나 중복 병상은 없다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리와 얼굴의 윤곽이 참조에 대체로 맞는다. 얼굴과 맨몸에는 흙먼지와 긁힌 흔적이 뚜렷하지만, 참조의 회색 셔츠가 사라지고 상체가 벗겨져 있다. 의사는 밝은색 긴소매 옷과 손 일부만 보여 중년 여부와 얼굴 정체성은 확인할 수 없다. 아이와 찰리는 화면 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 두 손가락이 실제로 소매 천을 감싸고 있으며 잡힌 부분을 중심으로 천에 당김 주름이 생긴다. 의사의 소매는 오른쪽 몸통과 이어지고 손은 소매 끝 아래로 내려와 있어 떠 있는 옷이 아니다. 현우의 하체와 지지면은 잘려 있지만 상체 자세에 불가능한 부유나 관절 연결은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우의 눈은 오른쪽 위 의사의 화면 밖 얼굴을 향한다. 두 손은 얼굴 아래에서 의사의 소매 끝 부근을 양쪽으로 움켜쥐며, 팔과 상체도 그 의사 쪽으로 기울어 있다. 카메라를 바라보지 않는다.",
        "built_space": "하단 배경에 병상 하나의 흰 침구와 밝은색 끝판 일부가 보이고 환자는 가려져 있다. 뒤에는 목재 문 하나와 벽면 설비 패널 하나가 보인다. 현우는 병상 바로 옆, 의사는 오른쪽 전경에 있어 대화 위치가 자연스럽다. 참조의 밝은 벽·침구 및 따뜻한 기가 약간 도는 실내 조명과 대체로 이어지며, 설비 중복이나 불가능한 반사는 없다.",
        "entities": "현우의 앳된 동아시아계 남성 외모, 검은 머리, 마른 체격과 회색 단추 셔츠가 참조에 잘 맞는다. 얼굴과 손의 오염 및 상처도 유지된다. 의사는 회색 긴소매 겉옷을 입은 몸통·팔·손 일부만 보여 직업과 중년 나이를 외모만으로 확정할 수는 없다. 아이는 드러나지 않고 찰리도 프레임 밖이다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "현우의 양손이 의사 소매를 직접 감싸며 천이 손 사이에서 조여진다. 의사의 손은 소매 끝에서 자연스럽게 이어지고 팔은 오른쪽 몸통에 연결된다. 현우의 팔꿈치와 손목 굽힘은 소매를 붙잡아 당기는 동작으로 가능하다. 하체 지지는 프레임 밖이라 확인되지 않지만, 공중에 뜬 몸이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.349
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.099
   },
   "violations": {
    "B": [
     "[gemini-pro] 캐릭터 레퍼런스의 의상 지침 위반 (셔츠 누락 및 상의 탈의)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1099
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스에 주어진 캐릭터의 외형과 의상을 정확히 반영했으며, 양손으로 옷소매를 쥐는 절박한 동작을 자연스럽게 연출했습니다."
   },
   {
    "label": "B",
    "score": 1099,
    "verdict_ko": "캐릭터 레퍼런스를 통해 확정된 의상(회색 셔츠)을 무시하고 상의를 탈의한 상태로 묘사하여 지침에 어긋납니다.  ★위반: [gemini-pro] 캐릭터 레퍼런스의 의상 지침 위반 (셔츠 누락 및 상의 탈의)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S76sh1_sel.png",
    "asset_id": "54c5c9c6-bf6d-4a2a-a7bb-ec8803fa3d2a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-c81d-77cb-a281-587eeb1beaef",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S76sh1"
  }
 },
 "S76sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:38:29.876717+00:00",
  "fingerprint": "666f144f55d212542300fcad6d9ff61f48f21f1ee128b4295c78b56f5eaa53e8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S76sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S76sh6_sel.png",
  "source_sha256": "b4aadc29f8779e73b35cb45b02e28a9147e084a3d946cee6059ece11ed89a013",
  "file": "S76sh6_cine.png",
  "staged_sha256": "92d57fea36a010b4a004734efb13464de6d2293fae7638cf1e3e2ecf8b3b76f2",
  "latency_ms": 14645
 },
 "S76sh12::signage": {
  "fp": "8b8ede2d2c856f26",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S76sh12": {
  "input_fingerprint": "a34a0f74be9e0446",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 커다란 금속 손을 번쩍 치켜든 찰리의 역동적인 상체.\n\nLOCATION (lock): In the group gathered beside the clinic bed, inside the nighttime-lit patient room. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Beside 찰리's position) — Only a small edge is retained at the lower margin, with the patient outside the framing; used as Spatial continuity with the bedside scene without competing with the raised arm.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same restrained clinic illumination, using controlled highlights to articulate the metal hand without adding an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged chest and bodywork have not yet been repaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 커다란 금속 손을 번쩍 치켜든 찰리의 역동적인 상체.\n\nLOCATION (lock): In the group gathered beside the clinic bed, inside the nighttime-lit patient room. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Beside 찰리's position) — Only a small edge is retained at the lower margin, with the patient outside the framing; used as Spatial continuity with the bedside scene without competing with the raised arm.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same restrained clinic illumination, using controlled highlights to articulate the metal hand without adding an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged chest and bodywork have not yet been repaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 커다란 금속 손을 번쩍 치켜든 찰리의 역동적인 상체.\n\nLOCATION (lock): In the group gathered beside the clinic bed, inside the nighttime-lit patient room. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Beside 찰리's position) — Only a small edge is retained at the lower margin, with the patient outside the framing; used as Spatial continuity with the bedside scene without competing with the raised arm.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same restrained clinic illumination, using controlled highlights to articulate the metal hand without adding an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged chest and bodywork have not yet been repaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 오른팔을 허공으로 뻗고 있으며 시선은 정면 아래를 향함.",
    "built_space": "병실 내부. 천장, 벽면 의료 패널, 나무 문이 보이며 하단에 침대와 환자의 일부가 나타남.",
    "entities": "찰리의 기본적인 외형은 일치하나 지시된 흉부 파손 흔적이 없음.",
    "hard_violations": [
     "[gemini-pro] 환자를 프레임 외부로 제외하라는 지시를 어기고 화면 하단에 환자를 포함함."
    ],
    "physics": "팔을 치켜든 자세와 로봇의 구조가 안정적으로 지지됨."
   },
   {
    "label": "B",
    "direction": "찰리가 오른팔을 허공으로 번쩍 들고 있으며 시선은 정면을 향함.",
    "built_space": "병실 내부. 천장, 수액 걸이, 의료 패널, 커튼이 있으며 화면 하단에 침대 프레임만 보임.",
    "entities": "찰리의 외형이 일치하며 지시된 흉부의 파손 흔적이 명확하게 묘사됨. 프레임 내에 환자가 없음.",
    "hard_violations": [],
    "physics": "동적이고 무거운 로봇의 상체 움직임이 자연스럽게 구현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "침대 모서리만 남긴 구도와 흉부 파손 상태 등 모든 프롬프트 지시사항을 충실히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "흉부 파손 상태가 누락되었으며, 환자를 프레임에서 제외하라는 명시적 구도 지시를 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 오른팔을 허공으로 뻗고 있으며 시선은 정면 아래를 향함.",
        "built_space": "병실 내부. 천장, 벽면 의료 패널, 나무 문이 보이며 하단에 침대와 환자의 일부가 나타남.",
        "entities": "찰리의 기본적인 외형은 일치하나 지시된 흉부 파손 흔적이 없음.",
        "hard_violations": [
         "환자를 프레임 외부로 제외하라는 지시를 어기고 화면 하단에 환자를 포함함."
        ],
        "physics": "팔을 치켜든 자세와 로봇의 구조가 안정적으로 지지됨."
       },
       {
        "label": "B",
        "direction": "찰리가 오른팔을 허공으로 번쩍 들고 있으며 시선은 정면을 향함.",
        "built_space": "병실 내부. 천장, 수액 걸이, 의료 패널, 커튼이 있으며 화면 하단에 침대 프레임만 보임.",
        "entities": "찰리의 외형이 일치하며 지시된 흉부의 파손 흔적이 명확하게 묘사됨. 프레임 내에 환자가 없음.",
        "hard_violations": [],
        "physics": "동적이고 무거운 로봇의 상체 움직임이 자연스럽게 구현됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "침대 모서리만 남긴 구도와 흉부 파손 상태 등 모든 프롬프트 지시사항을 충실히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "흉부 파손 상태가 누락되었으며, 환자를 프레임에서 제외하라는 명시적 구도 지시를 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 오른팔을 허공으로 뻗고 있으며 시선은 정면 아래를 향함.",
        "built_space": "병실 내부. 천장, 벽면 의료 패널, 나무 문이 보이며 하단에 침대와 환자의 일부가 나타남.",
        "entities": "찰리의 기본적인 외형은 일치하나 지시된 흉부 파손 흔적이 없음.",
        "hard_violations": [
         "환자를 프레임 외부로 제외하라는 지시를 어기고 화면 하단에 환자를 포함함."
        ],
        "physics": "팔을 치켜든 자세와 로봇의 구조가 안정적으로 지지됨."
       },
       {
        "label": "B",
        "direction": "찰리가 오른팔을 허공으로 번쩍 들고 있으며 시선은 정면을 향함.",
        "built_space": "병실 내부. 천장, 수액 걸이, 의료 패널, 커튼이 있으며 화면 하단에 침대 프레임만 보임.",
        "entities": "찰리의 외형이 일치하며 지시된 흉부의 파손 흔적이 명확하게 묘사됨. 프레임 내에 환자가 없음.",
        "hard_violations": [],
        "physics": "동적이고 무거운 로봇의 상체 움직임이 자연스럽게 구현됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "큰 금속 손을 허공으로 치켜든 역동적인 상체와 미수리 흉부 손상을 잘 구현했으나, 참조의 밀짚모자와 화려한 목도리가 빠졌다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "상체 중심 미디엄 숏과 위로 치켜든 손은 맞지만, 몸통의 동세가 약하고 흉부 손상이 표면 흠집 정도이며 참조의 모자와 목도리도 없다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 왼쪽의 큰 금속 손이 머리보다 높이 올라가 있으며 손끝은 천장 쪽 허공을 향한다. 손바닥은 카메라 쪽으로 비스듬히 열려 있다. 얼굴은 들린 팔 쪽으로 조금 돌아가 있지만 눈의 정확한 시선은 어두운 눈구멍 때문에 판별하기 어렵다. 허공으로 손을 치켜든다는 동작 방향은 맞는다.",
        "built_space": "베이지색 벽과 왼쪽 목재 문틀, 콘센트 패널 세 구획이 보이는 벽면 의료 설비대 하나, 벽등 하나가 있다. 왼쪽에는 수액대 하나, 뒤에는 모니터 하나, 오른쪽에는 천장 레일에 달린 커튼이 보인다. 침대의 먼 쪽 프레임과 가까운 쪽 난간이 하단에 걸리고 환자는 보이지 않는다. 찰리는 침대 너머 옆에 서 있는 배치로 읽힌다. 참조의 벽 재질과 의료 설비 분위기는 이어지지만, 수액대·모니터·커튼의 정확한 배치는 참조에서 확인할 수 없다. 거울이나 문제 될 반사는 없다.",
        "entities": "찰리 한 명만 보인다. 육중한 어깨, 긴 기계 팔, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴과 큰 금속 손은 참조의 정체성과 대체로 맞는다. 흉부 한쪽에는 깊게 벌어진 파손 자국이 있어 미수리 상태가 명확하다. 참조의 밀짚모자와 다색 목도리는 화면에 있어야 할 부위에서 누락됐다. 다리와 장화는 상체 프레이밍 밖이므로 평가 대상이 아니다. 다른 사람이나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "들린 손은 손목·전완·팔꿈치·상완을 통해 어깨에 연결되어 있고, 굽힌 팔꿈치와 기울어진 상체가 팔을 들어 올리는 동작을 뒷받침한다. 반대 팔도 몸통에 정상적으로 연결되어 있다. 하체와 발은 프레임 밖이지만 몸이 공중에 떠 있는 징후는 없다. 모니터는 받침 위에 있고 커튼은 천장 레일에 매달려 있어 지지 관계가 자연스럽다."
       },
       {
        "label": "B",
        "direction": "화면 왼쪽의 금속 손이 머리 위로 올라가고 전완과 손끝은 거의 수직으로 천장을 향한다. 열린 손바닥은 카메라를 향한다. 얼굴은 정면에서 약간 화면 왼쪽으로 돌아가 있으며, 망 형태의 눈 부분 때문에 정확한 시선 목표는 확인하기 어렵다. 손을 허공으로 치켜든 방향 자체는 정확하다.",
        "built_space": "베이지색 벽, 중앙 뒤쪽의 목재 문 또는 벽판 한 구획, 오른쪽 의료 설비대 하나와 벽등 하나가 보인다. 설비대에는 콘센트 패널 두 구획과 화면 끝에 잘린 추가 구획이 있다. 천장에는 사각 설비판 하나와 원형 설비 하나가 보인다. 침대 난간과 이불의 작은 가장자리만 하단에 남아 있고 환자는 프레임 밖이다. 찰리는 침대 옆에 선 것으로 읽히며, 참조의 목재·벽·의료 설비 재질과 대체로 이어진다. 불가능한 반사는 없다.",
        "entities": "찰리 한 명만 등장하며 샌드 베이지 기계 장갑, 큰 어깨와 긴 팔, 흰 마스크형 얼굴, 거대한 금속 손이 참조와 대체로 일치한다. 얼굴에는 점 형태의 구멍들이 두드러진다. 장갑에 긁힘과 마모는 있지만 흉부는 대부분 온전해서 미수리 손상이 분명하지 않다. 참조의 밀짚모자와 다색 목도리는 빠졌다. 하체 의상은 프레임 밖이다. 다른 사람이나 명확히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "큰 손은 손목 관절과 전완에 연결되어 있고, 굽힌 팔꿈치와 상완이 어깨까지 이어져 기계 팔의 지지가 성립한다. 몸통은 비교적 곧게 서 있어 동세는 약하지만 물리적으로 불가능한 자세는 아니다. 발은 화면 밖이며 부유를 나타내는 증거는 없다. 하단 이불은 침대 위에 놓여 있고 벽등과 의료 설비대는 벽에 고정되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "큰 금속 손을 허공으로 치켜든 역동적인 상체와 미수리 흉부 손상을 잘 구현했으나, 참조의 밀짚모자와 화려한 목도리가 빠졌다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "상체 중심 미디엄 숏과 위로 치켜든 손은 맞지만, 몸통의 동세가 약하고 흉부 손상이 표면 흠집 정도이며 참조의 모자와 목도리도 없다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 왼쪽의 큰 금속 손이 머리보다 높이 올라가 있으며 손끝은 천장 쪽 허공을 향한다. 손바닥은 카메라 쪽으로 비스듬히 열려 있다. 얼굴은 들린 팔 쪽으로 조금 돌아가 있지만 눈의 정확한 시선은 어두운 눈구멍 때문에 판별하기 어렵다. 허공으로 손을 치켜든다는 동작 방향은 맞는다.",
        "built_space": "베이지색 벽과 왼쪽 목재 문틀, 콘센트 패널 세 구획이 보이는 벽면 의료 설비대 하나, 벽등 하나가 있다. 왼쪽에는 수액대 하나, 뒤에는 모니터 하나, 오른쪽에는 천장 레일에 달린 커튼이 보인다. 침대의 먼 쪽 프레임과 가까운 쪽 난간이 하단에 걸리고 환자는 보이지 않는다. 찰리는 침대 너머 옆에 서 있는 배치로 읽힌다. 참조의 벽 재질과 의료 설비 분위기는 이어지지만, 수액대·모니터·커튼의 정확한 배치는 참조에서 확인할 수 없다. 거울이나 문제 될 반사는 없다.",
        "entities": "찰리 한 명만 보인다. 육중한 어깨, 긴 기계 팔, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴과 큰 금속 손은 참조의 정체성과 대체로 맞는다. 흉부 한쪽에는 깊게 벌어진 파손 자국이 있어 미수리 상태가 명확하다. 참조의 밀짚모자와 다색 목도리는 화면에 있어야 할 부위에서 누락됐다. 다리와 장화는 상체 프레이밍 밖이므로 평가 대상이 아니다. 다른 사람이나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "들린 손은 손목·전완·팔꿈치·상완을 통해 어깨에 연결되어 있고, 굽힌 팔꿈치와 기울어진 상체가 팔을 들어 올리는 동작을 뒷받침한다. 반대 팔도 몸통에 정상적으로 연결되어 있다. 하체와 발은 프레임 밖이지만 몸이 공중에 떠 있는 징후는 없다. 모니터는 받침 위에 있고 커튼은 천장 레일에 매달려 있어 지지 관계가 자연스럽다."
       },
       {
        "label": "A",
        "direction": "화면 왼쪽의 금속 손이 머리 위로 올라가고 전완과 손끝은 거의 수직으로 천장을 향한다. 열린 손바닥은 카메라를 향한다. 얼굴은 정면에서 약간 화면 왼쪽으로 돌아가 있으며, 망 형태의 눈 부분 때문에 정확한 시선 목표는 확인하기 어렵다. 손을 허공으로 치켜든 방향 자체는 정확하다.",
        "built_space": "베이지색 벽, 중앙 뒤쪽의 목재 문 또는 벽판 한 구획, 오른쪽 의료 설비대 하나와 벽등 하나가 보인다. 설비대에는 콘센트 패널 두 구획과 화면 끝에 잘린 추가 구획이 있다. 천장에는 사각 설비판 하나와 원형 설비 하나가 보인다. 침대 난간과 이불의 작은 가장자리만 하단에 남아 있고 환자는 프레임 밖이다. 찰리는 침대 옆에 선 것으로 읽히며, 참조의 목재·벽·의료 설비 재질과 대체로 이어진다. 불가능한 반사는 없다.",
        "entities": "찰리 한 명만 등장하며 샌드 베이지 기계 장갑, 큰 어깨와 긴 팔, 흰 마스크형 얼굴, 거대한 금속 손이 참조와 대체로 일치한다. 얼굴에는 점 형태의 구멍들이 두드러진다. 장갑에 긁힘과 마모는 있지만 흉부는 대부분 온전해서 미수리 손상이 분명하지 않다. 참조의 밀짚모자와 다색 목도리는 빠졌다. 하체 의상은 프레임 밖이다. 다른 사람이나 명확히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "큰 손은 손목 관절과 전완에 연결되어 있고, 굽힌 팔꿈치와 상완이 어깨까지 이어져 기계 팔의 지지가 성립한다. 몸통은 비교적 곧게 서 있어 동세는 약하지만 물리적으로 불가능한 자세는 아니다. 발은 화면 밖이며 부유를 나타내는 증거는 없다. 하단 이불은 침대 위에 놓여 있고 벽등과 의료 설비대는 벽에 고정되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.286,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.036,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 환자를 프레임 외부로 제외하라는 지시를 어기고 화면 하단에 환자를 포함함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1036
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "침대 모서리만 남긴 구도와 흉부 파손 상태 등 모든 프롬프트 지시사항을 충실히 구현함."
   },
   {
    "label": "A",
    "score": 1036,
    "verdict_ko": "흉부 파손 상태가 누락되었으며, 환자를 프레임에서 제외하라는 명시적 구도 지시를 위반함.  ★위반: [gemini-pro] 환자를 프레임 외부로 제외하라는 지시를 어기고 화면 하단에 환자를 포함함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S76sh6_sel.png",
    "asset_id": "1bd20102-b394-42bb-a202-a7d6b42a76b7",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-c9c3-7f1c-b156-79de28c99c0d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S76sh6"
  }
 },
 "S76sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:39:34.014205+00:00",
  "fingerprint": "031e1faeddd3d49505a769967e02329f1a6c55109d8cf206bcc36a4e475aa5ba",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S76sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S76sh12_sel.png",
  "source_sha256": "760248c0c11748b908a65528f4af5ceeaa87ab9e76b44428f639d0f4b3664301",
  "file": "S76sh12_cine.png",
  "staged_sha256": "42148a0631b2cc5dac604ae58490b15b295d5c5722896fa3630b0f094e3acfa5",
  "latency_ms": 13479
 },
 "S77sh22::signage": {
  "fp": "01b988ad1f408053",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S77sh22": {
  "input_fingerprint": "d2e86e6203459020",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 작동하는 비행기들 한가운데 서서 양팔을 허리에 얹은 채 당당한 포즈를 취한 찰리의 전신.\n\nLOCATION (lock): Inside the abandoned airfield's aircraft-storage hangar, among newly running aircraft and bulbs fluctuating brightly overhead. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating light aircraft (Several previously disabled aircraft have started and are emitting black exhaust) — Different oblique sides of the aircraft flank 찰리 and recede behind him; used as Visible proof of the repair, distributed across the composition with realistic scale; Hangar floor (Visible around 찰리 and the aircraft); used as Ground plane that anchors the full-body pose and aircraft spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Existing nighttime hangar illumination provides controlled separation through the aircrafts' black exhaust, keeping 찰리 readable without carrying the earlier electrical flicker forward as a new effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Several formerly nonfunctional light aircraft now have running engines and emit black exhaust; a bulb has burst during the power surge. Charlie remains physically battered, with his chest opening exposed despite successfully powering the aircraft.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 작동하는 비행기들 한가운데 서서 양팔을 허리에 얹은 채 당당한 포즈를 취한 찰리의 전신.\n\nLOCATION (lock): Inside the abandoned airfield's aircraft-storage hangar, among newly running aircraft and bulbs fluctuating brightly overhead. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating light aircraft (Several previously disabled aircraft have started and are emitting black exhaust) — Different oblique sides of the aircraft flank 찰리 and recede behind him; used as Visible proof of the repair, distributed across the composition with realistic scale; Hangar floor (Visible around 찰리 and the aircraft); used as Ground plane that anchors the full-body pose and aircraft spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Existing nighttime hangar illumination provides controlled separation through the aircrafts' black exhaust, keeping 찰리 readable without carrying the earlier electrical flicker forward as a new effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Several formerly nonfunctional light aircraft now have running engines and emit black exhaust; a bulb has burst during the power surge. Charlie remains physically battered, with his chest opening exposed despite successfully powering the aircraft.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 작동하는 비행기들 한가운데 서서 양팔을 허리에 얹은 채 당당한 포즈를 취한 찰리의 전신.\n\nLOCATION (lock): Inside the abandoned airfield's aircraft-storage hangar, among newly running aircraft and bulbs fluctuating brightly overhead. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating light aircraft (Several previously disabled aircraft have started and are emitting black exhaust) — Different oblique sides of the aircraft flank 찰리 and recede behind him; used as Visible proof of the repair, distributed across the composition with realistic scale; Hangar floor (Visible around 찰리 and the aircraft); used as Ground plane that anchors the full-body pose and aircraft spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Existing nighttime hangar illumination provides controlled separation through the aircrafts' black exhaust, keeping 찰리 readable without carrying the earlier electrical flicker forward as a new effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Several formerly nonfunctional light aircraft now have running engines and emit black exhaust; a bulb has burst during the power surge. Charlie remains physically battered, with his chest opening exposed despite successfully powering the aircraft.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선과 몸의 정면이 카메라를 향하고 있으며, 양옆과 뒤편에 경비행기들이 위치함.",
    "built_space": "이전 샷과 일치하는 격납고 내부로, 켜져 있는 천장 조명과 콘크리트 바닥, 철제 벽면 구조가 식별됨.",
    "entities": "찰리의 로봇 외형과 골격은 일치하지만 레퍼런스의 화려한 의상들은 보이지 않음. 프롬프트대로 좌측 가슴 장갑이 열려 내부 기계가 노출됨. 배경의 비행기들은 매연을 뿜으며 작동 중임.",
    "hard_violations": [],
    "physics": "찰리는 양발로 지면에 안정적으로 서서 허리에 손을 얹고 있으며, 프로펠러의 회전과 매연의 상승이 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "찰리는 카메라를 정면으로 응시하고 있으며, 양옆에 경비행기들이 대칭에 가까운 구도로 배치됨.",
    "built_space": "이전 샷과 유사한 격납고로, 천장에 여러 개의 전등이 파손 없이 밝게 켜져 있음.",
    "entities": "찰리의 로봇 형태는 나타나나 레퍼런스의 의상(모자, 셔츠, 신발)이 없으며, 가슴에는 열린 구조 대신 검은 그을음만 존재함. 경비행기들은 매연을 뿜고 있음.",
    "hard_violations": [],
    "physics": "찰리의 두 발이 바닥에 닿아 중력을 받고 있으며, 양손을 골반에 얹은 자세와 비행기의 배기 가스 방향이 정상적으로 작동함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 의상(셔츠, 모자, 장화)이 누락되었으나, 와이드 샷 프레이밍을 준수하고 지시된 '가슴 장갑이 열린 상태'를 정확히 구현하여 가장 우수함."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "레퍼런스의 주요 의상이 누락되었으며, 가슴 내부가 노출되어야 한다는 구체적인 지시를 어기고 그을음만 표현해 감점됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선과 몸의 정면이 카메라를 향하고 있으며, 양옆과 뒤편에 경비행기들이 위치함.",
        "built_space": "이전 샷과 일치하는 격납고 내부로, 켜져 있는 천장 조명과 콘크리트 바닥, 철제 벽면 구조가 식별됨.",
        "entities": "찰리의 로봇 외형과 골격은 일치하지만 레퍼런스의 화려한 의상들은 보이지 않음. 프롬프트대로 좌측 가슴 장갑이 열려 내부 기계가 노출됨. 배경의 비행기들은 매연을 뿜으며 작동 중임.",
        "hard_violations": [],
        "physics": "찰리는 양발로 지면에 안정적으로 서서 허리에 손을 얹고 있으며, 프로펠러의 회전과 매연의 상승이 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "찰리는 카메라를 정면으로 응시하고 있으며, 양옆에 경비행기들이 대칭에 가까운 구도로 배치됨.",
        "built_space": "이전 샷과 유사한 격납고로, 천장에 여러 개의 전등이 파손 없이 밝게 켜져 있음.",
        "entities": "찰리의 로봇 형태는 나타나나 레퍼런스의 의상(모자, 셔츠, 신발)이 없으며, 가슴에는 열린 구조 대신 검은 그을음만 존재함. 경비행기들은 매연을 뿜고 있음.",
        "hard_violations": [],
        "physics": "찰리의 두 발이 바닥에 닿아 중력을 받고 있으며, 양손을 골반에 얹은 자세와 비행기의 배기 가스 방향이 정상적으로 작동함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 의상(셔츠, 모자, 장화)이 누락되었으나, 와이드 샷 프레이밍을 준수하고 지시된 '가슴 장갑이 열린 상태'를 정확히 구현하여 가장 우수함."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "레퍼런스의 주요 의상이 누락되었으며, 가슴 내부가 노출되어야 한다는 구체적인 지시를 어기고 그을음만 표현해 감점됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선과 몸의 정면이 카메라를 향하고 있으며, 양옆과 뒤편에 경비행기들이 위치함.",
        "built_space": "이전 샷과 일치하는 격납고 내부로, 켜져 있는 천장 조명과 콘크리트 바닥, 철제 벽면 구조가 식별됨.",
        "entities": "찰리의 로봇 외형과 골격은 일치하지만 레퍼런스의 화려한 의상들은 보이지 않음. 프롬프트대로 좌측 가슴 장갑이 열려 내부 기계가 노출됨. 배경의 비행기들은 매연을 뿜으며 작동 중임.",
        "hard_violations": [],
        "physics": "찰리는 양발로 지면에 안정적으로 서서 허리에 손을 얹고 있으며, 프로펠러의 회전과 매연의 상승이 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "찰리는 카메라를 정면으로 응시하고 있으며, 양옆에 경비행기들이 대칭에 가까운 구도로 배치됨.",
        "built_space": "이전 샷과 유사한 격납고로, 천장에 여러 개의 전등이 파손 없이 밝게 켜져 있음.",
        "entities": "찰리의 로봇 형태는 나타나나 레퍼런스의 의상(모자, 셔츠, 신발)이 없으며, 가슴에는 열린 구조 대신 검은 그을음만 존재함. 경비행기들은 매연을 뿜고 있음.",
        "hard_violations": [],
        "physics": "찰리의 두 발이 바닥에 닿아 중력을 받고 있으며, 양손을 골반에 얹은 자세와 비행기의 배기 가스 방향이 정상적으로 작동함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "비행기들 가운데 양손을 허리에 댄 전신 구도는 충실하지만, 가슴 개구부가 닫힌 원형 장치처럼 보이고 참고 의상도 빠져 있다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "당당한 전신 자세와 작동 중인 비행기들의 배치에 더해 열린 가슴과 내부 부품까지 보여, 손상 상태의 연속성을 더 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 눈은 거의 카메라 정면을 향하며, 양팔을 벌려 굽힌 뒤 양손을 허리 앞쪽에 댔다. 좌우 앞 비행기의 기수와 프로펠러 축은 중앙 통로 쪽으로 조금 틀어진 채 카메라 방향을 향한다. 뒤쪽 기체들도 앞쪽을 향하며, 검은 배기는 기체 주변에서 위로 퍼진다. 지시된 별도의 시선 표적은 없다.",
        "built_space": "철골 박공지붕, 골판 금속 벽, 높은 측창, 얼룩진 콘크리트 바닥이 보인다. 천장에는 켜진 등기구 여섯 개가 명확히 보인다. 좌우 전경에 비행기 두 대, 그 뒤와 중앙에 추가로 세 대가 부분적으로 보여 총 다섯 대가 식별된다. 찰리는 기체 사이 중앙 통로에 서 있고 발밑 바닥도 충분히 드러난다. 참고의 낡은 금속 격납고와 재질은 부합하지만, 참고가 좁은 구도여서 전체 등기구 수의 일치 여부는 확정할 수 없다.",
        "entities": "찰리 한 명만 있으며 이전 장면의 남성은 없다. 흰 분절형 마스크 얼굴, 어두운 눈구멍, 샌드 베이지 장갑, 넓은 어깨와 육중한 팔은 참고와 맞는다. 가려진 기계형 외형이라 인간의 나이·성별·민족성은 판별되지 않는다. 참고의 밀짚모자, 화려한 꽃무늬 옷감, 초록 장화는 없고 다리는 장갑으로 덮였다. 가슴에 검게 탄 흔적은 있지만 중앙 원형 부품은 막힌 표면처럼 보여 노출된 개구부 표현이 약하다. 낡은 경비행기, 회전하는 프로펠러와 검은 배기가 있으며, 터진 전구는 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 양발은 바닥에 닿고 벌린 다리가 몸무게를 지탱한다. 양손도 허리 장갑에 접촉하며 팔꿈치 굽힘은 가능한 자세다. 비행기는 보이는 착륙바퀴로 바닥에 지지되고, 프로펠러의 회전 흐림과 상승하는 배기는 엔진 작동에 부합한다. 일부 날개는 원근상 겹치지만 명백한 관통이나 지지 없는 부유는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 몸과 얼굴을 화면 오른쪽으로 약간 돌리고 카메라 오른편을 바라본다. 양손은 허리 양옆에 놓였고 팔꿈치는 바깥으로 벌어져 있다. 왼쪽 비행기는 거의 정면, 오른쪽 두 비행기는 카메라와 중앙 통로 쪽을 향한 서로 다른 사선으로 보인다. 검은 배기는 엔진 부근에서 위로 올라간다. 특정 대상을 바라보라는 지시는 없어 이 시선은 장면과 충돌하지 않는다.",
        "built_space": "철골 지붕과 기둥, 금속 벽판, 양쪽 높은 창, 오염된 콘크리트 바닥이 보인다. 중앙 천장등 하나와 왼쪽·오른쪽·뒤쪽 벽등 각 하나로 켜진 등기구 네 개가 식별된다. 왼쪽 한 대, 오른쪽 두 대, 중앙 깊숙한 곳의 부분적인 한 대까지 비행기 네 대가 보인다. 찰리는 중앙보다 왼쪽의 빈 바닥에 서 있고 기체들이 양옆과 뒤로 배치된다. 참고의 차가운 벽 색과 국소적인 따뜻한 전구 조명이 잘 이어지며, 참고에서 가려진 공간의 정확한 설비 수는 검증할 수 없다.",
        "entities": "찰리 외에 다른 인물은 없다. 흰 마스크형 얼굴과 베이지 장갑, 큰 어깨와 팔은 참고의 정체성과 부합하지만, 다리는 비교적 길어 고릴라형 비율은 약하다. 가려진 외형이므로 인간의 나이·성별·민족성은 확인할 수 없다. 밀짚모자, 꽃무늬 옷감과 초록 장화는 생략됐다. 가슴 중앙의 어두운 원형 구멍과 옆의 파손된 장갑 안쪽 부품이 보여 가슴 개방 상태가 A보다 분명하다. 낡은 경비행기 여러 대와 프로펠러 회전 흔적, 검은 배기가 보인다. 터진 전구는 식별되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 양발은 콘크리트 바닥에 붙어 있고, 벌어진 다리와 약간 돌아간 몸통이 안정적으로 체중을 받는다. 양손은 허리와 접촉한다. 비행기들은 착륙바퀴로 지지되고 왼쪽 앞바퀴에는 고임목도 보인다. 프로펠러는 기수 축에 연결되어 있으며 배기는 위로 확산된다. 날개와 인물의 화면상 겹침은 깊이 차이로 설명 가능하고, 명백히 떠 있거나 지지 없는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "비행기들 가운데 양손을 허리에 댄 전신 구도는 충실하지만, 가슴 개구부가 닫힌 원형 장치처럼 보이고 참고 의상도 빠져 있다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "당당한 전신 자세와 작동 중인 비행기들의 배치에 더해 열린 가슴과 내부 부품까지 보여, 손상 상태의 연속성을 더 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 눈은 거의 카메라 정면을 향하며, 양팔을 벌려 굽힌 뒤 양손을 허리 앞쪽에 댔다. 좌우 앞 비행기의 기수와 프로펠러 축은 중앙 통로 쪽으로 조금 틀어진 채 카메라 방향을 향한다. 뒤쪽 기체들도 앞쪽을 향하며, 검은 배기는 기체 주변에서 위로 퍼진다. 지시된 별도의 시선 표적은 없다.",
        "built_space": "철골 박공지붕, 골판 금속 벽, 높은 측창, 얼룩진 콘크리트 바닥이 보인다. 천장에는 켜진 등기구 여섯 개가 명확히 보인다. 좌우 전경에 비행기 두 대, 그 뒤와 중앙에 추가로 세 대가 부분적으로 보여 총 다섯 대가 식별된다. 찰리는 기체 사이 중앙 통로에 서 있고 발밑 바닥도 충분히 드러난다. 참고의 낡은 금속 격납고와 재질은 부합하지만, 참고가 좁은 구도여서 전체 등기구 수의 일치 여부는 확정할 수 없다.",
        "entities": "찰리 한 명만 있으며 이전 장면의 남성은 없다. 흰 분절형 마스크 얼굴, 어두운 눈구멍, 샌드 베이지 장갑, 넓은 어깨와 육중한 팔은 참고와 맞는다. 가려진 기계형 외형이라 인간의 나이·성별·민족성은 판별되지 않는다. 참고의 밀짚모자, 화려한 꽃무늬 옷감, 초록 장화는 없고 다리는 장갑으로 덮였다. 가슴에 검게 탄 흔적은 있지만 중앙 원형 부품은 막힌 표면처럼 보여 노출된 개구부 표현이 약하다. 낡은 경비행기, 회전하는 프로펠러와 검은 배기가 있으며, 터진 전구는 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 양발은 바닥에 닿고 벌린 다리가 몸무게를 지탱한다. 양손도 허리 장갑에 접촉하며 팔꿈치 굽힘은 가능한 자세다. 비행기는 보이는 착륙바퀴로 바닥에 지지되고, 프로펠러의 회전 흐림과 상승하는 배기는 엔진 작동에 부합한다. 일부 날개는 원근상 겹치지만 명백한 관통이나 지지 없는 부유는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 몸과 얼굴을 화면 오른쪽으로 약간 돌리고 카메라 오른편을 바라본다. 양손은 허리 양옆에 놓였고 팔꿈치는 바깥으로 벌어져 있다. 왼쪽 비행기는 거의 정면, 오른쪽 두 비행기는 카메라와 중앙 통로 쪽을 향한 서로 다른 사선으로 보인다. 검은 배기는 엔진 부근에서 위로 올라간다. 특정 대상을 바라보라는 지시는 없어 이 시선은 장면과 충돌하지 않는다.",
        "built_space": "철골 지붕과 기둥, 금속 벽판, 양쪽 높은 창, 오염된 콘크리트 바닥이 보인다. 중앙 천장등 하나와 왼쪽·오른쪽·뒤쪽 벽등 각 하나로 켜진 등기구 네 개가 식별된다. 왼쪽 한 대, 오른쪽 두 대, 중앙 깊숙한 곳의 부분적인 한 대까지 비행기 네 대가 보인다. 찰리는 중앙보다 왼쪽의 빈 바닥에 서 있고 기체들이 양옆과 뒤로 배치된다. 참고의 차가운 벽 색과 국소적인 따뜻한 전구 조명이 잘 이어지며, 참고에서 가려진 공간의 정확한 설비 수는 검증할 수 없다.",
        "entities": "찰리 외에 다른 인물은 없다. 흰 마스크형 얼굴과 베이지 장갑, 큰 어깨와 팔은 참고의 정체성과 부합하지만, 다리는 비교적 길어 고릴라형 비율은 약하다. 가려진 외형이므로 인간의 나이·성별·민족성은 확인할 수 없다. 밀짚모자, 꽃무늬 옷감과 초록 장화는 생략됐다. 가슴 중앙의 어두운 원형 구멍과 옆의 파손된 장갑 안쪽 부품이 보여 가슴 개방 상태가 A보다 분명하다. 낡은 경비행기 여러 대와 프로펠러 회전 흔적, 검은 배기가 보인다. 터진 전구는 식별되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 양발은 콘크리트 바닥에 붙어 있고, 벌어진 다리와 약간 돌아간 몸통이 안정적으로 체중을 받는다. 양손은 허리와 접촉한다. 비행기들은 착륙바퀴로 지지되고 왼쪽 앞바퀴에는 고임목도 보인다. 프로펠러는 기수 축에 연결되어 있으며 배기는 위로 확산된다. 날개와 인물의 화면상 겹침은 깊이 차이로 설명 가능하고, 명백히 떠 있거나 지지 없는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.589
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.589
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1589
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스의 의상(셔츠, 모자, 장화)이 누락되었으나, 와이드 샷 프레이밍을 준수하고 지시된 '가슴 장갑이 열린 상태'를 정확히 구현하여 가장 우수함."
   },
   {
    "label": "B",
    "score": 1589,
    "verdict_ko": "레퍼런스의 주요 의상이 누락되었으며, 가슴 내부가 노출되어야 한다는 구체적인 지시를 어기고 그을음만 표현해 감점됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5_sel.png",
    "asset_id": "dc271dc3-8964-47de-87f3-4c3068e964c2",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-cb6d-717c-8725-c82f7cfcf3b0",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S75sh5"
  }
 },
 "S77sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:40:39.638969+00:00",
  "fingerprint": "01560577659b0d1e3bea3306d4b7b2851b28d9f11da7b204f1a642d1619d87b2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S77sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S77sh22_sel.png",
  "source_sha256": "af49892f100545ebcd58f29afb68e717b13c8d0e5fae9a111c8ab926553c24b3",
  "file": "S77sh22_cine.png",
  "staged_sha256": "26e9c91c6405c23b9249faebd4a0455514809bc6bf9a32f1dacba7f28d3a601f",
  "latency_ms": 10260
 },
 "S77sh43::signage": {
  "fp": "735e290f8cbaf775",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::fe03d1436e8ce9a0": {
  "subjects": [],
  "subject_text": "목포공항 격납고\n거대한 창고 문과 넓은 바닥을 갖춘 낡은 격납고. 높고 깊은 실내 위로 전구 조명이 설치되어 있다.",
  "identity": "canonical",
  "scope_id": "L257",
  "scope_role": "location_interior",
  "scope_sha": "a9d7f00b16b22d2b"
 },
 "S77sh43::bgfirst_bg": {
  "input_fingerprint": "0cf81d7fb6a2acd1",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43__bgfirst_bg.png",
  "asset_id": "aadf8628-2de5-41d1-ab0c-9466e0ee03d4",
  "input_asset_ids": [
   "5d426069-57f7-4d5a-b31e-f5848a2bcb37",
   "6a2d2684-afeb-4aa6-b80b-232ee22c7f48"
  ]
 },
 "S77sh43": {
  "input_fingerprint": "0b4683f4cf911378",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is now daylight, and the helicopter is airborne and receding from the airport. Charlie remains on the ground with his unrepaired body damage and exposed chest opening. 현우: He remains on the ground at the airport, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is now daylight, and the helicopter is airborne and receding from the airport. Charlie remains on the ground with his unrepaired body damage and exposed chest opening. 현우: He remains on the ground at the airport, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is now daylight, and the helicopter is airborne and receding from the airport. Charlie remains on the ground with his unrepaired body damage and exposed chest opening. 현우: He remains on the ground at the airport, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43__bgfirst_bg.png",
     "asset_id": "aadf8628-2de5-41d1-ab0c-9466e0ee03d4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S77sh43.png",
     "asset_id": "5d426069-57f7-4d5a-b31e-f5848a2bcb37",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L257B01.png",
     "asset_id": "6a2d2684-afeb-4aa6-b80b-232ee22c7f48",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1096323>",
     "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 우측 상단의 멀어지는 헬리콥터를 향하고 있으며, 찰리가 헬리콥터를 향해 오른팔을 들고 있음.",
    "built_space": "프레임 좌측에 격납고가 있고, 바닥은 넓은 활주로로 구성됨.",
    "entities": "현우는 레퍼런스와 달리 반팔 티셔츠를 입고 있음. 찰리는 모자와 셔츠가 없고 카툰식 2D 외곽선 렌더링으로 표현되어 실사감을 상실함. 우측 상단에 헬리콥터가 있음.",
    "hard_violations": [],
    "physics": "인물들은 지면에 두 발로 서 있으며, 찰리의 팔은 정상적으로 몸에 붙어 지지된 채로 들려 있음."
   },
   {
    "label": "B",
    "direction": "찰리와 현우 모두 우측 상단의 헬리콥터를 향해 서 있으며, 찰리의 오른팔이 헬리콥터 쪽으로 뻗어 있음.",
    "built_space": "좌측에 격납고 건물이 위치하고, 하단에는 두 인물이 서 있는 콘크리트 활주로가 펼쳐져 있음.",
    "entities": "현우는 뒷모습으로 레퍼런스와 일치하는 회색 긴팔 셔츠와 바지를 입고 있음. 찰리는 밀짚모자와 화려한 셔츠가 누락되었으나 베이지색 금속 장갑과 등 부분의 파손이 보임. 우측 상단에 헬리콥터가 존재함.",
    "hard_violations": [],
    "physics": "두 인물 모두 바닥에 안정적으로 딛고 서 있으며, 찰리의 들린 팔은 몸체 관절에 자연스럽게 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 실사 렌더링, 야간 조명, 현우의 의상을 잘 구현했으나 찰리의 옷차림(모자와 셔츠)이 누락됨."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캐릭터가 2D 카툰 렌더링으로 묘사되어 실사 요건을 크게 위반했으며, 현우의 의상과 찰리의 소품이 일치하지 않음."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리와 현우 모두 우측 상단의 헬리콥터를 향해 서 있으며, 찰리의 오른팔이 헬리콥터 쪽으로 뻗어 있음.",
        "built_space": "좌측에 격납고 건물이 위치하고, 하단에는 두 인물이 서 있는 콘크리트 활주로가 펼쳐져 있음.",
        "entities": "현우는 뒷모습으로 레퍼런스와 일치하는 회색 긴팔 셔츠와 바지를 입고 있음. 찰리는 밀짚모자와 화려한 셔츠가 누락되었으나 베이지색 금속 장갑과 등 부분의 파손이 보임. 우측 상단에 헬리콥터가 존재함.",
        "hard_violations": [],
        "physics": "두 인물 모두 바닥에 안정적으로 딛고 서 있으며, 찰리의 들린 팔은 몸체 관절에 자연스럽게 지지되어 있음."
       },
       {
        "label": "A",
        "direction": "두 인물이 우측 상단의 멀어지는 헬리콥터를 향하고 있으며, 찰리가 헬리콥터를 향해 오른팔을 들고 있음.",
        "built_space": "프레임 좌측에 격납고가 있고, 바닥은 넓은 활주로로 구성됨.",
        "entities": "현우는 레퍼런스와 달리 반팔 티셔츠를 입고 있음. 찰리는 모자와 셔츠가 없고 카툰식 2D 외곽선 렌더링으로 표현되어 실사감을 상실함. 우측 상단에 헬리콥터가 있음.",
        "hard_violations": [],
        "physics": "인물들은 지면에 두 발로 서 있으며, 찰리의 팔은 정상적으로 몸에 붙어 지지된 채로 들려 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 실사 렌더링, 야간 조명, 현우의 의상을 잘 구현했으나 찰리의 옷차림(모자와 셔츠)이 누락됨."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캐릭터가 2D 카툰 렌더링으로 묘사되어 실사 요건을 크게 위반했으며, 현우의 의상과 찰리의 소품이 일치하지 않음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리와 현우 모두 우측 상단의 헬리콥터를 향해 서 있으며, 찰리의 오른팔이 헬리콥터 쪽으로 뻗어 있음.",
        "built_space": "좌측에 격납고 건물이 위치하고, 하단에는 두 인물이 서 있는 콘크리트 활주로가 펼쳐져 있음.",
        "entities": "현우는 뒷모습으로 레퍼런스와 일치하는 회색 긴팔 셔츠와 바지를 입고 있음. 찰리는 밀짚모자와 화려한 셔츠가 누락되었으나 베이지색 금속 장갑과 등 부분의 파손이 보임. 우측 상단에 헬리콥터가 존재함.",
        "hard_violations": [],
        "physics": "두 인물 모두 바닥에 안정적으로 딛고 서 있으며, 찰리의 들린 팔은 몸체 관절에 자연스럽게 지지되어 있음."
       },
       {
        "label": "A",
        "direction": "두 인물이 우측 상단의 멀어지는 헬리콥터를 향하고 있으며, 찰리가 헬리콥터를 향해 오른팔을 들고 있음.",
        "built_space": "프레임 좌측에 격납고가 있고, 바닥은 넓은 활주로로 구성됨.",
        "entities": "현우는 레퍼런스와 달리 반팔 티셔츠를 입고 있음. 찰리는 모자와 셔츠가 없고 카툰식 2D 외곽선 렌더링으로 표현되어 실사감을 상실함. 우측 상단에 헬리콥터가 있음.",
        "hard_violations": [],
        "physics": "인물들은 지면에 두 발로 서 있으며, 찰리의 팔은 정상적으로 몸에 붙어 지지된 채로 들려 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "야간 와이드숏에서 두 인물의 뒷모습과 우측 상단 헬리콥터를 향한 작별 동작을 구현하며, 현우의 긴소매 셔츠까지 참조에 맞아 B보다 충실하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지상에 남은 두 인물과 멀어지는 헬리콥터의 배치는 충실하지만, 현우의 긴소매 셔츠를 반소매 티셔츠로 바꿨고 찰리의 참조 의상도 빠졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개를 오른쪽으로 돌려 우측 상단의 헬리콥터 쪽을 바라본다. 찰리도 같은 하늘 방향으로 몸을 두고 오른팔을 높이 들어 손바닥을 펼친다. 손짓의 대상은 헬리콥터로 읽힌다. 헬리콥터는 꼬리가 왼쪽, 기수가 오른쪽으로 놓이고 아래쪽과 후측면이 보여 오른쪽 먼 공간으로 떠나는 배치에 부합한다.",
        "built_space": "왼쪽에 골판 금속 외벽과 상부 창열을 가진 격납고 한 동, 열린 대형 출입구 한 곳이 보인다. 출입구 안쪽 천장등 한 개, 외벽 상단등 한 개와 모서리의 투광등 두 개가 식별된다. 두 인물은 문 밖 콘크리트 계류장에 나란히 서 있으며 건물이나 설비와 겹치지 않는다. 균열 난 바닥, 희미한 노란 유도선, 먼 산 능선은 장소 참조와 부합한다. 밤으로 바꾼 조명은 명시된 야간 잠금을 따른다.",
        "entities": "현우 한 명, 찰리 한 개체, 헬리콥터 한 대가 보이며 추가 인물이나 읽을 수 있는 글자는 없다. 현우의 검은 헝클어진 머리, 마른 청년 체형, 회색 긴소매 셔츠와 녹색 카고바지, 어두운 신발은 참조와 가깝다. 얼굴 대부분이 가려져 정확한 나이와 얼굴 일치는 확인할 수 없다. 찰리는 베이지 장갑판, 육중한 금속 팔, 상대적으로 짧은 다리, 어두운 바지와 녹색 장화를 갖췄지만 참조의 밀짚모자와 화려한 천 의상이 없다. 등에 원형 파손 구멍이 보인다. 가슴과 흰 마스크형 얼굴은 뒤쪽 구도라 확인할 수 없다.",
        "hard_violations": [],
        "physics": "두 인물 모두 양발이 계류장 바닥에 닿고 접지 그림자가 이어진다. 찰리의 들어 올린 팔은 어깨와 팔꿈치 관절에 연결되어 있으며 무거운 팔을 들어 작별 인사하는 자세로 성립한다. 다만 정지된 손들기와 구별할 만큼 강한 흔들림은 드러나지 않는다. 헬리콥터는 회전 흐림이 보이는 주회전날개로 비행 중이며, 지지 없이 떠 있는 별도 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우와 찰리는 모두 등을 보인 채 우측 상단의 헬리콥터가 있는 열린 하늘을 향한다. 찰리는 오른팔을 오른쪽 위로 뻗고 손을 펼쳐 헬리콥터를 향한 작별 인사를 한다. 헬리콥터의 꼬리는 왼쪽, 기수는 오른쪽이며 하부와 후측면이 보여 멀어지는 방향과 맞는다.",
        "built_space": "왼쪽에 금속 외벽 격납고 한 동과 열린 대형 출입구 한 곳, 상부 창열이 있다. 출입구 안쪽 조명 한 개, 외벽 중간의 작은 조명 한 개와 모서리 투광등 두 개가 보인다. 찰리는 격납고 가까운 바깥쪽, 현우는 그 오른쪽의 열린 계류장에 서 있다. 노란 곡선 유도선과 먼 산 능선이 참조 장소를 이어 준다. 바닥은 참조보다 넓게 젖어 있으나 조명과 인물의 반사가 아래쪽으로 이어지는 것은 가능한 배치다. 야간 잠금도 충족한다.",
        "entities": "현우와 찰리, 헬리콥터가 각각 하나씩 보이고 추가 인물이나 읽을 수 있는 글자는 없다. 현우는 헝클어진 검은 머리와 가는 청년 체형, 녹색 카고바지는 맞지만 참조의 회색 긴소매 단추 셔츠 대신 반소매 티셔츠를 입었다. 뒷모습이므로 얼굴과 정확한 연령은 확인할 수 없다. 찰리는 베이지 금속 장갑과 긴 팔, 짧은 다리, 녹색 장화를 갖췄으며 등에 그을음과 노출된 기계부가 있다. 참조의 밀짚모자와 화려한 천 의상은 없고 바지는 더 푸르게 보인다. 가슴 손상과 얼굴은 구도상 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리는 두 장화를 벌려 바닥에 디디고, 현우도 두 발로 지면을 지탱한다. 찰리의 오른팔은 연결된 어깨·팔꿈치 관절을 통해 올라가 있으며 몸통의 기울기와 벌린 다리가 동작을 지지한다. 힘찬 인사로 읽힐 수 있지만 손의 이동 자체는 정지 화면에서 확정할 수 없다. 헬리콥터는 회전날개의 흐림이 보여 양력으로 떠 있는 상태가 성립한다. 바닥 반사도 인물의 접지와 모순되지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "야간 와이드숏에서 두 인물의 뒷모습과 우측 상단 헬리콥터를 향한 작별 동작을 구현하며, 현우의 긴소매 셔츠까지 참조에 맞아 B보다 충실하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지상에 남은 두 인물과 멀어지는 헬리콥터의 배치는 충실하지만, 현우의 긴소매 셔츠를 반소매 티셔츠로 바꿨고 찰리의 참조 의상도 빠졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개를 오른쪽으로 돌려 우측 상단의 헬리콥터 쪽을 바라본다. 찰리도 같은 하늘 방향으로 몸을 두고 오른팔을 높이 들어 손바닥을 펼친다. 손짓의 대상은 헬리콥터로 읽힌다. 헬리콥터는 꼬리가 왼쪽, 기수가 오른쪽으로 놓이고 아래쪽과 후측면이 보여 오른쪽 먼 공간으로 떠나는 배치에 부합한다.",
        "built_space": "왼쪽에 골판 금속 외벽과 상부 창열을 가진 격납고 한 동, 열린 대형 출입구 한 곳이 보인다. 출입구 안쪽 천장등 한 개, 외벽 상단등 한 개와 모서리의 투광등 두 개가 식별된다. 두 인물은 문 밖 콘크리트 계류장에 나란히 서 있으며 건물이나 설비와 겹치지 않는다. 균열 난 바닥, 희미한 노란 유도선, 먼 산 능선은 장소 참조와 부합한다. 밤으로 바꾼 조명은 명시된 야간 잠금을 따른다.",
        "entities": "현우 한 명, 찰리 한 개체, 헬리콥터 한 대가 보이며 추가 인물이나 읽을 수 있는 글자는 없다. 현우의 검은 헝클어진 머리, 마른 청년 체형, 회색 긴소매 셔츠와 녹색 카고바지, 어두운 신발은 참조와 가깝다. 얼굴 대부분이 가려져 정확한 나이와 얼굴 일치는 확인할 수 없다. 찰리는 베이지 장갑판, 육중한 금속 팔, 상대적으로 짧은 다리, 어두운 바지와 녹색 장화를 갖췄지만 참조의 밀짚모자와 화려한 천 의상이 없다. 등에 원형 파손 구멍이 보인다. 가슴과 흰 마스크형 얼굴은 뒤쪽 구도라 확인할 수 없다.",
        "hard_violations": [],
        "physics": "두 인물 모두 양발이 계류장 바닥에 닿고 접지 그림자가 이어진다. 찰리의 들어 올린 팔은 어깨와 팔꿈치 관절에 연결되어 있으며 무거운 팔을 들어 작별 인사하는 자세로 성립한다. 다만 정지된 손들기와 구별할 만큼 강한 흔들림은 드러나지 않는다. 헬리콥터는 회전 흐림이 보이는 주회전날개로 비행 중이며, 지지 없이 떠 있는 별도 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우와 찰리는 모두 등을 보인 채 우측 상단의 헬리콥터가 있는 열린 하늘을 향한다. 찰리는 오른팔을 오른쪽 위로 뻗고 손을 펼쳐 헬리콥터를 향한 작별 인사를 한다. 헬리콥터의 꼬리는 왼쪽, 기수는 오른쪽이며 하부와 후측면이 보여 멀어지는 방향과 맞는다.",
        "built_space": "왼쪽에 금속 외벽 격납고 한 동과 열린 대형 출입구 한 곳, 상부 창열이 있다. 출입구 안쪽 조명 한 개, 외벽 중간의 작은 조명 한 개와 모서리 투광등 두 개가 보인다. 찰리는 격납고 가까운 바깥쪽, 현우는 그 오른쪽의 열린 계류장에 서 있다. 노란 곡선 유도선과 먼 산 능선이 참조 장소를 이어 준다. 바닥은 참조보다 넓게 젖어 있으나 조명과 인물의 반사가 아래쪽으로 이어지는 것은 가능한 배치다. 야간 잠금도 충족한다.",
        "entities": "현우와 찰리, 헬리콥터가 각각 하나씩 보이고 추가 인물이나 읽을 수 있는 글자는 없다. 현우는 헝클어진 검은 머리와 가는 청년 체형, 녹색 카고바지는 맞지만 참조의 회색 긴소매 단추 셔츠 대신 반소매 티셔츠를 입었다. 뒷모습이므로 얼굴과 정확한 연령은 확인할 수 없다. 찰리는 베이지 금속 장갑과 긴 팔, 짧은 다리, 녹색 장화를 갖췄으며 등에 그을음과 노출된 기계부가 있다. 참조의 밀짚모자와 화려한 천 의상은 없고 바지는 더 푸르게 보인다. 가슴 손상과 얼굴은 구도상 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리는 두 장화를 벌려 바닥에 디디고, 현우도 두 발로 지면을 지탱한다. 찰리의 오른팔은 연결된 어깨·팔꿈치 관절을 통해 올라가 있으며 몸통의 기울기와 벌린 다리가 동작을 지지한다. 힘찬 인사로 읽힐 수 있지만 손의 이동 자체는 정지 화면에서 확정할 수 없다. 헬리콥터는 회전날개의 흐림이 보여 양력으로 떠 있는 상태가 성립한다. 바닥 반사도 인물의 접지와 모순되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.304,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.304,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1304
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 실사 렌더링, 야간 조명, 현우의 의상을 잘 구현했으나 찰리의 옷차림(모자와 셔츠)이 누락됨."
   },
   {
    "label": "A",
    "score": 1304,
    "verdict_ko": "캐릭터가 2D 카툰 렌더링으로 묘사되어 실사 요건을 크게 위반했으며, 현우의 의상과 찰리의 소품이 일치하지 않음."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L257B01.png",
    "asset_id": "6a2d2684-afeb-4aa6-b80b-232ee22c7f48",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-cd13-7c14-b426-52249a9f261c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43__bgfirst_bg.png",
   "bg_asset_id": "aadf8628-2de5-41d1-ab0c-9466e0ee03d4",
   "bg_record_key": "S77sh43::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S77sh43::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:41:51.247135+00:00",
  "fingerprint": "e9d4c57c473552812ebf459a1d248df0fddf293b600d986e6a71e8cfd4ce81bf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S77sh43_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S77sh43_sel.png",
  "source_sha256": "1bd75b0d425132790421c517e9e4fd97b158b12296d0c634bbb96e20d9b1deaf",
  "file": "S77sh43_cine.png",
  "staged_sha256": "c124e55d72a5ac91204f7be0ea763a197148230809954676f46eca949050599d",
  "latency_ms": 8287
 },
 "S77sh53::signage": {
  "fp": "b019ba9531380923",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S77sh53": {
  "input_fingerprint": "508c8bc710bf43be",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 다시 앞을 향해 씩씩하게 걷는 도중 다리를 뻗어 딛은 mid-action 상태의 현우와 그 뒤를 따르며 한쪽 다리가 들린 mid-action 순간의 찰리가 멀어지는 뒷모습.\n\nLOCATION (lock): On the open ground leading away from the abandoned airfield's helicopter departure area in daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route continuing away from the camera in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Airport departure route (현우 and 찰리 are walking away along it) — The visible route recedes from the lower foreground toward the upper-center distance; used as Depth anchor for their increasing separation from the stationary camera; Backpacks (Worn by both departing figures) — Their outward-facing backs are visible against the figures' rear silhouettes; used as Small narrative details supporting the shared onward journey.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established restrained daylight and even tonal continuity, allowing the farewell's tenderness to persist without a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The helicopter has disappeared into the distance. Charlie is now wearing a backpack over his still-damaged body, with no repair of the chest opening established. 현우: He walks away from the airport wearing a backpack, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 다시 앞을 향해 씩씩하게 걷는 도중 다리를 뻗어 딛은 mid-action 상태의 현우와 그 뒤를 따르며 한쪽 다리가 들린 mid-action 순간의 찰리가 멀어지는 뒷모습.\n\nLOCATION (lock): On the open ground leading away from the abandoned airfield's helicopter departure area in daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route continuing away from the camera in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Airport departure route (현우 and 찰리 are walking away along it) — The visible route recedes from the lower foreground toward the upper-center distance; used as Depth anchor for their increasing separation from the stationary camera; Backpacks (Worn by both departing figures) — Their outward-facing backs are visible against the figures' rear silhouettes; used as Small narrative details supporting the shared onward journey.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established restrained daylight and even tonal continuity, allowing the farewell's tenderness to persist without a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The helicopter has disappeared into the distance. Charlie is now wearing a backpack over his still-damaged body, with no repair of the chest opening established. 현우: He walks away from the airport wearing a backpack, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 다시 앞을 향해 씩씩하게 걷는 도중 다리를 뻗어 딛은 mid-action 상태의 현우와 그 뒤를 따르며 한쪽 다리가 들린 mid-action 순간의 찰리가 멀어지는 뒷모습.\n\nLOCATION (lock): On the open ground leading away from the abandoned airfield's helicopter departure area in daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route continuing away from the camera in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Airport departure route (현우 and 찰리 are walking away along it) — The visible route recedes from the lower foreground toward the upper-center distance; used as Depth anchor for their increasing separation from the stationary camera; Backpacks (Worn by both departing figures) — Their outward-facing backs are visible against the figures' rear silhouettes; used as Small narrative details supporting the shared onward journey.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established restrained daylight and even tonal continuity, allowing the farewell's tenderness to persist without a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The helicopter has disappeared into the distance. Charlie is now wearing a backpack over his still-damaged body, with no repair of the chest opening established. 현우: He walks away from the airport wearing a backpack, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 찰리 모두 카메라를 등지고 활주로를 따라 멀어지는 방향을 향하고 있음.",
    "built_space": "이전 샷과 동일하게 왼쪽에 격납고가 위치해 있으며, 야간의 콘크리트 활주로 배경이 올바르게 배치됨.",
    "entities": "현우와 찰리 모두 지시된 배낭을 메고 뒷모습을 보이고 있으나, 화면 하단 중앙에 촬영용 카메라 장비가 잘못 렌더링되어 노출됨.",
    "hard_violations": [
     "[gemini-pro] 화면 하단 중앙에 프롬프트에 없는 촬영용 카메라 장비가 노출됨 (leaked marker/object)",
     "[gpt-high] 화면 하단 중앙에 프롬프트에 없는 모니터와 손잡이 형태의 촬영 장비가 노출되어, 장면 밖 장비가 최종 스틸에 유입되었다."
    ],
    "physics": "찰리는 오른발을 들고 걷는 자세를 취하고 있으나, 현우는 양발이 바닥에 닿아 있어 걷는 동작의 역동성이 떨어짐."
   },
   {
    "label": "B",
    "direction": "현우와 찰리 모두 카메라를 등지고 활주로 방향으로 멀어지며 걷고 있음.",
    "built_space": "왼쪽의 격납고와 바닥의 활주로 형태, 야간 조명 등 이전 샷의 구조와 환경이 일관되게 유지됨.",
    "entities": "현우는 배낭을 메고 있으나, 찰리는 프롬프트의 지시(양쪽 모두 배낭 착용)와 달리 배낭을 메지 않고 등 부분의 파손된 구멍이 그대로 노출됨.",
    "hard_violations": [],
    "physics": "찰리와 현우 모두 한쪽 다리를 들고 딛으려는 mid-action의 걷기 자세를 자연스럽게 보여주며, 지탱하는 발의 접지가 안정적임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 배낭 착용 지시가 누락되었으나, 두 인물의 걷는 동작이 지시된 대로 자연스럽게 구현되었고 치명적인 오류가 없습니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "두 인물 모두 배낭을 착용했지만, 화면 하단에 촬영용 카메라 장비가 그대로 노출되는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 찰리 모두 카메라를 등지고 활주로를 따라 멀어지는 방향을 향하고 있음.",
        "built_space": "이전 샷과 동일하게 왼쪽에 격납고가 위치해 있으며, 야간의 콘크리트 활주로 배경이 올바르게 배치됨.",
        "entities": "현우와 찰리 모두 지시된 배낭을 메고 뒷모습을 보이고 있으나, 화면 하단 중앙에 촬영용 카메라 장비가 잘못 렌더링되어 노출됨.",
        "hard_violations": [
         "화면 하단 중앙에 프롬프트에 없는 촬영용 카메라 장비가 노출됨 (leaked marker/object)"
        ],
        "physics": "찰리는 오른발을 들고 걷는 자세를 취하고 있으나, 현우는 양발이 바닥에 닿아 있어 걷는 동작의 역동성이 떨어짐."
       },
       {
        "label": "B",
        "direction": "현우와 찰리 모두 카메라를 등지고 활주로 방향으로 멀어지며 걷고 있음.",
        "built_space": "왼쪽의 격납고와 바닥의 활주로 형태, 야간 조명 등 이전 샷의 구조와 환경이 일관되게 유지됨.",
        "entities": "현우는 배낭을 메고 있으나, 찰리는 프롬프트의 지시(양쪽 모두 배낭 착용)와 달리 배낭을 메지 않고 등 부분의 파손된 구멍이 그대로 노출됨.",
        "hard_violations": [],
        "physics": "찰리와 현우 모두 한쪽 다리를 들고 딛으려는 mid-action의 걷기 자세를 자연스럽게 보여주며, 지탱하는 발의 접지가 안정적임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 배낭 착용 지시가 누락되었으나, 두 인물의 걷는 동작이 지시된 대로 자연스럽게 구현되었고 치명적인 오류가 없습니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "두 인물 모두 배낭을 착용했지만, 화면 하단에 촬영용 카메라 장비가 그대로 노출되는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 찰리 모두 카메라를 등지고 활주로를 따라 멀어지는 방향을 향하고 있음.",
        "built_space": "이전 샷과 동일하게 왼쪽에 격납고가 위치해 있으며, 야간의 콘크리트 활주로 배경이 올바르게 배치됨.",
        "entities": "현우와 찰리 모두 지시된 배낭을 메고 뒷모습을 보이고 있으나, 화면 하단 중앙에 촬영용 카메라 장비가 잘못 렌더링되어 노출됨.",
        "hard_violations": [
         "화면 하단 중앙에 프롬프트에 없는 촬영용 카메라 장비가 노출됨 (leaked marker/object)"
        ],
        "physics": "찰리는 오른발을 들고 걷는 자세를 취하고 있으나, 현우는 양발이 바닥에 닿아 있어 걷는 동작의 역동성이 떨어짐."
       },
       {
        "label": "B",
        "direction": "현우와 찰리 모두 카메라를 등지고 활주로 방향으로 멀어지며 걷고 있음.",
        "built_space": "왼쪽의 격납고와 바닥의 활주로 형태, 야간 조명 등 이전 샷의 구조와 환경이 일관되게 유지됨.",
        "entities": "현우는 배낭을 메고 있으나, 찰리는 프롬프트의 지시(양쪽 모두 배낭 착용)와 달리 배낭을 메지 않고 등 부분의 파손된 구멍이 그대로 노출됨.",
        "hard_violations": [],
        "physics": "찰리와 현우 모두 한쪽 다리를 들고 딛으려는 mid-action의 걷기 자세를 자연스럽게 보여주며, 지탱하는 발의 접지가 안정적임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "밤의 비행장을 떠나는 뒷모습과 접지된 보행은 맞지만, 인물이 다소 크게 잡혔고 찰리의 배낭이 누락되었으며 뒤따르는 간격이 약하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "넓은 출발 경로와 두 배낭은 잘 보이지만, 화면 하단에 요구하지 않은 촬영 장비가 노출되어 프레임을 사용할 수 없다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 찰리 모두 머리와 몸을 카메라 반대편 비행장 원경으로 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없다. 두 사람의 보행 방향은 멀어지는 동작에 맞지만, 찰리는 현우의 뒤보다는 오른쪽에 나란히 선 인상이 강하다.",
        "built_space": "왼쪽에 낡은 금속 격납고 한 동의 외벽과 상부 창열이 있고, 외벽 투광등 두 개가 보인다. 격납고 내부는 화면 밖이다. 갈라진 콘크리트 바닥과 노란 곡선 표시, 먼 산과 공항 불빛이 이전 장면과 연결된다. 두 사람은 구조물 밖의 열린 지면에 있으며, 이동 공간은 전경에서 중앙 원경으로 이어지지만 인물이 화면 높이의 상당 부분을 차지한다.",
        "entities": "등장 대상은 현우와 찰리 둘뿐이며 헬기는 없다. 현우는 검은 헝클어진 머리, 마른 체격, 회색 셔츠, 올리브색 바지와 낡은 신발을 유지하고 배낭 하나를 멘다. 얼굴이 가려져 정확한 나이와 민족적 외모는 확인할 수 없다. 찰리는 베이지색 각진 장갑, 긴 팔, 짙은 바지와 초록 장화를 유지하지만 배낭이 없다. 등 쪽 원형 개구부는 보이나 가슴 손상 상태는 이 시점에서 판정할 수 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우는 오른발로 지면을 딛고 왼발을 뒤로 들어 밑창이 보이며, 찰리는 왼발을 딛고 오른발을 든다. 각각 지지발이 있어 보행 중 체중 이동으로 가능한 자세다. 현우의 배낭은 어깨끈으로 지지되며, 지지 없이 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "두 사람은 카메라에 등을 보이고 중앙 원경의 열린 비행장으로 이동한다. 얼굴과 눈은 보이지 않는다. 찰리는 현우의 왼쪽에 거의 나란히 있어, 현우 뒤를 따르는 관계가 뚜렷하지 않다.",
        "built_space": "왼쪽 격납고 한 동에 열린 출입구 하나, 내부 철골과 천장등 하나, 외벽 투광등 두 개가 보인다. 낡은 금속 벽과 넓은 균열 콘크리트 바닥, 노란 바닥선 및 원경 불빛이 장소 참조와 부합한다. 두 사람을 작게 배치해 전경부터 중앙 원경까지 출발 경로가 넓게 보인다. 다만 화면 맨 아래 중앙에 모니터와 손잡이가 달린 촬영 장비 일부가 끼어 있다.",
        "entities": "현우와 찰리 두 대상이 있으며 둘 다 배낭을 착용한다. 현우의 검은 머리, 회색 셔츠와 올리브색 바지, 찰리의 베이지 장갑과 긴 팔, 짙은 바지와 초록 장화가 참조에 대응한다. 뒷모습이므로 현우의 얼굴 정체성과 찰리의 마스크 얼굴 및 가슴 손상은 확인할 수 없다. 헬기나 추가 인물은 없으나, 요구되지 않은 촬영 장비가 전경에 추가되었다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "화면 하단 중앙에 프롬프트에 없는 모니터와 손잡이 형태의 촬영 장비가 노출되어, 장면 밖 장비가 최종 스틸에 유입되었다."
        ],
        "physics": "찰리는 오른발을 지면에 두고 왼발을 들어 올렸으며, 현우도 오른발을 딛고 왼발을 뒤로 움직이는 보행 자세다. 두 몸 모두 지지발이 있고 배낭은 어깨끈에 걸려 있어 부유하지 않는다. 하단 장비는 프레임 경계에 잘려 지지부가 보이지 않지만, 공중에 떠 있다고 판단할 근거는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "밤의 비행장을 떠나는 뒷모습과 접지된 보행은 맞지만, 인물이 다소 크게 잡혔고 찰리의 배낭이 누락되었으며 뒤따르는 간격이 약하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "넓은 출발 경로와 두 배낭은 잘 보이지만, 화면 하단에 요구하지 않은 촬영 장비가 노출되어 프레임을 사용할 수 없다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 찰리 모두 머리와 몸을 카메라 반대편 비행장 원경으로 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없다. 두 사람의 보행 방향은 멀어지는 동작에 맞지만, 찰리는 현우의 뒤보다는 오른쪽에 나란히 선 인상이 강하다.",
        "built_space": "왼쪽에 낡은 금속 격납고 한 동의 외벽과 상부 창열이 있고, 외벽 투광등 두 개가 보인다. 격납고 내부는 화면 밖이다. 갈라진 콘크리트 바닥과 노란 곡선 표시, 먼 산과 공항 불빛이 이전 장면과 연결된다. 두 사람은 구조물 밖의 열린 지면에 있으며, 이동 공간은 전경에서 중앙 원경으로 이어지지만 인물이 화면 높이의 상당 부분을 차지한다.",
        "entities": "등장 대상은 현우와 찰리 둘뿐이며 헬기는 없다. 현우는 검은 헝클어진 머리, 마른 체격, 회색 셔츠, 올리브색 바지와 낡은 신발을 유지하고 배낭 하나를 멘다. 얼굴이 가려져 정확한 나이와 민족적 외모는 확인할 수 없다. 찰리는 베이지색 각진 장갑, 긴 팔, 짙은 바지와 초록 장화를 유지하지만 배낭이 없다. 등 쪽 원형 개구부는 보이나 가슴 손상 상태는 이 시점에서 판정할 수 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우는 오른발로 지면을 딛고 왼발을 뒤로 들어 밑창이 보이며, 찰리는 왼발을 딛고 오른발을 든다. 각각 지지발이 있어 보행 중 체중 이동으로 가능한 자세다. 현우의 배낭은 어깨끈으로 지지되며, 지지 없이 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "두 사람은 카메라에 등을 보이고 중앙 원경의 열린 비행장으로 이동한다. 얼굴과 눈은 보이지 않는다. 찰리는 현우의 왼쪽에 거의 나란히 있어, 현우 뒤를 따르는 관계가 뚜렷하지 않다.",
        "built_space": "왼쪽 격납고 한 동에 열린 출입구 하나, 내부 철골과 천장등 하나, 외벽 투광등 두 개가 보인다. 낡은 금속 벽과 넓은 균열 콘크리트 바닥, 노란 바닥선 및 원경 불빛이 장소 참조와 부합한다. 두 사람을 작게 배치해 전경부터 중앙 원경까지 출발 경로가 넓게 보인다. 다만 화면 맨 아래 중앙에 모니터와 손잡이가 달린 촬영 장비 일부가 끼어 있다.",
        "entities": "현우와 찰리 두 대상이 있으며 둘 다 배낭을 착용한다. 현우의 검은 머리, 회색 셔츠와 올리브색 바지, 찰리의 베이지 장갑과 긴 팔, 짙은 바지와 초록 장화가 참조에 대응한다. 뒷모습이므로 현우의 얼굴 정체성과 찰리의 마스크 얼굴 및 가슴 손상은 확인할 수 없다. 헬기나 추가 인물은 없으나, 요구되지 않은 촬영 장비가 전경에 추가되었다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "화면 하단 중앙에 프롬프트에 없는 모니터와 손잡이 형태의 촬영 장비가 노출되어, 장면 밖 장비가 최종 스틸에 유입되었다."
        ],
        "physics": "찰리는 오른발을 지면에 두고 왼발을 들어 올렸으며, 현우도 오른발을 딛고 왼발을 뒤로 움직이는 보행 자세다. 두 몸 모두 지지발이 있고 배낭은 어깨끈에 걸려 있어 부유하지 않는다. 하단 장비는 프레임 경계에 잘려 지지부가 보이지 않지만, 공중에 떠 있다고 판단할 근거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.429,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.179,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면 하단 중앙에 프롬프트에 없는 촬영용 카메라 장비가 노출됨 (leaked marker/object)",
     "[gpt-high] 화면 하단 중앙에 프롬프트에 없는 모니터와 손잡이 형태의 촬영 장비가 노출되어, 장면 밖 장비가 최종 스틸에 유입되었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "찰리의 배낭 착용 지시가 누락되었으나, 두 인물의 걷는 동작이 지시된 대로 자연스럽게 구현되었고 치명적인 오류가 없습니다."
   },
   {
    "label": "A",
    "score": 179,
    "verdict_ko": "두 인물 모두 배낭을 착용했지만, 화면 하단에 촬영용 카메라 장비가 그대로 노출되는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 화면 하단 중앙에 프롬프트에 없는 촬영용 카메라 장비가 노출됨 (leaked marker/object) / [gpt-high] 화면 하단 중앙에 프롬프트에 없는 모니터와 손잡이 형태의 촬영 장비가 노출되어, 장면 밖 장비가 최종 스틸에 유입되었다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43_sel.png",
    "asset_id": "4d87f145-d8de-481b-9869-506e0788c8b2",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1096323>",
    "asset_id": "11be7b7a-d72b-40b1-a881-ecc8892413a2",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-d067-7546-95c7-49955461b982",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S77sh43"
  }
 },
 "S77sh53::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:42:50.469774+00:00",
  "fingerprint": "56c452f6ec1e876ef47fc2e12575e4680dd34d33643b4e6bf7c4b00e24349f01",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S77sh53_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S77sh53_sel.png",
  "source_sha256": "ffdb62080cfa1f8bcc869238871ed14ae1c83e5092c5f30748772d1e206aa1b3",
  "file": "S77sh53_cine.png",
  "staged_sha256": "0c184d3d0e2b576c928d38d6002db5261b145ab6bc32da0b2919612f2f00c7ab",
  "latency_ms": 10560
 },
 "S78sh2::signage": {
  "fp": "ae9740d88c7c9b67",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::e1bbaed3c4359ef2": {
  "subjects": [],
  "subject_text": "목포 항구 부둣가와 생선 가판대\n출렁이는 바다와 맞닿은 낡은 부두. 가장자리에 굵은 닻줄과 계류 시설이 있고 가까운 가판대에는 생선이 늘어서 있다.",
  "identity": "canonical",
  "scope_id": "L260",
  "scope_role": "location_exterior",
  "scope_sha": "9499f1d955760266"
 },
 "groupbg::harbor_waiting_quay": {
  "input_fingerprint": "177135af6ca670da",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "harbor_waiting_quay",
    "tags": [
     "S78sh2"
    ]
   },
   "context_sig": "1bb958259d7437ee"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 커다란 모포를 뒤집어쓴 찰리와 현우가 부둣가에 걸터앉아있다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 커다란 모포를 뒤집어쓴 찰리와 현우가 부둣가에 걸터앉아있다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_waiting_quay_bf7967.png",
  "asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4",
  "input_asset_ids": [
   "aed4a301-2447-4839-a8a5-7d2e86e4dc08"
  ],
  "origin_tag": "S78sh2",
  "place_text": "On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.",
  "origin_inputs": {
   "place_text": "On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.",
   "time_of_day_en": "dusk",
   "conti_asset_id": "aed4a301-2447-4839-a8a5-7d2e86e4dc08"
  }
 },
 "S78sh2::bgfirst_bg": {
  "input_fingerprint": "15e5c8425aab4c98",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh2__bgfirst_bg.png",
  "asset_id": "f461492e-6de8-4f80-be43-4a1003286582",
  "input_asset_ids": [
   "aed4a301-2447-4839-a8a5-7d2e86e4dc08",
   "0467ac8e-a464-46ad-b316-68bc8f56f1d4"
  ]
 },
 "S78sh2": {
  "input_fingerprint": "f71e4f2b255abab3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boats are moored along the harbor in the evening. Charlie sits on the quay under a large blanket, concealing his damaged body; the backpack brought from the airport has no stated removal. 현우: He sits on the quay, still dirty and battered, carrying the backpack from the airport.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boats are moored along the harbor in the evening. Charlie sits on the quay under a large blanket, concealing his damaged body; the backpack brought from the airport has no stated removal. 현우: He sits on the quay, still dirty and battered, carrying the backpack from the airport.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boats are moored along the harbor in the evening. Charlie sits on the quay under a large blanket, concealing his damaged body; the backpack brought from the airport has no stated removal. 현우: He sits on the quay, still dirty and battered, carrying the backpack from the airport.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh2__bgfirst_bg.png",
     "asset_id": "f461492e-6de8-4f80-be43-4a1003286582",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S78sh2.png",
     "asset_id": "aed4a301-2447-4839-a8a5-7d2e86e4dc08",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_waiting_quay_bf7967.png",
     "asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우는 화면 오른쪽의 정박된 배들을 향해 시선을 두고 있으며, 찰리는 고개를 숙여 자신의 손 부근을 바라보고 있음.",
    "built_space": "부둣가 모서리에 인물들이 앉아 있고, 화면 왼쪽에 어시장의 천막과 조명, 오른쪽에 바다와 배들이 정확한 위치에 렌더링됨.",
    "entities": "현우는 모포를 어깨에만 둘렀고 머리가 노출됨. 찰리는 마스크를 썼으나, 레퍼런스의 로봇 몸체 대신 인간의 피부가 있는 맨손과 바지, 등산화를 착용함.",
    "hard_violations": [
     "[gemini-pro] 찰리의 로봇 몸체가 인간의 손과 발(신발)로 대체되어 생성됨 (physically impossible anatomy)"
    ],
    "physics": "두 인물 모두 부둣가 끝에 엉덩이를 대고 안정적으로 앉아 있으며, 배낭은 바닥에 놓여 있음."
   },
   {
    "label": "B",
    "direction": "두 인물 모두 카메라를 등지고 바다와 수평선 너머를 향해 시선을 두고 있음.",
    "built_space": "레퍼런스 이미지와 동일한 부둣가 모서리 환경이며, 왼쪽의 어시장 구조물과 오른쪽 바다의 배치가 요구사항을 잘 충족함.",
    "entities": "왼쪽 인물(찰리)은 모포를 머리끝까지 완전히 덮어쓰고 있으나, 오른쪽의 현우는 모포를 전혀 덮지 않은 채 배낭만 메고 있음.",
    "hard_violations": [],
    "physics": "부둣가 가장자리에 자연스럽게 걸터앉아 있으며, 왼쪽 인물의 배낭은 덮어쓴 모포 위로 어깨끈이 정상적으로 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우가 모포를 덮고 있지 않아 프롬프트 지시를 완전히 따르지는 못했으나, 배경 묘사와 구도가 우수하고 치명적인 캐릭터 왜곡이 없어 더 나은 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 인물 모두 모포를 머리끝까지 덮지 않았으며, 특히 찰리의 기계 몸체가 인간의 육체로 잘못 렌더링되는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽의 정박된 배들을 향해 시선을 두고 있으며, 찰리는 고개를 숙여 자신의 손 부근을 바라보고 있음.",
        "built_space": "부둣가 모서리에 인물들이 앉아 있고, 화면 왼쪽에 어시장의 천막과 조명, 오른쪽에 바다와 배들이 정확한 위치에 렌더링됨.",
        "entities": "현우는 모포를 어깨에만 둘렀고 머리가 노출됨. 찰리는 마스크를 썼으나, 레퍼런스의 로봇 몸체 대신 인간의 피부가 있는 맨손과 바지, 등산화를 착용함.",
        "hard_violations": [
         "찰리의 로봇 몸체가 인간의 손과 발(신발)로 대체되어 생성됨 (physically impossible anatomy)"
        ],
        "physics": "두 인물 모두 부둣가 끝에 엉덩이를 대고 안정적으로 앉아 있으며, 배낭은 바닥에 놓여 있음."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 카메라를 등지고 바다와 수평선 너머를 향해 시선을 두고 있음.",
        "built_space": "레퍼런스 이미지와 동일한 부둣가 모서리 환경이며, 왼쪽의 어시장 구조물과 오른쪽 바다의 배치가 요구사항을 잘 충족함.",
        "entities": "왼쪽 인물(찰리)은 모포를 머리끝까지 완전히 덮어쓰고 있으나, 오른쪽의 현우는 모포를 전혀 덮지 않은 채 배낭만 메고 있음.",
        "hard_violations": [],
        "physics": "부둣가 가장자리에 자연스럽게 걸터앉아 있으며, 왼쪽 인물의 배낭은 덮어쓴 모포 위로 어깨끈이 정상적으로 지지되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우가 모포를 덮고 있지 않아 프롬프트 지시를 완전히 따르지는 못했으나, 배경 묘사와 구도가 우수하고 치명적인 캐릭터 왜곡이 없어 더 나은 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 인물 모두 모포를 머리끝까지 덮지 않았으며, 특히 찰리의 기계 몸체가 인간의 육체로 잘못 렌더링되는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽의 정박된 배들을 향해 시선을 두고 있으며, 찰리는 고개를 숙여 자신의 손 부근을 바라보고 있음.",
        "built_space": "부둣가 모서리에 인물들이 앉아 있고, 화면 왼쪽에 어시장의 천막과 조명, 오른쪽에 바다와 배들이 정확한 위치에 렌더링됨.",
        "entities": "현우는 모포를 어깨에만 둘렀고 머리가 노출됨. 찰리는 마스크를 썼으나, 레퍼런스의 로봇 몸체 대신 인간의 피부가 있는 맨손과 바지, 등산화를 착용함.",
        "hard_violations": [
         "찰리의 로봇 몸체가 인간의 손과 발(신발)로 대체되어 생성됨 (physically impossible anatomy)"
        ],
        "physics": "두 인물 모두 부둣가 끝에 엉덩이를 대고 안정적으로 앉아 있으며, 배낭은 바닥에 놓여 있음."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 카메라를 등지고 바다와 수평선 너머를 향해 시선을 두고 있음.",
        "built_space": "레퍼런스 이미지와 동일한 부둣가 모서리 환경이며, 왼쪽의 어시장 구조물과 오른쪽 바다의 배치가 요구사항을 잘 충족함.",
        "entities": "왼쪽 인물(찰리)은 모포를 머리끝까지 완전히 덮어쓰고 있으나, 오른쪽의 현우는 모포를 전혀 덮지 않은 채 배낭만 메고 있음.",
        "hard_violations": [],
        "physics": "부둣가 가장자리에 자연스럽게 걸터앉아 있으며, 왼쪽 인물의 배낭은 덮어쓴 모포 위로 어깨끈이 정상적으로 지지되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "항구 장소와 찰리의 완전한 가림은 잘 지켰지만, 현우가 모포를 전혀 두르지 않았고 하체가 가려져 요청한 두 인물의 전신 연출이 약하다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "두 사람이 모포를 두르고 나란히 앉은 전신 와이드숏은 더 충실하지만, 현우의 머리가 노출되고 찰리의 얼굴·손·체형도 가림 지시와 정체성에 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 사람의 몸통은 카메라 반대편 바다 쪽을 향한다. 현우는 고개를 왼쪽 아래로 돌려 찰리 쪽 가까운 수면을 보는 듯하며, 오른쪽 정박선들을 직접 바라보지는 않는다. 찰리는 머리 전체가 가려져 시선을 확인할 수 없다. 무기나 방향성 있는 휴대 도구는 없다.",
        "built_space": "왼쪽에 천막 생선 좌판과 상자들, 바퀴 달린 운반대 한 대, 뒤로 이어지는 가로등 열이 있다. 오른쪽 전경에는 밧줄을 감은 큰 계선주 한 개가 있고, 부두 가장자리를 따라 작은 계선주들이 이어진다. 먼바다에는 방파제 두 구간과 왼쪽 등대 한 기, 오른쪽에는 여러 정박선이 보인다. 장소의 재료와 배치는 참조와 대체로 일치한다. 두 사람은 전경 석재 턱에 등을 보이고 앉아 있으며, 부두와 턱이 하체를 가려 머리부터 발까지의 형태는 확인되지 않는다. 배들은 수면 너머에 적절한 크기로 배치되어 있다.",
        "entities": "인물은 두 명뿐이다. 왼쪽 찰리는 큰 모포로 머리와 몸을 덮었고 옆에 배낭이 보인다. 얼굴과 장갑판은 가려져 확인할 수 없으며, 현우보다 넓은 등은 보인다. 오른쪽 현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로, 회색 셔츠와 올리브색 바지, 등에 멘 배낭이 참조와 부합한다. 다만 현우에게는 모포가 없고 얼굴의 부상 여부도 확인하기 어렵다. 바다, 생선 좌판, 정박선과 황혼은 구현되었고 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 석재 턱에 지지된다. 찰리의 모포는 머리와 어깨 위에서 내려와 턱에 걸쳐 있으며, 배낭도 몸과 턱에 기대어 있다. 현우의 배낭은 어깨끈으로 지지된다. 다리와 발의 접촉은 가려져 확인할 수 없지만 공중에 떠 있는 신체는 보이지 않는다. 배는 수면에 떠 있고 계선줄과 물 위 반사도 물리적으로 자연스럽다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 정박선들이 있는 항구 쪽을 향해 탐색 대상과 대체로 맞는다. 찰리는 고개를 숙여 자신의 맞잡은 손 또는 무릎 쪽을 향하며 배를 보지는 않는다. 두 사람의 다리는 카메라 쪽 부두 바닥을 향한다. 무기나 조작 중인 도구는 없다.",
        "built_space": "왼쪽 천막 좌판과 생선 상자, 운반대 한 대, 가로등 열, 먼 방파제 두 구간과 등대 한 기가 참조 장소를 유지한다. 오른쪽 전경에는 밧줄을 감은 큰 계선주 한 개가 있고, 두 사람 뒤쪽에도 계선주 한 개가 뚜렷하며 나머지는 부두선을 따라 멀어진다. 두 사람은 같은 석재 턱 위에 나란히 앉아 발을 전경 바닥에 놓는다. 머리부터 신발까지 화면 안에 들어오는 전신 와이드숏이다. 바다와 작은 정박선들은 인물 뒤쪽에 분리되어 있으며, 고정 시설의 불가능한 중복이나 반사는 보이지 않는다.",
        "entities": "두 명만 등장하고 각각 모포와 배낭을 지닌다. 현우는 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 얼굴의 상처와 더러움, 올리브색 바지와 어두운 신발로 참조에 비교적 가깝다. 그러나 모포가 어깨까지만 올라와 머리가 완전히 노출된다. 찰리는 모포를 머리에 썼지만 얼굴 일부가 드러나고, 맞잡은 맨손과 일반적인 바지·신발이 보인다. 육중한 고릴라형 장갑 몸체와 짧은 다리라는 정체성은 충분히 재현되지 않는다. 생선 좌판, 바다, 정박선과 황혼은 맞으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 석재 턱에, 신발은 아래 부두 바닥에 닿아 안정적으로 지지된다. 찰리의 손은 무릎 사이에서 맞잡혀 있고 현우의 팔은 모포 안에 가려져 있다. 모포는 어깨와 무릎을 따라 접혀 턱 위로 늘어지며, 배낭은 등과 턱에 기대어 있다. 지지 없이 떠 있는 신체나 물체는 없고 앉은 자세 자체는 가능하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "항구 장소와 찰리의 완전한 가림은 잘 지켰지만, 현우가 모포를 전혀 두르지 않았고 하체가 가려져 요청한 두 인물의 전신 연출이 약하다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "두 사람이 모포를 두르고 나란히 앉은 전신 와이드숏은 더 충실하지만, 현우의 머리가 노출되고 찰리의 얼굴·손·체형도 가림 지시와 정체성에 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "두 사람의 몸통은 카메라 반대편 바다 쪽을 향한다. 현우는 고개를 왼쪽 아래로 돌려 찰리 쪽 가까운 수면을 보는 듯하며, 오른쪽 정박선들을 직접 바라보지는 않는다. 찰리는 머리 전체가 가려져 시선을 확인할 수 없다. 무기나 방향성 있는 휴대 도구는 없다.",
        "built_space": "왼쪽에 천막 생선 좌판과 상자들, 바퀴 달린 운반대 한 대, 뒤로 이어지는 가로등 열이 있다. 오른쪽 전경에는 밧줄을 감은 큰 계선주 한 개가 있고, 부두 가장자리를 따라 작은 계선주들이 이어진다. 먼바다에는 방파제 두 구간과 왼쪽 등대 한 기, 오른쪽에는 여러 정박선이 보인다. 장소의 재료와 배치는 참조와 대체로 일치한다. 두 사람은 전경 석재 턱에 등을 보이고 앉아 있으며, 부두와 턱이 하체를 가려 머리부터 발까지의 형태는 확인되지 않는다. 배들은 수면 너머에 적절한 크기로 배치되어 있다.",
        "entities": "인물은 두 명뿐이다. 왼쪽 찰리는 큰 모포로 머리와 몸을 덮었고 옆에 배낭이 보인다. 얼굴과 장갑판은 가려져 확인할 수 없으며, 현우보다 넓은 등은 보인다. 오른쪽 현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로, 회색 셔츠와 올리브색 바지, 등에 멘 배낭이 참조와 부합한다. 다만 현우에게는 모포가 없고 얼굴의 부상 여부도 확인하기 어렵다. 바다, 생선 좌판, 정박선과 황혼은 구현되었고 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 석재 턱에 지지된다. 찰리의 모포는 머리와 어깨 위에서 내려와 턱에 걸쳐 있으며, 배낭도 몸과 턱에 기대어 있다. 현우의 배낭은 어깨끈으로 지지된다. 다리와 발의 접촉은 가려져 확인할 수 없지만 공중에 떠 있는 신체는 보이지 않는다. 배는 수면에 떠 있고 계선줄과 물 위 반사도 물리적으로 자연스럽다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 정박선들이 있는 항구 쪽을 향해 탐색 대상과 대체로 맞는다. 찰리는 고개를 숙여 자신의 맞잡은 손 또는 무릎 쪽을 향하며 배를 보지는 않는다. 두 사람의 다리는 카메라 쪽 부두 바닥을 향한다. 무기나 조작 중인 도구는 없다.",
        "built_space": "왼쪽 천막 좌판과 생선 상자, 운반대 한 대, 가로등 열, 먼 방파제 두 구간과 등대 한 기가 참조 장소를 유지한다. 오른쪽 전경에는 밧줄을 감은 큰 계선주 한 개가 있고, 두 사람 뒤쪽에도 계선주 한 개가 뚜렷하며 나머지는 부두선을 따라 멀어진다. 두 사람은 같은 석재 턱 위에 나란히 앉아 발을 전경 바닥에 놓는다. 머리부터 신발까지 화면 안에 들어오는 전신 와이드숏이다. 바다와 작은 정박선들은 인물 뒤쪽에 분리되어 있으며, 고정 시설의 불가능한 중복이나 반사는 보이지 않는다.",
        "entities": "두 명만 등장하고 각각 모포와 배낭을 지닌다. 현우는 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 얼굴의 상처와 더러움, 올리브색 바지와 어두운 신발로 참조에 비교적 가깝다. 그러나 모포가 어깨까지만 올라와 머리가 완전히 노출된다. 찰리는 모포를 머리에 썼지만 얼굴 일부가 드러나고, 맞잡은 맨손과 일반적인 바지·신발이 보인다. 육중한 고릴라형 장갑 몸체와 짧은 다리라는 정체성은 충분히 재현되지 않는다. 생선 좌판, 바다, 정박선과 황혼은 맞으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 석재 턱에, 신발은 아래 부두 바닥에 닿아 안정적으로 지지된다. 찰리의 손은 무릎 사이에서 맞잡혀 있고 현우의 팔은 모포 안에 가려져 있다. 모포는 어깨와 무릎을 따라 접혀 턱 위로 늘어지며, 배낭은 등과 턱에 기대어 있다. 지지 없이 떠 있는 신체나 물체는 없고 앉은 자세 자체는 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.8
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.8
   },
   "violations": {
    "A": [
     "[gemini-pro] 찰리의 로봇 몸체가 인간의 손과 발(신발)로 대체되어 생성됨 (physically impossible anatomy)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1800,
   "A": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1800,
    "verdict_ko": "현우가 모포를 덮고 있지 않아 프롬프트 지시를 완전히 따르지는 못했으나, 배경 묘사와 구도가 우수하고 치명적인 캐릭터 왜곡이 없어 더 나은 결과물입니다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "두 인물 모두 모포를 머리끝까지 덮지 않았으며, 특히 찰리의 기계 몸체가 인간의 육체로 잘못 렌더링되는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 찰리의 로봇 몸체가 인간의 손과 발(신발)로 대체되어 생성됨 (physically impossible anatomy)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_waiting_quay_bf7967.png",
    "asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-d216-71f5-bbe0-a7b84628716c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh2__bgfirst_bg.png",
   "bg_asset_id": "f461492e-6de8-4f80-be43-4a1003286582",
   "bg_record_key": "S78sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "harbor_waiting_quay",
   "groupbg_asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S78sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:44:03.607938+00:00",
  "fingerprint": "415d948ccc83084305b3e4aa87b37db557bbd5c9998be40b20a70ae0339a00ff",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S78sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S78sh2_sel.png",
  "source_sha256": "88f13421ef13a57e030392210dcaa9a9e04f04c804d4b7fc5b7e339eb335c6a9",
  "file": "S78sh2_cine.png",
  "staged_sha256": "dcb0f0b20df2a8b5ede5224aac34418e4413574d65d2527963ba80f8955c1625",
  "latency_ms": 11842
 },
 "S78sh13::signage": {
  "fp": "7eca9ae6b44448ef",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::harbor_repair_berth": {
  "input_fingerprint": "b1bcb0ab845f2c42",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "harbor_repair_berth",
    "tags": [
     "S78sh13"
    ]
   },
   "context_sig": "b8355d48debd79dc"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 선원이 어디론가로 현우를 인도한다.\n- 낡은 배를 정비하는 40대 초반의 크리스와 몇몇 선원들 보인다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 선원이 어디론가로 현우를 인도한다.\n- 낡은 배를 정비하는 40대 초반의 크리스와 몇몇 선원들 보인다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_repair_berth_41df90.png",
  "asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3",
  "input_asset_ids": [
   "4effe083-7d46-4dd7-a28a-33b44efc4e7b"
  ],
  "origin_tag": "S78sh13",
  "place_text": "At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.",
  "origin_inputs": {
   "place_text": "At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.",
   "time_of_day_en": "dusk",
   "conti_asset_id": "4effe083-7d46-4dd7-a28a-33b44efc4e7b"
  }
 },
 "S78sh13::bgfirst_bg": {
  "input_fingerprint": "3ae8336ec28768d4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh13__bgfirst_bg.png",
  "asset_id": "640ad084-58c6-44a6-b7fc-1706ec522e95",
  "input_asset_ids": [
   "4effe083-7d46-4dd7-a28a-33b44efc4e7b",
   "7aa8b28b-26f7-41d5-b9ce-185137890bc3"
  ]
 },
 "S78sh13": {
  "input_fingerprint": "70eb343a2504cf48",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old boat remains moored for maintenance. Charlie's large blanket conceals his damaged body, and his travel backpack has not been discarded. 크리스: He is beside the old boat, where he has been carrying out maintenance.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 크리스 right now, so 크리스's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 크리스: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 크리스 (한국인 남성, 40대 초반, 중년 초입의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old boat remains moored for maintenance. Charlie's large blanket conceals his damaged body, and his travel backpack has not been discarded. 크리스: He is beside the old boat, where he has been carrying out maintenance.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 크리스 right now, so 크리스's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 크리스: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 크리스 (한국인 남성, 40대 초반, 중년 초입의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old boat remains moored for maintenance. Charlie's large blanket conceals his damaged body, and his travel backpack has not been discarded. 크리스: He is beside the old boat, where he has been carrying out maintenance.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 크리스 right now, so 크리스's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 크리스: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 크리스 (한국인 남성, 40대 초반, 중년 초입의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh13__bgfirst_bg.png",
     "asset_id": "640ad084-58c6-44a6-b7fc-1706ec522e95",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S78sh13.png",
     "asset_id": "4effe083-7d46-4dd7-a28a-33b44efc4e7b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 크리스: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1199325>",
     "asset_id": "303f2ae8-9f48-4964-8a8b-c24edec958ba",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_repair_berth_41df90.png",
     "asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 크리스: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1199325>",
     "asset_id": "303f2ae8-9f48-4964-8a8b-c24edec958ba",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "크리스의 시선은 화면 좌측 전경에 서 있는 두 인물(현우와 찰리)을 정확히 향하고 있음.",
    "built_space": "항구 야외 작업장이며, 우측 뒤편으로 낡은 배의 선체가 비스듬하게 올바른 비율로 배치되어 있음.",
    "entities": "크리스(레퍼런스 외모 일치), 전경의 뒷모습 두 인물(한 명은 모포를 두름), 손에 든 용접기 모두 확인됨.",
    "hard_violations": [
     "[gpt-high] 크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
    ],
    "physics": "크리스가 한 손에 용접기를 들고 서 있는 자세 및 전경 인물들의 서 있는 모습에서 물리적 지지 및 무게 중심이 자연스러움."
   },
   {
    "label": "B",
    "direction": "크리스의 시선이 지시된 대상이 아닌 카메라 렌즈를 직접 향하고 있음.",
    "built_space": "낡은 배가 배경 전체를 차지하고 있으며, 인물 앞쪽에 정체불명의 금속판이 배치되어 있음.",
    "entities": "크리스(레퍼런스 일치)와 장갑 낀 손에 든 용접기는 보이나, 현우와 찰리는 프레임 내에 없음.",
    "hard_violations": [],
    "physics": "크리스가 금속판 위에 턱을 괴고 있으며, 지시와 달리 용접기를 밑으로 내리지 않고 얼굴 앞까지 끌어올려 판 위에 얹어둔 상태임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 얼굴 클로즈업보다 프레임이 다소 넓어졌으나, 내린 용접기와 전경의 인물들을 향한 시선 등 프롬프트의 상황을 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업 프레이밍에 가깝지만, 용접기를 얼굴 높이로 들어 올리고 렌즈를 정면 응시하여 '내린 용접기'와 주시 대상에 대한 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "크리스의 시선은 화면 좌측 전경에 서 있는 두 인물(현우와 찰리)을 정확히 향하고 있음.",
        "built_space": "항구 야외 작업장이며, 우측 뒤편으로 낡은 배의 선체가 비스듬하게 올바른 비율로 배치되어 있음.",
        "entities": "크리스(레퍼런스 외모 일치), 전경의 뒷모습 두 인물(한 명은 모포를 두름), 손에 든 용접기 모두 확인됨.",
        "hard_violations": [],
        "physics": "크리스가 한 손에 용접기를 들고 서 있는 자세 및 전경 인물들의 서 있는 모습에서 물리적 지지 및 무게 중심이 자연스러움."
       },
       {
        "label": "B",
        "direction": "크리스의 시선이 지시된 대상이 아닌 카메라 렌즈를 직접 향하고 있음.",
        "built_space": "낡은 배가 배경 전체를 차지하고 있으며, 인물 앞쪽에 정체불명의 금속판이 배치되어 있음.",
        "entities": "크리스(레퍼런스 일치)와 장갑 낀 손에 든 용접기는 보이나, 현우와 찰리는 프레임 내에 없음.",
        "hard_violations": [],
        "physics": "크리스가 금속판 위에 턱을 괴고 있으며, 지시와 달리 용접기를 밑으로 내리지 않고 얼굴 앞까지 끌어올려 판 위에 얹어둔 상태임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 얼굴 클로즈업보다 프레임이 다소 넓어졌으나, 내린 용접기와 전경의 인물들을 향한 시선 등 프롬프트의 상황을 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업 프레이밍에 가깝지만, 용접기를 얼굴 높이로 들어 올리고 렌즈를 정면 응시하여 '내린 용접기'와 주시 대상에 대한 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "크리스의 시선은 화면 좌측 전경에 서 있는 두 인물(현우와 찰리)을 정확히 향하고 있음.",
        "built_space": "항구 야외 작업장이며, 우측 뒤편으로 낡은 배의 선체가 비스듬하게 올바른 비율로 배치되어 있음.",
        "entities": "크리스(레퍼런스 외모 일치), 전경의 뒷모습 두 인물(한 명은 모포를 두름), 손에 든 용접기 모두 확인됨.",
        "hard_violations": [],
        "physics": "크리스가 한 손에 용접기를 들고 서 있는 자세 및 전경 인물들의 서 있는 모습에서 물리적 지지 및 무게 중심이 자연스러움."
       },
       {
        "label": "B",
        "direction": "크리스의 시선이 지시된 대상이 아닌 카메라 렌즈를 직접 향하고 있음.",
        "built_space": "낡은 배가 배경 전체를 차지하고 있으며, 인물 앞쪽에 정체불명의 금속판이 배치되어 있음.",
        "entities": "크리스(레퍼런스 일치)와 장갑 낀 손에 든 용접기는 보이나, 현우와 찰리는 프레임 내에 없음.",
        "hard_violations": [],
        "physics": "크리스가 금속판 위에 턱을 괴고 있으며, 지시와 달리 용접기를 밑으로 내리지 않고 얼굴 앞까지 끌어올려 판 위에 얹어둔 상태임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "얼굴 클로즈업, 어두운 저녁빛, 손으로 잡아 내려놓은 용접기는 요구에 가깝지만, 렌즈를 향한 시선과 작업면에 지나치게 가까운 얼굴 때문에 두 방문자를 주시하는 순간은 덜 명확하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "방문자를 향한 차가운 시선과 항구 장소는 분명하지만, 두 인물을 추가하고 상반신·항구 전경으로 넓혀 크리스만의 얼굴 클로즈업이라는 핵심 구도를 어겼다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "크리스의 얼굴과 두 눈은 거의 렌즈 정면을 향한다. 현우와 찰리는 화면 밖이므로 실제 시선이 두 사람에게 닿는지는 확인되지 않는다. 용접 토치 끝은 화면 오른쪽 아래의 철판을 향하며 사람을 겨누지 않는다.",
        "built_space": "뒤에는 녹슨 흰색 선체와 파란 띠, 왼쪽의 원형 창 하나, 상단 난간 일부가 비스듬하게 보인다. 장소 사진의 배 재질과 색은 이어지며 배 전체나 항구 전경으로 확대하지 않았다. 전경에는 금속 작업면 하나와 그 위의 작은 철판 하나가 있고, 크리스는 얼굴을 그 작업면 높이 가까이 낮추고 있다. 계류 상태와 부두의 나머지 설비는 이 크롭에서 확인되지 않는다.",
        "entities": "보이는 사람은 크리스 한 명이다. 검은 머리의 중년 초입 동아시아계 남성으로, 참고 인물의 얼굴 윤곽과 눈·코 형태에 대체로 가깝다. 눈은 정상적인 사람의 눈이며 표정은 굳어 있다. 남색 계열 옷이지만 참고의 흰 셔츠와 니트 조합보다는 작업복처럼 보인다. 장갑 낀 손에 용접 토치가 있고, 현우·찰리·모포·배낭은 얼굴 중심 구도 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "용접 토치 손잡이는 크리스의 장갑 낀 손이 잡고 있고, 노즐과 철판은 금속 작업면에 놓여 있다. 손과 팔은 화면 왼쪽의 소매로 연결되어 소유 관계가 자연스럽다. 얼굴을 작업면에 가깝게 낮춘 자세는 가능하지만 방문자를 살피는 동작으로는 다소 어색하다. 하체가 잘렸다는 이유로 부유한다고 볼 근거는 없다."
       },
       {
        "label": "B",
        "direction": "크리스는 화면 왼쪽 전경의 검은 머리 남성을 향해 눈을 돌리고 있다. 그 옆에는 모포를 뒤집어쓴 인물이 있어 방문자들을 살피는 관계가 읽힌다. 다만 모포 속 인물을 직접 응시하는 순간은 아니다. 허리 가까이 잡은 토치의 노즐은 왼쪽 위를 향하며 방문자들의 얼굴을 직접 겨누지는 않는다.",
        "built_space": "오른쪽에 녹슨 흰색 배, 파란 띠, 원형 창 두 개, 난간, 기둥과 켜진 등 하나가 보인다. 중앙 부두에는 계선주 하나와 계류 밧줄이 있고, 뒤로 바다·방파제·등대 하나가 보인다. 장소 사진의 구조는 잘 이어지지만, 선체의 작은 사선 조각만 남기는 대신 항구와 부두를 넓게 보여준다. 크리스는 오른쪽 배 옆에 서 있고 두 방문자는 왼쪽 전경을 차지한다.",
        "entities": "크리스는 참고와 닮은 검은 머리의 중년 초입 동아시아계 남성이며, 회녹색 작업 재킷과 밝은 속옷을 입어 참고 의상과 다르다. 맨손에 용접 토치를 들고 있다. 추가로 뒷머리와 어깨가 보이는 남성 한 명, 회색 모포로 머리와 몸을 덮은 인물 한 명이 있다. 전경 남성에게 배낭 끈처럼 보이는 부분은 있으나 찰리의 배낭인지는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
        ],
        "physics": "크리스의 손가락이 토치 손잡이를 감싸고 손목과 소매가 자연스럽게 이어져 도구의 지지가 분명하다. 토치는 작동하지 않고 낮게 들려 있다. 인물들의 하체는 프레임 밖이지만 부두에 서 있는 상체 배치로 읽히며 부유나 불가능한 관절은 보이지 않는다. 모포는 안쪽 신체를 따라 어깨에서 아래로 드리워진다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 클로즈업, 어두운 저녁빛, 손으로 잡아 내려놓은 용접기는 요구에 가깝지만, 렌즈를 향한 시선과 작업면에 지나치게 가까운 얼굴 때문에 두 방문자를 주시하는 순간은 덜 명확하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "방문자를 향한 차가운 시선과 항구 장소는 분명하지만, 두 인물을 추가하고 상반신·항구 전경으로 넓혀 크리스만의 얼굴 클로즈업이라는 핵심 구도를 어겼다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "크리스의 얼굴과 두 눈은 거의 렌즈 정면을 향한다. 현우와 찰리는 화면 밖이므로 실제 시선이 두 사람에게 닿는지는 확인되지 않는다. 용접 토치 끝은 화면 오른쪽 아래의 철판을 향하며 사람을 겨누지 않는다.",
        "built_space": "뒤에는 녹슨 흰색 선체와 파란 띠, 왼쪽의 원형 창 하나, 상단 난간 일부가 비스듬하게 보인다. 장소 사진의 배 재질과 색은 이어지며 배 전체나 항구 전경으로 확대하지 않았다. 전경에는 금속 작업면 하나와 그 위의 작은 철판 하나가 있고, 크리스는 얼굴을 그 작업면 높이 가까이 낮추고 있다. 계류 상태와 부두의 나머지 설비는 이 크롭에서 확인되지 않는다.",
        "entities": "보이는 사람은 크리스 한 명이다. 검은 머리의 중년 초입 동아시아계 남성으로, 참고 인물의 얼굴 윤곽과 눈·코 형태에 대체로 가깝다. 눈은 정상적인 사람의 눈이며 표정은 굳어 있다. 남색 계열 옷이지만 참고의 흰 셔츠와 니트 조합보다는 작업복처럼 보인다. 장갑 낀 손에 용접 토치가 있고, 현우·찰리·모포·배낭은 얼굴 중심 구도 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "용접 토치 손잡이는 크리스의 장갑 낀 손이 잡고 있고, 노즐과 철판은 금속 작업면에 놓여 있다. 손과 팔은 화면 왼쪽의 소매로 연결되어 소유 관계가 자연스럽다. 얼굴을 작업면에 가깝게 낮춘 자세는 가능하지만 방문자를 살피는 동작으로는 다소 어색하다. 하체가 잘렸다는 이유로 부유한다고 볼 근거는 없다."
       },
       {
        "label": "A",
        "direction": "크리스는 화면 왼쪽 전경의 검은 머리 남성을 향해 눈을 돌리고 있다. 그 옆에는 모포를 뒤집어쓴 인물이 있어 방문자들을 살피는 관계가 읽힌다. 다만 모포 속 인물을 직접 응시하는 순간은 아니다. 허리 가까이 잡은 토치의 노즐은 왼쪽 위를 향하며 방문자들의 얼굴을 직접 겨누지는 않는다.",
        "built_space": "오른쪽에 녹슨 흰색 배, 파란 띠, 원형 창 두 개, 난간, 기둥과 켜진 등 하나가 보인다. 중앙 부두에는 계선주 하나와 계류 밧줄이 있고, 뒤로 바다·방파제·등대 하나가 보인다. 장소 사진의 구조는 잘 이어지지만, 선체의 작은 사선 조각만 남기는 대신 항구와 부두를 넓게 보여준다. 크리스는 오른쪽 배 옆에 서 있고 두 방문자는 왼쪽 전경을 차지한다.",
        "entities": "크리스는 참고와 닮은 검은 머리의 중년 초입 동아시아계 남성이며, 회녹색 작업 재킷과 밝은 속옷을 입어 참고 의상과 다르다. 맨손에 용접 토치를 들고 있다. 추가로 뒷머리와 어깨가 보이는 남성 한 명, 회색 모포로 머리와 몸을 덮은 인물 한 명이 있다. 전경 남성에게 배낭 끈처럼 보이는 부분은 있으나 찰리의 배낭인지는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
        ],
        "physics": "크리스의 손가락이 토치 손잡이를 감싸고 손목과 소매가 자연스럽게 이어져 도구의 지지가 분명하다. 토치는 작동하지 않고 낮게 들려 있다. 인물들의 하체는 프레임 밖이지만 부두에 서 있는 상체 배치로 읽히며 부유나 불가능한 관절은 보이지 않는다. 모포는 안쪽 신체를 따라 어깨에서 아래로 드리워진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.571
   },
   "violations": {
    "A": [
     "[gpt-high] 크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "요구된 얼굴 클로즈업보다 프레임이 다소 넓어졌으나, 내린 용접기와 전경의 인물들을 향한 시선 등 프롬프트의 상황을 매우 충실하게 구현했습니다.  ★위반: [gpt-high] 크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "클로즈업 프레이밍에 가깝지만, 용접기를 얼굴 높이로 들어 올리고 렌즈를 정면 응시하여 '내린 용접기'와 주시 대상에 대한 지시를 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_repair_berth_41df90.png",
    "asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 크리스: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1199325>",
    "asset_id": "303f2ae8-9f48-4964-8a8b-c24edec958ba",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-d6ef-7fd1-a884-2dcc76b80470",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh13__bgfirst_bg.png",
   "bg_asset_id": "640ad084-58c6-44a6-b7fc-1706ec522e95",
   "bg_record_key": "S78sh13::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "harbor_repair_berth",
   "groupbg_asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S78sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:32:35.525077+00:00",
  "fingerprint": "d434835cb30aa30bafd6297afc894fe5b598318a8ddcebbcce54675a8a2d2275",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S78sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S78sh13_sel.png",
  "source_sha256": "a1a8ea4b4190de3815c6de2df799147cc4f74674e528a5867b4554e502b4df6b",
  "file": "S78sh13_cine.png",
  "staged_sha256": "48e3a6b7546bb37ee141a797d692a974028fa050556930e37ebb4832efe0af5d",
  "latency_ms": 12043
 },
 "S78sh17::signage": {
  "fp": "8bc706f713ce35a3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S78sh17": {
  "input_fingerprint": "d5e4a5bba528fc14",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 어두운 바다를 가르는 하얀 물살을 남긴 채, 부둣가에서 멀어진 위치에 떠 있는 낡은 배의 뒷모습 풀샷.\n\nLOCATION (lock): On the dark harbor water just beyond the quay, where the small departing boat leaves a white wake. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Departing boat in the upper-center of the frame, background, moves toward open sea beyond the upper frame; Trailing wake in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Departing old boat (Moving away from the harbor) — Stern and one side remain visible in rear three-quarter view; used as Small, complete subject against the surrounding sea; Wake (White disturbed water trailing behind the boat); used as Connects the distant stern to the lower foreground; Sea (Dark, with water disturbed along the departure path); used as Provides open space around the diminishing vessel.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark evening sea and white wake provide restrained tonal separation around the receding boat.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small motorboat is leaving the harbor under power, churning a wake. Charlie is aboard, still covered by the blanket, with his damaged body and travel backpack unchanged.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 어두운 바다를 가르는 하얀 물살을 남긴 채, 부둣가에서 멀어진 위치에 떠 있는 낡은 배의 뒷모습 풀샷.\n\nLOCATION (lock): On the dark harbor water just beyond the quay, where the small departing boat leaves a white wake. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Departing boat in the upper-center of the frame, background, moves toward open sea beyond the upper frame; Trailing wake in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Departing old boat (Moving away from the harbor) — Stern and one side remain visible in rear three-quarter view; used as Small, complete subject against the surrounding sea; Wake (White disturbed water trailing behind the boat); used as Connects the distant stern to the lower foreground; Sea (Dark, with water disturbed along the departure path); used as Provides open space around the diminishing vessel.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark evening sea and white wake provide restrained tonal separation around the receding boat.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small motorboat is leaving the harbor under power, churning a wake. Charlie is aboard, still covered by the blanket, with his damaged body and travel backpack unchanged.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 어두운 바다를 가르는 하얀 물살을 남긴 채, 부둣가에서 멀어진 위치에 떠 있는 낡은 배의 뒷모습 풀샷.\n\nLOCATION (lock): On the dark harbor water just beyond the quay, where the small departing boat leaves a white wake. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Departing boat in the upper-center of the frame, background, moves toward open sea beyond the upper frame; Trailing wake in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Departing old boat (Moving away from the harbor) — Stern and one side remain visible in rear three-quarter view; used as Small, complete subject against the surrounding sea; Wake (White disturbed water trailing behind the boat); used as Connects the distant stern to the lower foreground; Sea (Dark, with water disturbed along the departure path); used as Provides open space around the diminishing vessel.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark evening sea and white wake provide restrained tonal separation around the receding boat.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small motorboat is leaving the harbor under power, churning a wake. Charlie is aboard, still covered by the blanket, with his damaged body and travel backpack unchanged.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "배의 선수가 화면 상단의 탁 트인 먼 바다를 향해 올바르게 나아감.",
    "built_space": "인공 구조물 없이 어두운 바다와 배만 단독으로 배치됨.",
    "entities": "낡은 배, 하얀 물살, 어두운 바다가 프롬프트대로 묘사되었고 인물은 없음.",
    "hard_violations": [],
    "physics": "배가 수면 위에 안정적으로 떠서 항해하며 자연스러운 항적을 만듦."
   },
   {
    "label": "B",
    "direction": "배가 화면 상단으로 향하지만 뱃머리 방향에 육지와 항구가 가로막고 있어 먼 바다 방향과 모순됨.",
    "built_space": "화면 상단 배경에 방파제, 건물, 항구 조명 시설이 위치함.",
    "entities": "낡은 배와 하얀 항적이 있으나, 선실 부근에 사람의 실루엣이 나타남.",
    "hard_violations": [
     "[gemini-pro] 명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
     "[gemini-pro] 배가 먼 바다가 아닌 항구를 향하는 구도적 모순",
     "[gpt-high] 선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
     "[gpt-high] 장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
    ],
    "physics": "배가 물 위에 떠서 이동하며 물살을 일으킴."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물 배제 규칙을 정확히 따랐으며, 명시된 대로 화면 상단의 먼 바다를 향해 나아가는 낡은 배의 구도를 성공적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "명시된 인물 배제 규칙을 어기고 선상에 사람을 묘사했으며, 배가 먼 바다가 아닌 항구 육지 쪽을 향하고 있어 치명적인 오류를 범함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "배의 선수가 화면 상단의 탁 트인 먼 바다를 향해 올바르게 나아감.",
        "built_space": "인공 구조물 없이 어두운 바다와 배만 단독으로 배치됨.",
        "entities": "낡은 배, 하얀 물살, 어두운 바다가 프롬프트대로 묘사되었고 인물은 없음.",
        "hard_violations": [],
        "physics": "배가 수면 위에 안정적으로 떠서 항해하며 자연스러운 항적을 만듦."
       },
       {
        "label": "B",
        "direction": "배가 화면 상단으로 향하지만 뱃머리 방향에 육지와 항구가 가로막고 있어 먼 바다 방향과 모순됨.",
        "built_space": "화면 상단 배경에 방파제, 건물, 항구 조명 시설이 위치함.",
        "entities": "낡은 배와 하얀 항적이 있으나, 선실 부근에 사람의 실루엣이 나타남.",
        "hard_violations": [
         "명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
         "배가 먼 바다가 아닌 항구를 향하는 구도적 모순"
        ],
        "physics": "배가 물 위에 떠서 이동하며 물살을 일으킴."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물 배제 규칙을 정확히 따랐으며, 명시된 대로 화면 상단의 먼 바다를 향해 나아가는 낡은 배의 구도를 성공적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "명시된 인물 배제 규칙을 어기고 선상에 사람을 묘사했으며, 배가 먼 바다가 아닌 항구 육지 쪽을 향하고 있어 치명적인 오류를 범함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "배의 선수가 화면 상단의 탁 트인 먼 바다를 향해 올바르게 나아감.",
        "built_space": "인공 구조물 없이 어두운 바다와 배만 단독으로 배치됨.",
        "entities": "낡은 배, 하얀 물살, 어두운 바다가 프롬프트대로 묘사되었고 인물은 없음.",
        "hard_violations": [],
        "physics": "배가 수면 위에 안정적으로 떠서 항해하며 자연스러운 항적을 만듦."
       },
       {
        "label": "B",
        "direction": "배가 화면 상단으로 향하지만 뱃머리 방향에 육지와 항구가 가로막고 있어 먼 바다 방향과 모순됨.",
        "built_space": "화면 상단 배경에 방파제, 건물, 항구 조명 시설이 위치함.",
        "entities": "낡은 배와 하얀 항적이 있으나, 선실 부근에 사람의 실루엣이 나타남.",
        "hard_violations": [
         "명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
         "배가 먼 바다가 아닌 항구를 향하는 구도적 모순"
        ],
        "physics": "배가 물 위에 떠서 이동하며 물살을 일으킴."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "선미와 흰 항적의 방향은 맞지만, 선내 인물이 노출되어 인물 금지 조건을 어기고 배도 요구된 원경의 작은 피사체보다 크게 보입니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "인물 없이 상단 중앙의 낡은 배를 후방 사선으로 보여주며, 어두운 바다와 하단 전경으로 이어지는 흰 항적이 지정된 구도를 충실히 구현합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽으로 향하고, 선미와 한쪽 측면이 보인다. 흰 항적은 선미에서 하단 중앙 전경으로 이어져 카메라에서 멀어지는 운항 방향과 맞는다. 선내 인물의 시선은 어두워 판별하기 어렵다.",
        "built_space": "배에는 조타실 하나와 뒤쪽 갑판을 덮는 차양 하나가 보인다. 차양 아래 뒤쪽에 담요를 두른 인물 형상이 있고, 그 앞 조타실에도 별도의 어두운 인물 실루엣이 보인다. 상단 배경에는 부두 건물 여러 동, 정박선 여러 척, 방파제와 탑형 시설이 추가되어 있으며, 배와 바다 중심의 제한된 장소 묘사보다 많은 구조물을 도입했다.",
        "entities": "낡은 소형 동력선 한 척, 어두운 바다, 흰 항적과 해질녘의 낮은 조도는 일치한다. 그러나 담요를 두른 사람과 조타실의 사람 형상이 보여 인물 금지 조건에 맞지 않는다. 얼굴·연령·민족성·신체 손상과 배낭 상태는 확인할 수 없다. 선미의 표식은 선명하게 읽히지 않는다.",
        "hard_violations": [
         "선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
         "장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
        ],
        "physics": "선체는 수면에 잠겨 부력으로 지지되고, 선미 바로 뒤의 거품과 양옆으로 퍼지는 물결은 동력 운항으로 설명된다. 뒤쪽 인물은 갑판의 좌석 부근에 앉아 있고 담요는 몸 위에 걸쳐져 있다. 조타실 인물은 하체가 가려져 있으나 공중에 떠 있는 정황은 없다."
       },
       {
        "label": "B",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽의 열린 바다로 향한다. 카메라에는 선미와 한쪽 측면이 함께 보이며, 항적은 선미에서 하단 중앙과 오른쪽 전경으로 퍼져 멀어지는 배의 진행 방향과 일치한다.",
        "built_space": "배에는 조타실 하나, 선미의 열린 갑판 하나와 측면 난간이 보인다. 배 전체가 상단 중앙에 작게 들어오고 주위는 바다로 비어 있다. 별도의 부두 건물이나 정박선을 추가하지 않았으며, 사람이 서거나 앉은 모습과 부자연스러운 반사는 보이지 않는다.",
        "entities": "마모된 소형 동력선 한 척, 어두운 바다와 선미부터 이어지는 흰 항적이 확인된다. 살아 있는 사람이나 얼굴은 보이지 않는다. 찰리의 담요·손상된 몸·배낭은 화면에서 식별되지 않으며, 인물을 노출하지 않는 이 구도에서는 이를 불일치로 볼 근거가 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "배는 수면에 정상적으로 잠겨 부력으로 지지된다. 선미에 붙어 시작하는 거품과 뒤로 넓어지는 항적은 추진하며 떠나는 동력선의 움직임에 부합한다. 조타실과 난간은 선체에 고정되어 있고, 지지 없이 떠 있는 물체나 인체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "선미와 흰 항적의 방향은 맞지만, 선내 인물이 노출되어 인물 금지 조건을 어기고 배도 요구된 원경의 작은 피사체보다 크게 보입니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "인물 없이 상단 중앙의 낡은 배를 후방 사선으로 보여주며, 어두운 바다와 하단 전경으로 이어지는 흰 항적이 지정된 구도를 충실히 구현합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽으로 향하고, 선미와 한쪽 측면이 보인다. 흰 항적은 선미에서 하단 중앙 전경으로 이어져 카메라에서 멀어지는 운항 방향과 맞는다. 선내 인물의 시선은 어두워 판별하기 어렵다.",
        "built_space": "배에는 조타실 하나와 뒤쪽 갑판을 덮는 차양 하나가 보인다. 차양 아래 뒤쪽에 담요를 두른 인물 형상이 있고, 그 앞 조타실에도 별도의 어두운 인물 실루엣이 보인다. 상단 배경에는 부두 건물 여러 동, 정박선 여러 척, 방파제와 탑형 시설이 추가되어 있으며, 배와 바다 중심의 제한된 장소 묘사보다 많은 구조물을 도입했다.",
        "entities": "낡은 소형 동력선 한 척, 어두운 바다, 흰 항적과 해질녘의 낮은 조도는 일치한다. 그러나 담요를 두른 사람과 조타실의 사람 형상이 보여 인물 금지 조건에 맞지 않는다. 얼굴·연령·민족성·신체 손상과 배낭 상태는 확인할 수 없다. 선미의 표식은 선명하게 읽히지 않는다.",
        "hard_violations": [
         "선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
         "장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
        ],
        "physics": "선체는 수면에 잠겨 부력으로 지지되고, 선미 바로 뒤의 거품과 양옆으로 퍼지는 물결은 동력 운항으로 설명된다. 뒤쪽 인물은 갑판의 좌석 부근에 앉아 있고 담요는 몸 위에 걸쳐져 있다. 조타실 인물은 하체가 가려져 있으나 공중에 떠 있는 정황은 없다."
       },
       {
        "label": "A",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽의 열린 바다로 향한다. 카메라에는 선미와 한쪽 측면이 함께 보이며, 항적은 선미에서 하단 중앙과 오른쪽 전경으로 퍼져 멀어지는 배의 진행 방향과 일치한다.",
        "built_space": "배에는 조타실 하나, 선미의 열린 갑판 하나와 측면 난간이 보인다. 배 전체가 상단 중앙에 작게 들어오고 주위는 바다로 비어 있다. 별도의 부두 건물이나 정박선을 추가하지 않았으며, 사람이 서거나 앉은 모습과 부자연스러운 반사는 보이지 않는다.",
        "entities": "마모된 소형 동력선 한 척, 어두운 바다와 선미부터 이어지는 흰 항적이 확인된다. 살아 있는 사람이나 얼굴은 보이지 않는다. 찰리의 담요·손상된 몸·배낭은 화면에서 식별되지 않으며, 인물을 노출하지 않는 이 구도에서는 이를 불일치로 볼 근거가 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "배는 수면에 정상적으로 잠겨 부력으로 지지된다. 선미에 붙어 시작하는 거품과 뒤로 넓어지는 항적은 추진하며 떠나는 동력선의 움직임에 부합한다. 조타실과 난간은 선체에 고정되어 있고, 지지 없이 떠 있는 물체나 인체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.651
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.401
   },
   "violations": {
    "B": [
     "[gemini-pro] 명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
     "[gemini-pro] 배가 먼 바다가 아닌 항구를 향하는 구도적 모순",
     "[gpt-high] 선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
     "[gpt-high] 장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 401
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "인물 배제 규칙을 정확히 따랐으며, 명시된 대로 화면 상단의 먼 바다를 향해 나아가는 낡은 배의 구도를 성공적으로 구현함."
   },
   {
    "label": "B",
    "score": 401,
    "verdict_ko": "명시된 인물 배제 규칙을 어기고 선상에 사람을 묘사했으며, 배가 먼 바다가 아닌 항구 육지 쪽을 향하고 있어 치명적인 오류를 범함.  ★위반: [gemini-pro] 명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성 / [gemini-pro] 배가 먼 바다가 아닌 항구를 향하는 구도적 모순 / [gpt-high] 선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다. / [gpt-high] 장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-dbc0-72f2-9e91-44e492509d60",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S78sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:52:44.466196+00:00",
  "fingerprint": "3bfc1aeb855b481a554a465f8647a1253dee84f964eba36c55f956985bbbb523",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S78sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S78sh17_sel.png",
  "source_sha256": "4b146f4966b201e9eb79bda26cdd6b094b7e1c64bf4d9a8e11b55fa7ec889813",
  "file": "S78sh17_cine.png",
  "staged_sha256": "986982dba03a996a80b3be4cc33055bc82e4640ac21f5c1e72d296f29c8ff8b9",
  "latency_ms": 9225
 },
 "S79sh10::signage": {
  "fp": "8ad1272ac5edb949",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S79sh10": {
  "input_fingerprint": "0c29929adfa7793f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small two-person cabin door is closed, and the sofa remains the resting place. Charlie reclines there with his accumulated body damage; the blanket and travel backpack have no stated removal. 현우: He reclines on the sofa with his legs stretched out, still dirty and battered; his travel backpack remains with him.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small two-person cabin door is closed, and the sofa remains the resting place. Charlie reclines there with his accumulated body damage; the blanket and travel backpack have no stated removal. 현우: He reclines on the sofa with his legs stretched out, still dirty and battered; his travel backpack remains with him.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small two-person cabin door is closed, and the sofa remains the resting place. Charlie reclines there with his accumulated body damage; the blanket and travel backpack have no stated removal. 현우: He reclines on the sofa with his legs stretched out, still dirty and battered; his travel backpack remains with him.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10__bgfirst_bg.png",
     "asset_id": "9ca43f94-b62c-4185-bbd5-d874b7069c57",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S79sh10.png",
     "asset_id": "471ec738-cab9-44f7-b95c-4082e4ab2709",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L263B02.png",
     "asset_id": "a3a7948a-4ed6-4034-bdbc-66720a85c15b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 전경 좌측의 현우를 향해 고개를 돌리고 시선을 맞추고 있음.",
    "built_space": "레퍼런스와 동일한 선실 내부로 소파, 뒤편 수납장, 창문, 지도가 정확한 위치에 배치됨.",
    "entities": "찰리는 흰색 마스크, 베이지색 장갑, 담요를 착용하고 붉은 디지털 눈과 입을 표시함; 현우는 헝클어진 머리로 전경에 어깨와 뒷모습이 걸쳐 있음.",
    "hard_violations": [],
    "physics": "두 인물 모두 소파에 체중을 싣고 자연스럽게 기대어 앉아 있음."
   },
   {
    "label": "B",
    "direction": "찰리는 화면 우측을 응시하며, 배경의 현우는 시선이 아래를 향함.",
    "built_space": "소파는 존재하나 우측 배경에 레퍼런스에 없는 2층 침대 구조물이 렌더링됨.",
    "entities": "찰리는 눈 이모티콘을 띄웠으나 얼굴 전체가 베이지색임; 현우는 전경이 아닌 배경에 작게 배치됨.",
    "hard_violations": [
     "[gemini-pro] invented objects: 공간 레퍼런스에 존재하지 않는 2층 침대 구조물 생성",
     "[gemini-pro] person placed where the staging does not put them: 현우가 전경의 어깨 위치가 아닌 배경에 배치됨"
    ],
    "physics": "찰리와 현우 모두 각자의 위치에서 구조물에 앉아 체중을 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 흰색 마스크와 캐릭터 디자인을 정확히 반영했으며, 현우의 어깨 너머로 보이는 지정된 구도와 선실 배경을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 마스크가 흰색이 아닌 베이지색으로 렌더링되었고, 지정된 전경 구도를 무시한 채 레퍼런스에 없는 침대 구조물을 배경에 추가했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 전경 좌측의 현우를 향해 고개를 돌리고 시선을 맞추고 있음.",
        "built_space": "레퍼런스와 동일한 선실 내부로 소파, 뒤편 수납장, 창문, 지도가 정확한 위치에 배치됨.",
        "entities": "찰리는 흰색 마스크, 베이지색 장갑, 담요를 착용하고 붉은 디지털 눈과 입을 표시함; 현우는 헝클어진 머리로 전경에 어깨와 뒷모습이 걸쳐 있음.",
        "hard_violations": [],
        "physics": "두 인물 모두 소파에 체중을 싣고 자연스럽게 기대어 앉아 있음."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 우측을 응시하며, 배경의 현우는 시선이 아래를 향함.",
        "built_space": "소파는 존재하나 우측 배경에 레퍼런스에 없는 2층 침대 구조물이 렌더링됨.",
        "entities": "찰리는 눈 이모티콘을 띄웠으나 얼굴 전체가 베이지색임; 현우는 전경이 아닌 배경에 작게 배치됨.",
        "hard_violations": [
         "invented objects: 공간 레퍼런스에 존재하지 않는 2층 침대 구조물 생성",
         "person placed where the staging does not put them: 현우가 전경의 어깨 위치가 아닌 배경에 배치됨"
        ],
        "physics": "찰리와 현우 모두 각자의 위치에서 구조물에 앉아 체중을 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 흰색 마스크와 캐릭터 디자인을 정확히 반영했으며, 현우의 어깨 너머로 보이는 지정된 구도와 선실 배경을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 마스크가 흰색이 아닌 베이지색으로 렌더링되었고, 지정된 전경 구도를 무시한 채 레퍼런스에 없는 침대 구조물을 배경에 추가했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 전경 좌측의 현우를 향해 고개를 돌리고 시선을 맞추고 있음.",
        "built_space": "레퍼런스와 동일한 선실 내부로 소파, 뒤편 수납장, 창문, 지도가 정확한 위치에 배치됨.",
        "entities": "찰리는 흰색 마스크, 베이지색 장갑, 담요를 착용하고 붉은 디지털 눈과 입을 표시함; 현우는 헝클어진 머리로 전경에 어깨와 뒷모습이 걸쳐 있음.",
        "hard_violations": [],
        "physics": "두 인물 모두 소파에 체중을 싣고 자연스럽게 기대어 앉아 있음."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 우측을 응시하며, 배경의 현우는 시선이 아래를 향함.",
        "built_space": "소파는 존재하나 우측 배경에 레퍼런스에 없는 2층 침대 구조물이 렌더링됨.",
        "entities": "찰리는 눈 이모티콘을 띄웠으나 얼굴 전체가 베이지색임; 현우는 전경이 아닌 배경에 작게 배치됨.",
        "hard_violations": [
         "invented objects: 공간 레퍼런스에 존재하지 않는 2층 침대 구조물 생성",
         "person placed where the staging does not put them: 현우가 전경의 어깨 위치가 아닌 배경에 배치됨"
        ],
        "physics": "찰리와 현우 모두 각자의 위치에서 구조물에 앉아 체중을 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찰리의 낡은 얼굴과 비대칭 디지털 눈을 크게 잡아 핵심 클로즈업을 충족하지만, 흰 마스크의 색과 선실 배경의 일치도는 아쉽다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "선실 구조와 찰리의 외형·담요, 현우를 향한 반응은 충실하지만, 현우의 상체와 배경까지 넓게 담은 어깨너머 구도가 얼굴 클로즈업 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 화면 오른쪽으로 조금 돌아 있으나 디지털 눈은 대체로 카메라 쪽을 향한다. 오른쪽 가장자리의 현우와 눈을 맞추는 모습은 아니다. 현우는 화면 오른쪽 바깥으로 고개를 돌리고 있어 눈의 정확한 방향은 보이지 않는다. 질문에 대한 표정 반응은 읽히며, 지시문이 직접적인 눈맞춤까지 요구하지는 않는다.",
        "built_space": "회색 직물 소파 하나의 등받이가 찰리 뒤에서 현우 쪽까지 이어진다. 베이지색 벽 패널과 여러 사진, 오른쪽 뒤의 수평 선반 같은 구조가 보인다. 소파와 벽의 재질은 장소 참조와 유사하지만, 뒤쪽 구조는 참조의 지도·책장·창문 배치를 명확히 재현하지 않는다. 문과 주요 고정 설비는 대부분 화면 밖이라 개수나 개폐 상태를 확인할 수 없다.",
        "entities": "찰리 한 명과 현우의 머리·어깨 일부가 보인다. 찰리는 긁히고 벗겨진 샌드 베이지 장갑, 귀 옆 기계 장치, 안테나, 각진 얼굴을 갖췄다. 다만 참조의 흰 마스크 부분까지 베이지색에 가까워졌다. 서로 다른 모양의 흰 점형 디지털 눈은 복잡한 감정을 전달한다. 현우의 검은 머리와 짙은 상의는 참조와 부합하지만 얼굴이 잘려 나이와 얼굴 동일성은 확정하기 어렵다. 현우 곁에 여행 배낭이 보이며, 찰리의 담요는 이 크롭에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 머리는 노출된 기계 목으로 몸통에 연결되고, 어깨와 등은 소파 등받이에 기대어 있다. 현우의 어깨도 같은 소파 앞에 위치하며 배낭은 몸과 소파 사이에 받쳐진 것으로 보인다. 하체가 잘려 다리를 뻗은 상태는 판단할 수 없다. 떠 있는 신체나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 왼쪽 앞의 현우를 향해 머리를 기울이고 디지털 눈을 돌린다. 현우도 찰리 쪽으로 얼굴을 돌려 두 인물의 대화 방향이 명확하다. 찰리의 올라간 안쪽 눈썹과 처진 입 모양은 걱정스럽고 난처한 반응으로 읽힌다.",
        "built_space": "회색 소파 하나, 벽 지도 하나, 책장 하나, 상단 선풍기 하나, 커튼이 달린 야간 창문 하나가 보인다. 창문 아래에는 수납장과 금속 컵·보온병·바구니가 있어 장소 참조의 배치와 매우 가깝다. 찰리는 등받이에 기대고 현우는 같은 소파의 전경에서 몸을 돌리고 있다. 다만 배경과 현우의 상체가 넓게 드러나, 소파를 좁게 남기는 찰리 얼굴 클로즈업보다 넓은 구도다. 출입문은 화면 밖이다.",
        "entities": "찰리와 현우만 보인다. 찰리의 흰 각진 마스크, 샌드 베이지 장갑, 육중한 상체, 안테나와 회색 줄무늬 담요는 참조에 가깝다. 붉은 점형 눈과 흰 점형 눈썹·입이 얼굴 표면의 발광 요소로 표현되어 있다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로 보이며 피부에 때와 상처가 있다. 다만 참조의 남색 상의 대신 낡은 민소매 차림이고, 뒷모습 위주여서 정확한 얼굴 동일성은 확인하기 어렵다. 아래쪽에는 다른 담요가, 오른쪽에는 배낭 일부가 보인다. 배경 종이의 글씨는 판독되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 몸통은 소파에 기대어 지지되고 머리는 기계 목에 연결되어 있다. 줄무늬 담요는 어깨에 걸쳐 몸통을 따라 자연스럽게 처진다. 현우는 소파 전경에서 찰리 쪽으로 상체를 돌린 자세이며 공중에 뜬 부분은 없다. 다만 현우가 편히 뒤로 누운 상태보다는 상체를 일으킨 모습에 가깝고, 다리를 뻗었는지는 화면 밖이라 확인할 수 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "찰리의 낡은 얼굴과 비대칭 디지털 눈을 크게 잡아 핵심 클로즈업을 충족하지만, 흰 마스크의 색과 선실 배경의 일치도는 아쉽다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "선실 구조와 찰리의 외형·담요, 현우를 향한 반응은 충실하지만, 현우의 상체와 배경까지 넓게 담은 어깨너머 구도가 얼굴 클로즈업 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 화면 오른쪽으로 조금 돌아 있으나 디지털 눈은 대체로 카메라 쪽을 향한다. 오른쪽 가장자리의 현우와 눈을 맞추는 모습은 아니다. 현우는 화면 오른쪽 바깥으로 고개를 돌리고 있어 눈의 정확한 방향은 보이지 않는다. 질문에 대한 표정 반응은 읽히며, 지시문이 직접적인 눈맞춤까지 요구하지는 않는다.",
        "built_space": "회색 직물 소파 하나의 등받이가 찰리 뒤에서 현우 쪽까지 이어진다. 베이지색 벽 패널과 여러 사진, 오른쪽 뒤의 수평 선반 같은 구조가 보인다. 소파와 벽의 재질은 장소 참조와 유사하지만, 뒤쪽 구조는 참조의 지도·책장·창문 배치를 명확히 재현하지 않는다. 문과 주요 고정 설비는 대부분 화면 밖이라 개수나 개폐 상태를 확인할 수 없다.",
        "entities": "찰리 한 명과 현우의 머리·어깨 일부가 보인다. 찰리는 긁히고 벗겨진 샌드 베이지 장갑, 귀 옆 기계 장치, 안테나, 각진 얼굴을 갖췄다. 다만 참조의 흰 마스크 부분까지 베이지색에 가까워졌다. 서로 다른 모양의 흰 점형 디지털 눈은 복잡한 감정을 전달한다. 현우의 검은 머리와 짙은 상의는 참조와 부합하지만 얼굴이 잘려 나이와 얼굴 동일성은 확정하기 어렵다. 현우 곁에 여행 배낭이 보이며, 찰리의 담요는 이 크롭에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 머리는 노출된 기계 목으로 몸통에 연결되고, 어깨와 등은 소파 등받이에 기대어 있다. 현우의 어깨도 같은 소파 앞에 위치하며 배낭은 몸과 소파 사이에 받쳐진 것으로 보인다. 하체가 잘려 다리를 뻗은 상태는 판단할 수 없다. 떠 있는 신체나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 화면 왼쪽 앞의 현우를 향해 머리를 기울이고 디지털 눈을 돌린다. 현우도 찰리 쪽으로 얼굴을 돌려 두 인물의 대화 방향이 명확하다. 찰리의 올라간 안쪽 눈썹과 처진 입 모양은 걱정스럽고 난처한 반응으로 읽힌다.",
        "built_space": "회색 소파 하나, 벽 지도 하나, 책장 하나, 상단 선풍기 하나, 커튼이 달린 야간 창문 하나가 보인다. 창문 아래에는 수납장과 금속 컵·보온병·바구니가 있어 장소 참조의 배치와 매우 가깝다. 찰리는 등받이에 기대고 현우는 같은 소파의 전경에서 몸을 돌리고 있다. 다만 배경과 현우의 상체가 넓게 드러나, 소파를 좁게 남기는 찰리 얼굴 클로즈업보다 넓은 구도다. 출입문은 화면 밖이다.",
        "entities": "찰리와 현우만 보인다. 찰리의 흰 각진 마스크, 샌드 베이지 장갑, 육중한 상체, 안테나와 회색 줄무늬 담요는 참조에 가깝다. 붉은 점형 눈과 흰 점형 눈썹·입이 얼굴 표면의 발광 요소로 표현되어 있다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로 보이며 피부에 때와 상처가 있다. 다만 참조의 남색 상의 대신 낡은 민소매 차림이고, 뒷모습 위주여서 정확한 얼굴 동일성은 확인하기 어렵다. 아래쪽에는 다른 담요가, 오른쪽에는 배낭 일부가 보인다. 배경 종이의 글씨는 판독되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 몸통은 소파에 기대어 지지되고 머리는 기계 목에 연결되어 있다. 줄무늬 담요는 어깨에 걸쳐 몸통을 따라 자연스럽게 처진다. 현우는 소파 전경에서 찰리 쪽으로 상체를 돌린 자세이며 공중에 뜬 부분은 없다. 다만 현우가 편히 뒤로 누운 상태보다는 상체를 일으킨 모습에 가깝고, 다리를 뻗었는지는 화면 밖이라 확인할 수 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.321
   },
   "violations": {
    "B": [
     "[gemini-pro] invented objects: 공간 레퍼런스에 존재하지 않는 2층 침대 구조물 생성",
     "[gemini-pro] person placed where the staging does not put them: 현우가 전경의 어깨 위치가 아닌 배경에 배치됨"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "찰리의 흰색 마스크와 캐릭터 디자인을 정확히 반영했으며, 현우의 어깨 너머로 보이는 지정된 구도와 선실 배경을 충실히 구현했습니다."
   },
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "찰리의 마스크가 흰색이 아닌 베이지색으로 렌더링되었고, 지정된 전경 구도를 무시한 채 레퍼런스에 없는 침대 구조물을 배경에 추가했습니다.  ★위반: [gemini-pro] invented objects: 공간 레퍼런스에 존재하지 않는 2층 침대 구조물 생성 / [gemini-pro] person placed where the staging does not put them: 현우가 전경의 어깨 위치가 아닌 배경에 배치됨"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L263B02.png",
    "asset_id": "a3a7948a-4ed6-4034-bdbc-66720a85c15b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-dd5a-7aa3-8425-0bb530cc64ca",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10__bgfirst_bg.png",
   "bg_asset_id": "9ca43f94-b62c-4185-bbd5-d874b7069c57",
   "bg_record_key": "S79sh10::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C01"
  ]
 },
 "S79sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:45:15.361148+00:00",
  "fingerprint": "2270e89a0b53e423744e31fd2adc798b11793a068ed9947fdb4f80a72e8b6e88",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S79sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S79sh10_sel.png",
  "source_sha256": "979462f3ea3883f87b379229412a8ca135c23d0c1deaf08a4ff5688010ebceee",
  "file": "S79sh10_cine.png",
  "staged_sha256": "d0a85a7892b608f7f297d423cd0ba7d63febf4e181f2513a15cb5f7bebbc2f5a",
  "latency_ms": 11007
 },
 "S79sh12::signage": {
  "fp": "80220a2ca3440b3f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S79sh12": {
  "input_fingerprint": "ac9d162b596c91a7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낮게 울리는 진동 속, 깜짝 놀란 표정으로 두 눈을 번쩍 뜬 채 굳어 있는 현우의 고요한 얼굴.\n\nLOCATION (lock): On the sofa inside the small crew cabin aboard the boat, in low nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Still supporting 현우 as he wakes) — A cropped supporting section sits beneath and behind his head; used as Anchors the arrested face to his previous sleeping position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the cabin's ambient illumination unchanged across the temporal cut, with restrained contrast preserving the stillness of 현우's startled face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cabin door remains closed and the sofa remains in place. Charlie is still in the cabin with his damaged body, blanket and travel backpack. 현우: He has fallen asleep on the sofa and is now waking; his dirty clothing and accumulated injuries remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낮게 울리는 진동 속, 깜짝 놀란 표정으로 두 눈을 번쩍 뜬 채 굳어 있는 현우의 고요한 얼굴.\n\nLOCATION (lock): On the sofa inside the small crew cabin aboard the boat, in low nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Still supporting 현우 as he wakes) — A cropped supporting section sits beneath and behind his head; used as Anchors the arrested face to his previous sleeping position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the cabin's ambient illumination unchanged across the temporal cut, with restrained contrast preserving the stillness of 현우's startled face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cabin door remains closed and the sofa remains in place. Charlie is still in the cabin with his damaged body, blanket and travel backpack. 현우: He has fallen asleep on the sofa and is now waking; his dirty clothing and accumulated injuries remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낮게 울리는 진동 속, 깜짝 놀란 표정으로 두 눈을 번쩍 뜬 채 굳어 있는 현우의 고요한 얼굴.\n\nLOCATION (lock): On the sofa inside the small crew cabin aboard the boat, in low nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Still supporting 현우 as he wakes) — A cropped supporting section sits beneath and behind his head; used as Anchors the arrested face to his previous sleeping position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the cabin's ambient illumination unchanged across the temporal cut, with restrained contrast preserving the stillness of 현우's startled face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cabin door remains closed and the sofa remains in place. Charlie is still in the cabin with his damaged body, blanket and travel backpack. 현우: He has fallen asleep on the sofa and is now waking; his dirty clothing and accumulated injuries remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 정면을 향해 고정되어 있음.",
    "built_space": "회색 소파가 인물 뒤를 채우고 있으며, 뒤쪽 벽면의 선실 디테일은 흐릿하게 처리됨.",
    "entities": "현우의 얼굴과 표정은 묘사에 부합하나, 이전 샷의 민소매가 아닌 카라 셔츠를 착용함.",
    "hard_violations": [
     "[gemini-pro] 이전 샷과 동일해야 하는 의상 조건 위배 (민소매 대신 셔츠 착용)"
    ],
    "physics": "목과 상체 근육으로 머리를 지탱하며 소파 앞에 앉아 있음."
   },
   {
    "label": "B",
    "direction": "시선은 화면 오른쪽 위를 향해 놀란 채 멈춰 있음.",
    "built_space": "회색 소파에 기대어 있으며, 배경 벽면에 이전 샷과 동일한 위치의 배관 장치와 지도가 명확히 보임.",
    "entities": "현우의 얼굴, 표정 및 이전 샷과 정확히 일치하는 오염된 민소매 의상 착용.",
    "hard_violations": [],
    "physics": "등과 어깨가 소파 등받이에 기대어 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷의 의상(민소매)과 배경 구조물(지도, 기계 장치)을 완벽하게 유지하며 프롬프트의 연속성 조건을 훌륭하게 충족했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 구도는 좋으나, 이전 샷으로 고정된 의상(민소매) 대신 캐릭터 레퍼런스의 셔츠를 입고 있어 연속성 조건에 크게 위배됩니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면을 향해 고정되어 있음.",
        "built_space": "회색 소파가 인물 뒤를 채우고 있으며, 뒤쪽 벽면의 선실 디테일은 흐릿하게 처리됨.",
        "entities": "현우의 얼굴과 표정은 묘사에 부합하나, 이전 샷의 민소매가 아닌 카라 셔츠를 착용함.",
        "hard_violations": [
         "이전 샷과 동일해야 하는 의상 조건 위배 (민소매 대신 셔츠 착용)"
        ],
        "physics": "목과 상체 근육으로 머리를 지탱하며 소파 앞에 앉아 있음."
       },
       {
        "label": "B",
        "direction": "시선은 화면 오른쪽 위를 향해 놀란 채 멈춰 있음.",
        "built_space": "회색 소파에 기대어 있으며, 배경 벽면에 이전 샷과 동일한 위치의 배관 장치와 지도가 명확히 보임.",
        "entities": "현우의 얼굴, 표정 및 이전 샷과 정확히 일치하는 오염된 민소매 의상 착용.",
        "hard_violations": [],
        "physics": "등과 어깨가 소파 등받이에 기대어 지탱되고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷의 의상(민소매)과 배경 구조물(지도, 기계 장치)을 완벽하게 유지하며 프롬프트의 연속성 조건을 훌륭하게 충족했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 구도는 좋으나, 이전 샷으로 고정된 의상(민소매) 대신 캐릭터 레퍼런스의 셔츠를 입고 있어 연속성 조건에 크게 위배됩니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면을 향해 고정되어 있음.",
        "built_space": "회색 소파가 인물 뒤를 채우고 있으며, 뒤쪽 벽면의 선실 디테일은 흐릿하게 처리됨.",
        "entities": "현우의 얼굴과 표정은 묘사에 부합하나, 이전 샷의 민소매가 아닌 카라 셔츠를 착용함.",
        "hard_violations": [
         "이전 샷과 동일해야 하는 의상 조건 위배 (민소매 대신 셔츠 착용)"
        ],
        "physics": "목과 상체 근육으로 머리를 지탱하며 소파 앞에 앉아 있음."
       },
       {
        "label": "B",
        "direction": "시선은 화면 오른쪽 위를 향해 놀란 채 멈춰 있음.",
        "built_space": "회색 소파에 기대어 있으며, 배경 벽면에 이전 샷과 동일한 위치의 배관 장치와 지도가 명확히 보임.",
        "entities": "현우의 얼굴, 표정 및 이전 샷과 정확히 일치하는 오염된 민소매 의상 착용.",
        "hard_violations": [],
        "physics": "등과 어깨가 소파 등받이에 기대어 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 숏의 더러운 민소매와 낮은 조도는 잘 이어지지만, 어깨·가슴과 선실 배경을 넓게 담아 고요하게 굳은 얼굴 중심의 클로즈업이 덜 정확하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "소파에 지지된 얼굴과 번쩍 뜬 두 눈을 밀착해 담아 핵심 클로즈업에 더 충실하지만, 목 아래 보이는 셔츠 깃은 이전 숏의 민소매와 다르고 조명도 다소 밝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈은 카메라보다 높은 화면 밖 지점을 향한다. 바라보는 대상은 보이지 않으며, 지문도 특정 대상을 지정하지 않았다. 입을 벌리고 위를 응시하는 놀람이 보인다. 무기나 방향성 소품은 없다.",
        "built_space": "회색 직물 소파의 등받이 구획들이 머리와 양어깨 뒤를 채운다. 뒤쪽 벽에는 지도 한 장, 전선관이 연결된 사각 함 하나, 세로 금속 장치 하나가 보인다. 이전 숏의 재질과 선실 분위기는 유사하지만 지도와 금속 장치의 화면상 배치가 다르다. 문과 창은 프레임 밖이며, 반사나 명백한 설비 중복은 없다.",
        "entities": "젊은 동아시아계 남성 한 명이며 검은 머리와 얼굴 윤곽은 현우 참조와 대체로 맞는다. 국적은 외양만으로 확인할 수 없다. 더러운 회갈색 민소매, 피부의 때와 긁힌 흔적은 이전 숏의 상태를 따른다. 눈은 정상적인 홍채와 동공을 갖는다. 다른 인물이나 찰리의 신체는 없고, 담요와 배낭은 이 구도에서 보이지 않는다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "현우는 상체를 뒤로 기대고 있으며 등과 어깨 뒤의 소파 등받이가 몸을 받친다. 머리는 목 위에 자연스럽게 놓여 뒤쪽 쿠션 가까이에 있다. 잠에서 깨 그대로 멈춘 자세로 가능하며, 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "크게 뜬 두 눈은 거의 렌즈 방향의 정면을 향한다. 화면 안에 별도의 시선 대상은 없다. 특정 대상을 요구하지 않는 지문과 충돌하지 않으며, 눈꺼풀과 약간 열린 입으로 놀라 멈춘 순간을 표현한다. 방향성 소품은 없다.",
        "built_space": "회색 직물 소파의 넓은 등받이 구획 하나가 머리 바로 뒤를 받치고, 왼쪽 가장자리에 인접 구획 일부가 보인다. 위쪽에는 베이지색 벽과 오른쪽 상단의 액자 또는 지도 모서리 하나만 남는다. 얼굴 중심의 크롭에 필요한 소파 지지부가 충분히 보이며, 문·창·기타 설비는 프레임 밖이다. 불가능한 반사나 중복 설비는 없다.",
        "entities": "젊은 동아시아계 남성 한 명으로, 앳된 얼굴과 헝클어진 검은 머리가 현우 참조에 가깝다. 국적은 화면만으로 판단할 수 없다. 얼굴의 때, 작은 상처와 땀이 보이고 눈의 구조는 정상적이다. 다만 화면 아래 목 양옆에 회색 셔츠 깃처럼 보이는 옷감이 있어, 이전 숏에 고정된 민소매와 맞지 않는다. 다른 인물은 없으며 찰리·담요·배낭은 클로즈업 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리 뒤와 목 주변에 소파 등받이가 이어지고, 상체는 그쪽으로 기대어 있다. 목과 어깨의 연결이 자연스러우며 소파가 기대는 몸을 지지한다. 얼굴만 긴장한 채 잠에서 깨어 멈춘 자세가 가능하고, 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 숏의 더러운 민소매와 낮은 조도는 잘 이어지지만, 어깨·가슴과 선실 배경을 넓게 담아 고요하게 굳은 얼굴 중심의 클로즈업이 덜 정확하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "소파에 지지된 얼굴과 번쩍 뜬 두 눈을 밀착해 담아 핵심 클로즈업에 더 충실하지만, 목 아래 보이는 셔츠 깃은 이전 숏의 민소매와 다르고 조명도 다소 밝다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 눈은 카메라보다 높은 화면 밖 지점을 향한다. 바라보는 대상은 보이지 않으며, 지문도 특정 대상을 지정하지 않았다. 입을 벌리고 위를 응시하는 놀람이 보인다. 무기나 방향성 소품은 없다.",
        "built_space": "회색 직물 소파의 등받이 구획들이 머리와 양어깨 뒤를 채운다. 뒤쪽 벽에는 지도 한 장, 전선관이 연결된 사각 함 하나, 세로 금속 장치 하나가 보인다. 이전 숏의 재질과 선실 분위기는 유사하지만 지도와 금속 장치의 화면상 배치가 다르다. 문과 창은 프레임 밖이며, 반사나 명백한 설비 중복은 없다.",
        "entities": "젊은 동아시아계 남성 한 명이며 검은 머리와 얼굴 윤곽은 현우 참조와 대체로 맞는다. 국적은 외양만으로 확인할 수 없다. 더러운 회갈색 민소매, 피부의 때와 긁힌 흔적은 이전 숏의 상태를 따른다. 눈은 정상적인 홍채와 동공을 갖는다. 다른 인물이나 찰리의 신체는 없고, 담요와 배낭은 이 구도에서 보이지 않는다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "현우는 상체를 뒤로 기대고 있으며 등과 어깨 뒤의 소파 등받이가 몸을 받친다. 머리는 목 위에 자연스럽게 놓여 뒤쪽 쿠션 가까이에 있다. 잠에서 깨 그대로 멈춘 자세로 가능하며, 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "크게 뜬 두 눈은 거의 렌즈 방향의 정면을 향한다. 화면 안에 별도의 시선 대상은 없다. 특정 대상을 요구하지 않는 지문과 충돌하지 않으며, 눈꺼풀과 약간 열린 입으로 놀라 멈춘 순간을 표현한다. 방향성 소품은 없다.",
        "built_space": "회색 직물 소파의 넓은 등받이 구획 하나가 머리 바로 뒤를 받치고, 왼쪽 가장자리에 인접 구획 일부가 보인다. 위쪽에는 베이지색 벽과 오른쪽 상단의 액자 또는 지도 모서리 하나만 남는다. 얼굴 중심의 크롭에 필요한 소파 지지부가 충분히 보이며, 문·창·기타 설비는 프레임 밖이다. 불가능한 반사나 중복 설비는 없다.",
        "entities": "젊은 동아시아계 남성 한 명으로, 앳된 얼굴과 헝클어진 검은 머리가 현우 참조에 가깝다. 국적은 화면만으로 판단할 수 없다. 얼굴의 때, 작은 상처와 땀이 보이고 눈의 구조는 정상적이다. 다만 화면 아래 목 양옆에 회색 셔츠 깃처럼 보이는 옷감이 있어, 이전 숏에 고정된 민소매와 맞지 않는다. 다른 인물은 없으며 찰리·담요·배낭은 클로즈업 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리 뒤와 목 주변에 소파 등받이가 이어지고, 상체는 그쪽으로 기대어 있다. 목과 어깨의 연결이 자연스러우며 소파가 기대는 몸을 지지한다. 얼굴만 긴장한 채 잠에서 깨어 멈춘 자세가 가능하고, 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.875
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷과 동일해야 하는 의상 조건 위배 (민소매 대신 셔츠 착용)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "이전 샷의 의상(민소매)과 배경 구조물(지도, 기계 장치)을 완벽하게 유지하며 프롬프트의 연속성 조건을 훌륭하게 충족했습니다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "클로즈업 구도는 좋으나, 이전 샷으로 고정된 의상(민소매) 대신 캐릭터 레퍼런스의 셔츠를 입고 있어 연속성 조건에 크게 위배됩니다.  ★위반: [gemini-pro] 이전 샷과 동일해야 하는 의상 조건 위배 (민소매 대신 셔츠 착용)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10_sel.png",
    "asset_id": "b591baf9-3120-47ec-bc8e-6da887eef843",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-e0c0-7af1-94b8-ffe67c02d992",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S79sh10"
  }
 },
 "S79sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:46:07.574736+00:00",
  "fingerprint": "fc40a6e274a89ebccf5feb9d933f970c984e318de7405d756bcebd1136cdb7d8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S79sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S79sh12_sel.png",
  "source_sha256": "7a95266fb3850d20e7dae5013517c93b67670517e436a0daa1b31b8d95d4218e",
  "file": "S79sh12_cine.png",
  "staged_sha256": "175ce54bb0a7f6ce8a1d535a7e30ddc8f9d334f64981bd98318d8dbe4d03c286",
  "latency_ms": 10952
 },
 "S79sh14::signage": {
  "fp": "795724a02daeef62",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S79sh14": {
  "input_fingerprint": "7167e0790fc45f5c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 기울어진 바닥을 타고 선원실 벽 쪽으로 거칠게 미끄러지는 도중, 몸이 아래 방향으로 강하게 쏠린 mid-action 자세의 현우와 찰리 역동적인 전신.\n\nLOCATION (lock): Inside the boat's small crew cabin, along the floor sloping toward the wall as the vessel lists in nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Sofa at the start of the slide in the upper-left of the frame, background; Destination cabin wall in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Sofa (Behind the pair as they slide away) — Viewed obliquely from above at the uphill side of the composition; used as Marks the origin of their displacement; Cabin floor (Sharply tilted with the ship) — Its visible plane slopes through the composition toward lower right; used as Makes the direction of involuntary movement legible; Cabin wall (Ahead of the sliding pair, before contact) — The interior face is visible at the downhill end of their trajectory; used as Defines the imminent collision and remaining clearance.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the cabin's existing ambient illumination and controlled contrast as the physical tilt disrupts the previously quiet composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The vessel and cabin are now sharply heeled to one side, with the cabin door still closed unless subsequently opened. Charlie's damaged body is thrown toward the wall; no removal of his blanket or travel backpack is established. 현우: He is thrown toward the cabin wall as the floor tilts, still bearing his earlier injuries and dirty clothing. The cabin is visibly tilted by the vessel's violent roll.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 기울어진 바닥을 타고 선원실 벽 쪽으로 거칠게 미끄러지는 도중, 몸이 아래 방향으로 강하게 쏠린 mid-action 자세의 현우와 찰리 역동적인 전신.\n\nLOCATION (lock): Inside the boat's small crew cabin, along the floor sloping toward the wall as the vessel lists in nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Sofa at the start of the slide in the upper-left of the frame, background; Destination cabin wall in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Sofa (Behind the pair as they slide away) — Viewed obliquely from above at the uphill side of the composition; used as Marks the origin of their displacement; Cabin floor (Sharply tilted with the ship) — Its visible plane slopes through the composition toward lower right; used as Makes the direction of involuntary movement legible; Cabin wall (Ahead of the sliding pair, before contact) — The interior face is visible at the downhill end of their trajectory; used as Defines the imminent collision and remaining clearance.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the cabin's existing ambient illumination and controlled contrast as the physical tilt disrupts the previously quiet composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The vessel and cabin are now sharply heeled to one side, with the cabin door still closed unless subsequently opened. Charlie's damaged body is thrown toward the wall; no removal of his blanket or travel backpack is established. 현우: He is thrown toward the cabin wall as the floor tilts, still bearing his earlier injuries and dirty clothing. The cabin is visibly tilted by the vessel's violent roll.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 기울어진 바닥을 타고 선원실 벽 쪽으로 거칠게 미끄러지는 도중, 몸이 아래 방향으로 강하게 쏠린 mid-action 자세의 현우와 찰리 역동적인 전신.\n\nLOCATION (lock): Inside the boat's small crew cabin, along the floor sloping toward the wall as the vessel lists in nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Sofa at the start of the slide in the upper-left of the frame, background; Destination cabin wall in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Sofa (Behind the pair as they slide away) — Viewed obliquely from above at the uphill side of the composition; used as Marks the origin of their displacement; Cabin floor (Sharply tilted with the ship) — Its visible plane slopes through the composition toward lower right; used as Makes the direction of involuntary movement legible; Cabin wall (Ahead of the sliding pair, before contact) — The interior face is visible at the downhill end of their trajectory; used as Defines the imminent collision and remaining clearance.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the cabin's existing ambient illumination and controlled contrast as the physical tilt disrupts the previously quiet composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The vessel and cabin are now sharply heeled to one side, with the cabin door still closed unless subsequently opened. Charlie's damaged body is thrown toward the wall; no removal of his blanket or travel backpack is established. 현우: He is thrown toward the cabin wall as the floor tilts, still bearing his earlier injuries and dirty clothing. The cabin is visibly tilted by the vessel's violent roll.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우는 우측 벽을 향하고 있으며, 찰리는 정면 아래쪽 경사면을 향함.",
    "built_space": "좌측 상단 소파, 기울어진 바닥, 우측 벽 배치로 참고 이미지의 공간 및 조명 설정과 잘 부합함.",
    "entities": "현우의 얼굴과 오염된 의상, 찰리의 고릴라형 체형과 흰 마스크, 담요 모두 참고 이미지와 정확히 일치함.",
    "hard_violations": [
     "[gemini-pro] 지지대 없이 수평으로 공중에 완전히 떠 있는 현우의 자세 (물리 법칙 위반)",
     "[gemini-pro] 충돌 전(before contact) 여유 공간을 요구한 프레이밍 지시와 달리 이미 벽에 발이 닿아 있음",
     "[gpt-high] 현우의 골반과 몸통이 바닥 위에 떠 있는데 이를 받치는 손이나 좌면이 보이지 않는다. 두 발은 수직 벽에 닿아 있지만 몸을 매달아 지지할 접점은 없으며, 바닥을 타고 미끄러지는 동작에서 이 공중 자세로 이어지는 이륙이나 낙하도 드러나지 않는다."
    ],
    "physics": "찰리는 바닥에 무릎을 대고 지탱 중이나, 현우는 도약이나 착지 없이 공중에 떠 있으며 몸의 하중을 지지하는 요소가 전혀 없음."
   },
   {
    "label": "B",
    "direction": "현우와 찰리 모두 바닥의 기울기를 따라 우측 하단으로 몸과 시선이 쏠려 있음.",
    "built_space": "좌측 상단 소파와 우측 중경의 벽, 그리고 두 인물 앞의 충돌 전 여유 공간까지 지시된 프레임 구도를 완벽히 충족함.",
    "entities": "현우의 얼굴 및 의상, 찰리의 체형과 샌드 베이지 장갑판 등 모든 설정이 참고 이미지와 일치함.",
    "hard_violations": [],
    "physics": "현우는 두 발로 바닥을 딛고 가방을 짚어 지탱하며, 찰리 역시 바닥에 서서 안정적으로 체중을 지지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 프레임 구도와 충돌 전 여유 공간을 잘 구현했고 인물들이 바닥에 지탱되어 안정적이나, 찰리의 백팩이 분리되고 포즈가 다소 정적인 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우가 공중에 떠 있어 물리적 지지대 규칙을 위반했으며, 충돌 전 여유 공간 지시를 어기고 이미 벽에 닿아 있어 프레이밍 지시를 실패했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 우측 벽을 향하고 있으며, 찰리는 정면 아래쪽 경사면을 향함.",
        "built_space": "좌측 상단 소파, 기울어진 바닥, 우측 벽 배치로 참고 이미지의 공간 및 조명 설정과 잘 부합함.",
        "entities": "현우의 얼굴과 오염된 의상, 찰리의 고릴라형 체형과 흰 마스크, 담요 모두 참고 이미지와 정확히 일치함.",
        "hard_violations": [
         "지지대 없이 수평으로 공중에 완전히 떠 있는 현우의 자세 (물리 법칙 위반)",
         "충돌 전(before contact) 여유 공간을 요구한 프레이밍 지시와 달리 이미 벽에 발이 닿아 있음"
        ],
        "physics": "찰리는 바닥에 무릎을 대고 지탱 중이나, 현우는 도약이나 착지 없이 공중에 떠 있으며 몸의 하중을 지지하는 요소가 전혀 없음."
       },
       {
        "label": "B",
        "direction": "현우와 찰리 모두 바닥의 기울기를 따라 우측 하단으로 몸과 시선이 쏠려 있음.",
        "built_space": "좌측 상단 소파와 우측 중경의 벽, 그리고 두 인물 앞의 충돌 전 여유 공간까지 지시된 프레임 구도를 완벽히 충족함.",
        "entities": "현우의 얼굴 및 의상, 찰리의 체형과 샌드 베이지 장갑판 등 모든 설정이 참고 이미지와 일치함.",
        "hard_violations": [],
        "physics": "현우는 두 발로 바닥을 딛고 가방을 짚어 지탱하며, 찰리 역시 바닥에 서서 안정적으로 체중을 지지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 프레임 구도와 충돌 전 여유 공간을 잘 구현했고 인물들이 바닥에 지탱되어 안정적이나, 찰리의 백팩이 분리되고 포즈가 다소 정적인 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우가 공중에 떠 있어 물리적 지지대 규칙을 위반했으며, 충돌 전 여유 공간 지시를 어기고 이미 벽에 닿아 있어 프레이밍 지시를 실패했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 우측 벽을 향하고 있으며, 찰리는 정면 아래쪽 경사면을 향함.",
        "built_space": "좌측 상단 소파, 기울어진 바닥, 우측 벽 배치로 참고 이미지의 공간 및 조명 설정과 잘 부합함.",
        "entities": "현우의 얼굴과 오염된 의상, 찰리의 고릴라형 체형과 흰 마스크, 담요 모두 참고 이미지와 정확히 일치함.",
        "hard_violations": [
         "지지대 없이 수평으로 공중에 완전히 떠 있는 현우의 자세 (물리 법칙 위반)",
         "충돌 전(before contact) 여유 공간을 요구한 프레이밍 지시와 달리 이미 벽에 발이 닿아 있음"
        ],
        "physics": "찰리는 바닥에 무릎을 대고 지탱 중이나, 현우는 도약이나 착지 없이 공중에 떠 있으며 몸의 하중을 지지하는 요소가 전혀 없음."
       },
       {
        "label": "B",
        "direction": "현우와 찰리 모두 바닥의 기울기를 따라 우측 하단으로 몸과 시선이 쏠려 있음.",
        "built_space": "좌측 상단 소파와 우측 중경의 벽, 그리고 두 인물 앞의 충돌 전 여유 공간까지 지시된 프레임 구도를 완벽히 충족함.",
        "entities": "현우의 얼굴 및 의상, 찰리의 체형과 샌드 베이지 장갑판 등 모든 설정이 참고 이미지와 일치함.",
        "hard_violations": [],
        "physics": "현우는 두 발로 바닥을 딛고 가방을 짚어 지탱하며, 찰리 역시 바닥에 서서 안정적으로 체중을 지지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "좌상단 소파에서 우하단 벽으로 향하는 전신 구도와 신체 지지가 더 타당하지만, 찰리의 미끄러짐은 약하고 배낭이 등에서 바닥으로 옮겨졌다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리의 담요·배낭은 유지했지만 현우의 공중 자세를 설명할 지지가 불명확하고, 두 발이 이미 벽에 닿아 충돌 직전의 여유 공간 조건을 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 두 다리는 화면 우하단의 닫힌 문과 벽 쪽으로 뻗어 있고, 몸은 좌상단 소파에서 떨어져 나온 방향으로 기울어 있다. 현우의 시선은 목적지 벽보다 아래쪽 배낭과 바닥을 향한다. 찰리의 얼굴은 현우 쪽 아래를 향하며 몸도 우하단으로 기울지만, 발을 딛고 있어 거칠게 미끄러지는 운동은 현우보다 약하다.",
        "built_space": "좌상단에 회색 소파 한 개, 그 위쪽에 지도 한 장과 전기함·배관 묶음, 상단에 검은 화면 한 개와 천장등 한 개가 보인다. 오른쪽에는 타원형 창이 달린 닫힌 문 한 개와 여러 제어함이 있다. 두 인물의 전신이 소파와 오른쪽 벽 사이 바닥에 배치되어 있으며 발끝 앞에 좁은 간격이 남는다. 회색 직물과 베이지색 벽, 지도와 노출 배관은 이전 장면과 이어지지만, 추가로 드러난 설비와 목재 바닥은 참조의 좁은 화면만으로 동일성을 확인할 수 없다. 검증할 거울 반사는 없다.",
        "entities": "인물은 현우와 찰리 두 명뿐이다. 현우는 젊은 동아시아계 남성으로 검은 머리, 더러운 회갈색 민소매, 얼굴과 팔의 오염을 유지한다. 한국계 미국인이라는 국적은 외형만으로 판별할 수 없다. 찰리는 긴 중장갑 팔, 짧은 다리, 샌드 베이지 장갑과 흰 마스크형 얼굴을 갖춰 참조와 대체로 맞으며 줄무늬 담요도 걸치고 있다. 여행 배낭은 찰리의 등에 유지되지 않고 현우의 손 아래 바닥에 놓여 있다. 선명하게 읽히는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우는 두 손을 바닥의 배낭에 대고 아래쪽 부츠를 바닥에 접촉시켜 기운 몸을 받친다. 다른 다리가 들려 있어도 몸 전체가 무지지 상태는 아니다. 찰리는 두 발로 바닥을 딛고 있으며 담요는 어깨와 몸통에 걸려 있다. 배낭은 바닥에 놓여 있다. 경사진 바닥에서 손과 발로 버티며 밀리는 자세는 가능하지만, 찰리의 높은 상체와 발 디딤은 수동적으로 내던져진 상태보다 버티는 동작에 가깝다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 벽을 바라보고 한 팔을 그쪽으로 뻗으며 두 부츠도 같은 벽을 향한다. 찰리의 얼굴 역시 오른쪽 현우와 벽 쪽을 향하고, 다리는 우하단으로 놓여 있다. 이동 목적지는 명료하지만 현우의 두 발이 이미 벽에 닿아 있어 지시된 충돌 전 순간은 아니다.",
        "built_space": "좌상단 배경에 긴 회색 소파 한 개, 뒤 벽에 지도 한 장과 전기함·배관 묶음이 보인다. 중앙 뒤쪽에는 문과 문틀이 겹쳐 보이는 출입구가 있고, 오른쪽 벽에는 작은 창 한 개, 긴 굽은 배관 한 개, 상단 조명 한 개가 있다. 찰리는 바닥에 앉듯 놓이고 현우는 그 옆 벽 앞에 떠 있는 높이로 배치된다. 참조의 벽 재질과 소파 직물은 유사하지만 바닥은 A와 달리 회색의 거친 판재처럼 보이며, 출입구의 닫힌 상태는 명료하지 않다. 불가능한 거울 반사는 보이지 않는다.",
        "entities": "현우와 찰리만 보인다. 현우의 젊은 동아시아계 남성 외형, 검은 머리, 더러운 민소매와 녹색 계열 바지는 참조에 대체로 부합한다. 찰리의 흰 마스크형 얼굴, 베이지 장갑, 육중한 긴 팔과 짧은 다리도 맞는다. 줄무늬 담요가 몸에 걸려 있고 여행 배낭은 등 뒤에 유지되어 이 부분은 A보다 충실하다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "현우의 골반과 몸통이 바닥 위에 떠 있는데 이를 받치는 손이나 좌면이 보이지 않는다. 두 발은 수직 벽에 닿아 있지만 몸을 매달아 지지할 접점은 없으며, 바닥을 타고 미끄러지는 동작에서 이 공중 자세로 이어지는 이륙이나 낙하도 드러나지 않는다."
        ],
        "physics": "찰리는 골반과 다리, 바닥에 댄 장갑 손으로 체중을 지지한다. 담요는 몸에 걸리고 배낭은 등에 붙어 있어 지지가 설명된다. 반면 현우는 한 팔을 허공으로 뻗고 다른 손의 지지점은 보이지 않으며, 골반과 몸통 아래에도 받침이 없다. 벽에 댄 부츠 외에는 지지가 확인되지 않아, 바닥을 따라 밀리는 자세라기보다 몸통이 공중에 고정된 모습으로 읽힌다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "좌상단 소파에서 우하단 벽으로 향하는 전신 구도와 신체 지지가 더 타당하지만, 찰리의 미끄러짐은 약하고 배낭이 등에서 바닥으로 옮겨졌다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 담요·배낭은 유지했지만 현우의 공중 자세를 설명할 지지가 불명확하고, 두 발이 이미 벽에 닿아 충돌 직전의 여유 공간 조건을 어긴다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 두 다리는 화면 우하단의 닫힌 문과 벽 쪽으로 뻗어 있고, 몸은 좌상단 소파에서 떨어져 나온 방향으로 기울어 있다. 현우의 시선은 목적지 벽보다 아래쪽 배낭과 바닥을 향한다. 찰리의 얼굴은 현우 쪽 아래를 향하며 몸도 우하단으로 기울지만, 발을 딛고 있어 거칠게 미끄러지는 운동은 현우보다 약하다.",
        "built_space": "좌상단에 회색 소파 한 개, 그 위쪽에 지도 한 장과 전기함·배관 묶음, 상단에 검은 화면 한 개와 천장등 한 개가 보인다. 오른쪽에는 타원형 창이 달린 닫힌 문 한 개와 여러 제어함이 있다. 두 인물의 전신이 소파와 오른쪽 벽 사이 바닥에 배치되어 있으며 발끝 앞에 좁은 간격이 남는다. 회색 직물과 베이지색 벽, 지도와 노출 배관은 이전 장면과 이어지지만, 추가로 드러난 설비와 목재 바닥은 참조의 좁은 화면만으로 동일성을 확인할 수 없다. 검증할 거울 반사는 없다.",
        "entities": "인물은 현우와 찰리 두 명뿐이다. 현우는 젊은 동아시아계 남성으로 검은 머리, 더러운 회갈색 민소매, 얼굴과 팔의 오염을 유지한다. 한국계 미국인이라는 국적은 외형만으로 판별할 수 없다. 찰리는 긴 중장갑 팔, 짧은 다리, 샌드 베이지 장갑과 흰 마스크형 얼굴을 갖춰 참조와 대체로 맞으며 줄무늬 담요도 걸치고 있다. 여행 배낭은 찰리의 등에 유지되지 않고 현우의 손 아래 바닥에 놓여 있다. 선명하게 읽히는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우는 두 손을 바닥의 배낭에 대고 아래쪽 부츠를 바닥에 접촉시켜 기운 몸을 받친다. 다른 다리가 들려 있어도 몸 전체가 무지지 상태는 아니다. 찰리는 두 발로 바닥을 딛고 있으며 담요는 어깨와 몸통에 걸려 있다. 배낭은 바닥에 놓여 있다. 경사진 바닥에서 손과 발로 버티며 밀리는 자세는 가능하지만, 찰리의 높은 상체와 발 디딤은 수동적으로 내던져진 상태보다 버티는 동작에 가깝다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 벽을 바라보고 한 팔을 그쪽으로 뻗으며 두 부츠도 같은 벽을 향한다. 찰리의 얼굴 역시 오른쪽 현우와 벽 쪽을 향하고, 다리는 우하단으로 놓여 있다. 이동 목적지는 명료하지만 현우의 두 발이 이미 벽에 닿아 있어 지시된 충돌 전 순간은 아니다.",
        "built_space": "좌상단 배경에 긴 회색 소파 한 개, 뒤 벽에 지도 한 장과 전기함·배관 묶음이 보인다. 중앙 뒤쪽에는 문과 문틀이 겹쳐 보이는 출입구가 있고, 오른쪽 벽에는 작은 창 한 개, 긴 굽은 배관 한 개, 상단 조명 한 개가 있다. 찰리는 바닥에 앉듯 놓이고 현우는 그 옆 벽 앞에 떠 있는 높이로 배치된다. 참조의 벽 재질과 소파 직물은 유사하지만 바닥은 A와 달리 회색의 거친 판재처럼 보이며, 출입구의 닫힌 상태는 명료하지 않다. 불가능한 거울 반사는 보이지 않는다.",
        "entities": "현우와 찰리만 보인다. 현우의 젊은 동아시아계 남성 외형, 검은 머리, 더러운 민소매와 녹색 계열 바지는 참조에 대체로 부합한다. 찰리의 흰 마스크형 얼굴, 베이지 장갑, 육중한 긴 팔과 짧은 다리도 맞는다. 줄무늬 담요가 몸에 걸려 있고 여행 배낭은 등 뒤에 유지되어 이 부분은 A보다 충실하다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "현우의 골반과 몸통이 바닥 위에 떠 있는데 이를 받치는 손이나 좌면이 보이지 않는다. 두 발은 수직 벽에 닿아 있지만 몸을 매달아 지지할 접점은 없으며, 바닥을 타고 미끄러지는 동작에서 이 공중 자세로 이어지는 이륙이나 낙하도 드러나지 않는다."
        ],
        "physics": "찰리는 골반과 다리, 바닥에 댄 장갑 손으로 체중을 지지한다. 담요는 몸에 걸리고 배낭은 등에 붙어 있어 지지가 설명된다. 반면 현우는 한 팔을 허공으로 뻗고 다른 손의 지지점은 보이지 않으며, 골반과 몸통 아래에도 받침이 없다. 벽에 댄 부츠 외에는 지지가 확인되지 않아, 바닥을 따라 밀리는 자세라기보다 몸통이 공중에 고정된 모습으로 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.857,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.607,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지지대 없이 수평으로 공중에 완전히 떠 있는 현우의 자세 (물리 법칙 위반)",
     "[gemini-pro] 충돌 전(before contact) 여유 공간을 요구한 프레이밍 지시와 달리 이미 벽에 발이 닿아 있음",
     "[gpt-high] 현우의 골반과 몸통이 바닥 위에 떠 있는데 이를 받치는 손이나 좌면이 보이지 않는다. 두 발은 수직 벽에 닿아 있지만 몸을 매달아 지지할 접점은 없으며, 바닥을 타고 미끄러지는 동작에서 이 공중 자세로 이어지는 이륙이나 낙하도 드러나지 않는다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 607
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 프레임 구도와 충돌 전 여유 공간을 잘 구현했고 인물들이 바닥에 지탱되어 안정적이나, 찰리의 백팩이 분리되고 포즈가 다소 정적인 점이 아쉽습니다."
   },
   {
    "label": "A",
    "score": 607,
    "verdict_ko": "현우가 공중에 떠 있어 물리적 지지대 규칙을 위반했으며, 충돌 전 여유 공간 지시를 어기고 이미 벽에 닿아 있어 프레이밍 지시를 실패했습니다.  ★위반: [gemini-pro] 지지대 없이 수평으로 공중에 완전히 떠 있는 현우의 자세 (물리 법칙 위반) / [gemini-pro] 충돌 전(before contact) 여유 공간을 요구한 프레이밍 지시와 달리 이미 벽에 발이 닿아 있음 / [gpt-high] 현우의 골반과 몸통이 바닥 위에 떠 있는데 이를 받치는 손이나 좌면이 보이지 않는다. 두 발은 수직 벽에 닿아 있지만 몸을 매달아 지지할 접점은 없으며, 바닥을 타고 미끄러지는 동작에서 이 공중 자세로 이어지는 이륙이나 낙하도 드러나지 않는다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh12_sel.png",
    "asset_id": "2a6846f3-0dca-4a76-8163-8cbbf5be88db",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-e276-7167-9e79-483de213bb0b",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S79sh12"
  }
 },
 "S79sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:48:15.599992+00:00",
  "fingerprint": "ab4def78c3f3557a07017dec9e6455b99b36609d2396938de5d30ae98c7270b9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S79sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S79sh14_sel.png",
  "source_sha256": "3f8fe0c470c84c91f67dd957238f44de550228571b6e0f987ef3c80d08c5aaf0",
  "file": "S79sh14_cine.png",
  "staged_sha256": "84ca59505a378e684678a56daa5969a5e6d2c5ff08916eb96bd17f9277fa2db8",
  "latency_ms": 11092
 },
 "S80sh9::signage": {
  "fp": "6fb1eaa78cf58749",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::a955c71a887aa880": {
  "subjects": [],
  "subject_text": "크리스의 선박 갑판\n바다에 노출된 작은 선박의 낡은 갑판 공간. 바닥 위로 지지 기둥과 짐이 배치되고 실내로 연결되는 출입구가 있다.",
  "identity": "canonical",
  "scope_id": "L261",
  "scope_role": "location_exterior",
  "scope_sha": "2edff8e6cc4e90f6"
 },
 "S80sh9::bgfirst_bg": {
  "input_fingerprint": "42e829411c948878",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9__bgfirst_bg.png",
  "asset_id": "308f16b6-d6a7-43a8-94fa-ab3e5aa531be",
  "input_asset_ids": [
   "7e2e9a21-fd2a-469c-b057-4ecfe185964a",
   "a7ba2d1e-a8a6-411b-a6de-e6dbb6a4b132"
  ]
 },
 "S80sh9": {
  "input_fingerprint": "814700c7f13f6749",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Torrential rain, typhoon winds and huge waves batter the ship at night. Charlie's already damaged body is rain-soaked, with one hand braced around a shipboard post and the other extended over the edge. 현우: He is soaked and suspended at the ship's edge with one arm stretched upward, his earlier injuries still present.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Torrential rain, typhoon winds and huge waves batter the ship at night. Charlie's already damaged body is rain-soaked, with one hand braced around a shipboard post and the other extended over the edge. 현우: He is soaked and suspended at the ship's edge with one arm stretched upward, his earlier injuries still present.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Torrential rain, typhoon winds and huge waves batter the ship at night. Charlie's already damaged body is rain-soaked, with one hand braced around a shipboard post and the other extended over the edge. 현우: He is soaked and suspended at the ship's edge with one arm stretched upward, his earlier injuries still present.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9__bgfirst_bg.png",
     "asset_id": "308f16b6-d6a7-43a8-94fa-ab3e5aa531be",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S80sh9.png",
     "asset_id": "7e2e9a21-fd2a-469c-b057-4ecfe185964a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B01.png",
     "asset_id": "a7ba2d1e-a8a6-411b-a6de-e6dbb6a4b132",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "로봇 팔이 오른쪽을 향해 뻗어 인간의 팔뚝을 단단히 붙잡고 있음.",
    "built_space": "야간 선박 갑판. 좌측 상단에 기둥이 보이며, 대각선 난간이 갑판과 우측의 추락 공간을 명확히 분리함.",
    "entities": "찰리의 샌드 베이지색 금속 팔과 현우의 핏자국 묻은 회색 소매 팔이 지시사항과 일치함.",
    "hard_violations": [],
    "physics": "갑판에 고정된 로봇 팔이 난간 밖으로 추락하는 인간의 체중을 악력으로 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "로봇 팔이 오른쪽을 향해 뻗어 위로 향하는 인간의 손목을 잡고 있음.",
    "built_space": "좌측에 기둥이 있으나, 난간이 두 팔 뒤쪽 배경에 수평으로 위치하여 공간 분리에 실패함.",
    "entities": "찰리의 금속 팔과 현우의 회색 소매 의상이 일치함.",
    "hard_violations": [
     "[gemini-pro] 현우가 난간 밖 허공이 아닌 갑판 안쪽에 배치되어 공간적 연출 위반",
     "[gpt-high] 현우의 팔이 난간 바깥 아래의 추락 공간이 아니라 갑판 안쪽인 화면 왼쪽 아래에서 올라오도록 배치되어, 선외에 매달린 현우라는 필수 위치 관계를 뒤집는다."
    ],
    "physics": "두 팔이 모두 갑판 안쪽에 위치하여 허공으로 추락하는 물리적 상태와 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 카메라 앵글과 난간을 경계로 한 추락 공간 분리 연출을 매우 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물의 팔이 난간 안쪽 갑판에 위치하여 허공으로 떨어지는 핵심 상황 연출에 실패함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇 팔이 오른쪽을 향해 뻗어 인간의 팔뚝을 단단히 붙잡고 있음.",
        "built_space": "야간 선박 갑판. 좌측 상단에 기둥이 보이며, 대각선 난간이 갑판과 우측의 추락 공간을 명확히 분리함.",
        "entities": "찰리의 샌드 베이지색 금속 팔과 현우의 핏자국 묻은 회색 소매 팔이 지시사항과 일치함.",
        "hard_violations": [],
        "physics": "갑판에 고정된 로봇 팔이 난간 밖으로 추락하는 인간의 체중을 악력으로 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "로봇 팔이 오른쪽을 향해 뻗어 위로 향하는 인간의 손목을 잡고 있음.",
        "built_space": "좌측에 기둥이 있으나, 난간이 두 팔 뒤쪽 배경에 수평으로 위치하여 공간 분리에 실패함.",
        "entities": "찰리의 금속 팔과 현우의 회색 소매 의상이 일치함.",
        "hard_violations": [
         "현우가 난간 밖 허공이 아닌 갑판 안쪽에 배치되어 공간적 연출 위반"
        ],
        "physics": "두 팔이 모두 갑판 안쪽에 위치하여 허공으로 추락하는 물리적 상태와 맞지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 카메라 앵글과 난간을 경계로 한 추락 공간 분리 연출을 매우 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물의 팔이 난간 안쪽 갑판에 위치하여 허공으로 떨어지는 핵심 상황 연출에 실패함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "로봇 팔이 오른쪽을 향해 뻗어 인간의 팔뚝을 단단히 붙잡고 있음.",
        "built_space": "야간 선박 갑판. 좌측 상단에 기둥이 보이며, 대각선 난간이 갑판과 우측의 추락 공간을 명확히 분리함.",
        "entities": "찰리의 샌드 베이지색 금속 팔과 현우의 핏자국 묻은 회색 소매 팔이 지시사항과 일치함.",
        "hard_violations": [],
        "physics": "갑판에 고정된 로봇 팔이 난간 밖으로 추락하는 인간의 체중을 악력으로 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "로봇 팔이 오른쪽을 향해 뻗어 위로 향하는 인간의 손목을 잡고 있음.",
        "built_space": "좌측에 기둥이 있으나, 난간이 두 팔 뒤쪽 배경에 수평으로 위치하여 공간 분리에 실패함.",
        "entities": "찰리의 금속 팔과 현우의 회색 소매 의상이 일치함.",
        "hard_violations": [
         "현우가 난간 밖 허공이 아닌 갑판 안쪽에 배치되어 공간적 연출 위반"
        ],
        "physics": "두 팔이 모두 갑판 안쪽에 위치하여 허공으로 추락하는 물리적 상태와 맞지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "손목을 꽉 감싼 금속 손의 클로즈업은 명확하지만, 현우의 팔이 선외 추락 공간이 아닌 갑판 쪽에서 올라와 핵심 공간 배치를 어긴다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "갑판 쪽 찰리가 선외 현우의 손목을 붙드는 방향과 지지 관계가 맞고 상처도 유지되지만, 손의 비중이 작고 지지 기둥과 갑판이 지시보다 많이 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔은 왼쪽에서 오른쪽으로 뻗어 현우의 손목을 감싼다. 현우의 팔은 화면 왼쪽 아래 갑판 쪽에서 오른쪽 위로 올라가며, 손끝은 위쪽을 향한다. 붙잡는 대상은 정확히 손목이지만, 선외 아래로 떨어지는 사람에게 뻗은 구조 동작으로는 방향이 맞지 않는다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "녹슨 흰 지지 기둥 하나가 왼쪽 위에서 아래까지 이어지고, 밧줄과 고정 철물이 붙어 있다. 선측의 상부 난간과 하부 테두리가 팔 아래에서 오른쪽 아래로 비스듬히 지나간다. 기둥과 재질은 장소 참조에 부합하지만, 현우의 팔이 들어오는 왼쪽 아래는 젖은 갑판이 보이는 안쪽이다. 기둥도 작은 배경 조각보다 크게 노출된다.",
        "entities": "보이는 인물 부위는 현우의 맨손·팔·젖은 회색 소매와 찰리의 금속 손·팔뿐이다. 현우의 얼굴이 없어 나이와 한국계 미국인 정체성을 직접 확인할 수는 없으며, 팔의 체격과 회색 의상은 참조와 양립한다. 찰리의 모래색 각진 장갑판, 검은 관절, 마모 흔적은 참조에 가깝다. 현우의 기존 부상은 뚜렷하지 않다. 밤바다와 폭우가 보이고 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우의 팔이 난간 바깥 아래의 추락 공간이 아니라 갑판 안쪽인 화면 왼쪽 아래에서 올라오도록 배치되어, 선외에 매달린 현우라는 필수 위치 관계를 뒤집는다."
        ],
        "physics": "금속 손가락들이 손목 둘레에 실제로 접촉하며 감겨 있어 붙잡는 힘은 설득력 있다. 팔들은 각각 화면 밖으로 이어지며, 지지 없이 떠 있는 독립 물체는 없다. 찰리의 기둥을 잡은 다른 손과 두 인물의 몸은 화면 밖이므로 직접 검증할 수 없다. 문제는 지지 없는 부유가 아니라 현우의 팔이 갑판 쪽에서 올라오는 공간 배치다."
       },
       {
        "label": "B",
        "direction": "찰리의 팔은 왼쪽 위 갑판 쪽에서 오른쪽 아래 선측으로 뻗고, 현우의 팔은 오른쪽 선외에서 왼쪽 위의 붙잡힌 손목으로 이어진다. 현우의 손가락은 아래로 처져 있다. 금속 손이 향하는 대상은 현우의 손목이며, 추락하는 몸을 갑판 방향으로 붙드는 관계가 읽힌다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "왼쪽에 밧줄이 달린 녹슨 흰 지지 기둥 하나가 있고, 선측의 상부 난간과 하부 테두리가 팔 아래를 대각선으로 가른다. 가까운 난간에는 세 개의 세로 연결대가 구분된다. 젖은 갑판은 왼쪽, 파도와 현우의 팔이 이어지는 선외 공간은 오른쪽으로 나뉘어 장소 참조와 부합한다. 다만 기둥의 밑부분과 넓은 갑판까지 보여, 작은 기둥 일부만 남기라는 배경 지시보다 노출이 많다.",
        "entities": "찰리의 모래색 금속 팔과 관절, 왼쪽 가장자리의 회색 천이 참조와 일치한다. 현우는 젖은 회색 소매와 맨손·팔로 표현되며, 팔의 찰과상과 찢어지고 피 묻은 소매가 기존 부상을 보여 준다. 얼굴이 없어 정확한 나이·민족성·얼굴 동일성은 확인할 수 없다. 다른 사람은 없고, 밤의 폭우와 거친 파도, 젖은 금속과 천의 질감이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 손바닥과 굽힌 금속 손가락이 현우의 손목 부근에 접촉해 지지하며, 현우의 팔은 선외 화면 밖의 몸 쪽으로 이어진다. 일부 금속 손가락이 벌어져 있어 A보다 꽉 움켜쥔 인상은 약하지만, 손목을 붙드는 접촉 자체는 보인다. 찰리의 팔은 왼쪽 몸 쪽으로 연결되고 지지 기둥도 그 뒤에 있다. 기둥을 잡은 다른 손은 프레임 밖이며, 근거 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "손목을 꽉 감싼 금속 손의 클로즈업은 명확하지만, 현우의 팔이 선외 추락 공간이 아닌 갑판 쪽에서 올라와 핵심 공간 배치를 어긴다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "갑판 쪽 찰리가 선외 현우의 손목을 붙드는 방향과 지지 관계가 맞고 상처도 유지되지만, 손의 비중이 작고 지지 기둥과 갑판이 지시보다 많이 보인다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 팔은 왼쪽에서 오른쪽으로 뻗어 현우의 손목을 감싼다. 현우의 팔은 화면 왼쪽 아래 갑판 쪽에서 오른쪽 위로 올라가며, 손끝은 위쪽을 향한다. 붙잡는 대상은 정확히 손목이지만, 선외 아래로 떨어지는 사람에게 뻗은 구조 동작으로는 방향이 맞지 않는다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "녹슨 흰 지지 기둥 하나가 왼쪽 위에서 아래까지 이어지고, 밧줄과 고정 철물이 붙어 있다. 선측의 상부 난간과 하부 테두리가 팔 아래에서 오른쪽 아래로 비스듬히 지나간다. 기둥과 재질은 장소 참조에 부합하지만, 현우의 팔이 들어오는 왼쪽 아래는 젖은 갑판이 보이는 안쪽이다. 기둥도 작은 배경 조각보다 크게 노출된다.",
        "entities": "보이는 인물 부위는 현우의 맨손·팔·젖은 회색 소매와 찰리의 금속 손·팔뿐이다. 현우의 얼굴이 없어 나이와 한국계 미국인 정체성을 직접 확인할 수는 없으며, 팔의 체격과 회색 의상은 참조와 양립한다. 찰리의 모래색 각진 장갑판, 검은 관절, 마모 흔적은 참조에 가깝다. 현우의 기존 부상은 뚜렷하지 않다. 밤바다와 폭우가 보이고 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우의 팔이 난간 바깥 아래의 추락 공간이 아니라 갑판 안쪽인 화면 왼쪽 아래에서 올라오도록 배치되어, 선외에 매달린 현우라는 필수 위치 관계를 뒤집는다."
        ],
        "physics": "금속 손가락들이 손목 둘레에 실제로 접촉하며 감겨 있어 붙잡는 힘은 설득력 있다. 팔들은 각각 화면 밖으로 이어지며, 지지 없이 떠 있는 독립 물체는 없다. 찰리의 기둥을 잡은 다른 손과 두 인물의 몸은 화면 밖이므로 직접 검증할 수 없다. 문제는 지지 없는 부유가 아니라 현우의 팔이 갑판 쪽에서 올라오는 공간 배치다."
       },
       {
        "label": "A",
        "direction": "찰리의 팔은 왼쪽 위 갑판 쪽에서 오른쪽 아래 선측으로 뻗고, 현우의 팔은 오른쪽 선외에서 왼쪽 위의 붙잡힌 손목으로 이어진다. 현우의 손가락은 아래로 처져 있다. 금속 손이 향하는 대상은 현우의 손목이며, 추락하는 몸을 갑판 방향으로 붙드는 관계가 읽힌다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "왼쪽에 밧줄이 달린 녹슨 흰 지지 기둥 하나가 있고, 선측의 상부 난간과 하부 테두리가 팔 아래를 대각선으로 가른다. 가까운 난간에는 세 개의 세로 연결대가 구분된다. 젖은 갑판은 왼쪽, 파도와 현우의 팔이 이어지는 선외 공간은 오른쪽으로 나뉘어 장소 참조와 부합한다. 다만 기둥의 밑부분과 넓은 갑판까지 보여, 작은 기둥 일부만 남기라는 배경 지시보다 노출이 많다.",
        "entities": "찰리의 모래색 금속 팔과 관절, 왼쪽 가장자리의 회색 천이 참조와 일치한다. 현우는 젖은 회색 소매와 맨손·팔로 표현되며, 팔의 찰과상과 찢어지고 피 묻은 소매가 기존 부상을 보여 준다. 얼굴이 없어 정확한 나이·민족성·얼굴 동일성은 확인할 수 없다. 다른 사람은 없고, 밤의 폭우와 거친 파도, 젖은 금속과 천의 질감이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 손바닥과 굽힌 금속 손가락이 현우의 손목 부근에 접촉해 지지하며, 현우의 팔은 선외 화면 밖의 몸 쪽으로 이어진다. 일부 금속 손가락이 벌어져 있어 A보다 꽉 움켜쥔 인상은 약하지만, 손목을 붙드는 접촉 자체는 보인다. 찰리의 팔은 왼쪽 몸 쪽으로 연결되고 지지 기둥도 그 뒤에 있다. 기둥을 잡은 다른 손은 프레임 밖이며, 근거 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 현우가 난간 밖 허공이 아닌 갑판 안쪽에 배치되어 공간적 연출 위반",
     "[gpt-high] 현우의 팔이 난간 바깥 아래의 추락 공간이 아니라 갑판 안쪽인 화면 왼쪽 아래에서 올라오도록 배치되어, 선외에 매달린 현우라는 필수 위치 관계를 뒤집는다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 카메라 앵글과 난간을 경계로 한 추락 공간 분리 연출을 매우 정확히 구현함."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "인물의 팔이 난간 안쪽 갑판에 위치하여 허공으로 떨어지는 핵심 상황 연출에 실패함.  ★위반: [gemini-pro] 현우가 난간 밖 허공이 아닌 갑판 안쪽에 배치되어 공간적 연출 위반 / [gpt-high] 현우의 팔이 난간 바깥 아래의 추락 공간이 아니라 갑판 안쪽인 화면 왼쪽 아래에서 올라오도록 배치되어, 선외에 매달린 현우라는 필수 위치 관계를 뒤집는다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B01.png",
    "asset_id": "a7ba2d1e-a8a6-411b-a6de-e6dbb6a4b132",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-e429-7a80-953e-55e5ada9e777",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9__bgfirst_bg.png",
   "bg_asset_id": "308f16b6-d6a7-43a8-94fa-ab3e5aa531be",
   "bg_record_key": "S80sh9::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S80sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:49:41.885079+00:00",
  "fingerprint": "8f5c2014a3666f6c7e60a0245726c24d0ff55590f44dfa0685b47a2e9c1ac0d1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S80sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S80sh9_sel.png",
  "source_sha256": "a461b467bd990a4ad0c8d7fb55b29731f45e8811fd33099ce9834ff2c534d0f5",
  "file": "S80sh9_cine.png",
  "staged_sha256": "9f1ae917eb3add68b20c7d6129dc4719645eec31cc41c11d7c58ac98043ffb1e",
  "latency_ms": 9102
 },
 "S80sh15::signage": {
  "fp": "4420b6e6498fab5b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S80sh15": {
  "input_fingerprint": "894b06b3a046ab7a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 파도의 엄청난 힘에 의해 찰리의 금속 손아귀에서 거칠게 뜯겨져 나간 현우의 손 클로즈업.\n\nLOCATION (lock): At the open deck edge of the boat during the nighttime storm, where a breaking wave tears the two apart. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside the fall as the wave breaks the grip) — Retains the same diagonal deck-side view as the earlier hand close-up; used as Provides an unchanged reference for the separation and impending downward tilt; Wave water (Striking the ship and pulling the pair apart); used as Partially interrupts the surrounding space without concealing the gap between the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established nighttime storm illumination and rain-obscured contrast through the release, without adding a flash or new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ship is being struck by another immense wave amid the ongoing typhoon and torrential rain. Charlie's wet, damaged body remains at the ship's edge as the outstretched grip breaks. 현우: He is fully soaked and being swept away from the ship, no longer held by his extended hand.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 파도의 엄청난 힘에 의해 찰리의 금속 손아귀에서 거칠게 뜯겨져 나간 현우의 손 클로즈업.\n\nLOCATION (lock): At the open deck edge of the boat during the nighttime storm, where a breaking wave tears the two apart. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside the fall as the wave breaks the grip) — Retains the same diagonal deck-side view as the earlier hand close-up; used as Provides an unchanged reference for the separation and impending downward tilt; Wave water (Striking the ship and pulling the pair apart); used as Partially interrupts the surrounding space without concealing the gap between the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established nighttime storm illumination and rain-obscured contrast through the release, without adding a flash or new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ship is being struck by another immense wave amid the ongoing typhoon and torrential rain. Charlie's wet, damaged body remains at the ship's edge as the outstretched grip breaks. 현우: He is fully soaked and being swept away from the ship, no longer held by his extended hand.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 파도의 엄청난 힘에 의해 찰리의 금속 손아귀에서 거칠게 뜯겨져 나간 현우의 손 클로즈업.\n\nLOCATION (lock): At the open deck edge of the boat during the nighttime storm, where a breaking wave tears the two apart. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside the fall as the wave breaks the grip) — Retains the same diagonal deck-side view as the earlier hand close-up; used as Provides an unchanged reference for the separation and impending downward tilt; Wave water (Striking the ship and pulling the pair apart); used as Partially interrupts the surrounding space without concealing the gap between the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established nighttime storm illumination and rain-obscured contrast through the release, without adding a flash or new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ship is being struck by another immense wave amid the ongoing typhoon and torrential rain. Charlie's wet, damaged body remains at the ship's edge as the outstretched grip breaks. 현우: He is fully soaked and being swept away from the ship, no longer held by his extended hand.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "로봇의 팔은 오른쪽으로 뻗어 있고, 사람의 손은 파도에 휩쓸려 프레임 오른쪽 아래로 멀어지고 있음.",
    "built_space": "비 내리는 밤의 선박 갑판 가장자리. 금속 난간, 젖은 바닥, 뒷배경의 로프 등 레퍼런스 샷의 공간과 동일함.",
    "entities": "베이지색 장갑판의 로봇 팔(찰리)과 젖은 회색 셔츠를 입고 피가 묻은 인간의 팔(현우)이 정확히 일치함.",
    "hard_violations": [],
    "physics": "현우의 팔이 파도에 밀려 뒤로 젖혀지며 추락하는 힘이 잘 느껴지며, 두 손은 허공에서 완전히 분리되어 있음."
   },
   {
    "label": "B",
    "direction": "로봇의 팔과 사람의 손이 서로를 향해 뻗어 있으며 맞닿아 있음.",
    "built_space": "비 내리는 밤의 선박 갑판 가장자리. 금속 난간과 젖은 바닥, 덮개 등의 요소가 레퍼런스와 일치함.",
    "entities": "베이지색 장갑판의 로봇 팔과 회색 셔츠를 입고 상처 난 인간의 팔이 나타남.",
    "hard_violations": [
     "[gemini-pro] 두 손이 완전히 분리되어 허공에 있어야 한다는 지시를 위반하고 손가락이 닿아 있음."
    ],
    "physics": "강력한 파도에 의해 분리되는 순간임에도 불구하고 사람의 손이 여전히 로봇 손가락에 걸쳐 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "파도에 의해 두 손이 거칠게 뜯겨져 나간 분리 상태를 프롬프트의 지시대로 정확하게 연출했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손이 이미 뜯겨져 나가 분리되었다는 핵심 지시와 달리, 두 손가락이 여전히 맞닿아 있어 연출 의도에서 벗어났습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 팔은 오른쪽으로 뻗어 있고, 사람의 손은 파도에 휩쓸려 프레임 오른쪽 아래로 멀어지고 있음.",
        "built_space": "비 내리는 밤의 선박 갑판 가장자리. 금속 난간, 젖은 바닥, 뒷배경의 로프 등 레퍼런스 샷의 공간과 동일함.",
        "entities": "베이지색 장갑판의 로봇 팔(찰리)과 젖은 회색 셔츠를 입고 피가 묻은 인간의 팔(현우)이 정확히 일치함.",
        "hard_violations": [],
        "physics": "현우의 팔이 파도에 밀려 뒤로 젖혀지며 추락하는 힘이 잘 느껴지며, 두 손은 허공에서 완전히 분리되어 있음."
       },
       {
        "label": "B",
        "direction": "로봇의 팔과 사람의 손이 서로를 향해 뻗어 있으며 맞닿아 있음.",
        "built_space": "비 내리는 밤의 선박 갑판 가장자리. 금속 난간과 젖은 바닥, 덮개 등의 요소가 레퍼런스와 일치함.",
        "entities": "베이지색 장갑판의 로봇 팔과 회색 셔츠를 입고 상처 난 인간의 팔이 나타남.",
        "hard_violations": [
         "두 손이 완전히 분리되어 허공에 있어야 한다는 지시를 위반하고 손가락이 닿아 있음."
        ],
        "physics": "강력한 파도에 의해 분리되는 순간임에도 불구하고 사람의 손이 여전히 로봇 손가락에 걸쳐 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "파도에 의해 두 손이 거칠게 뜯겨져 나간 분리 상태를 프롬프트의 지시대로 정확하게 연출했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손이 이미 뜯겨져 나가 분리되었다는 핵심 지시와 달리, 두 손가락이 여전히 맞닿아 있어 연출 의도에서 벗어났습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 팔은 오른쪽으로 뻗어 있고, 사람의 손은 파도에 휩쓸려 프레임 오른쪽 아래로 멀어지고 있음.",
        "built_space": "비 내리는 밤의 선박 갑판 가장자리. 금속 난간, 젖은 바닥, 뒷배경의 로프 등 레퍼런스 샷의 공간과 동일함.",
        "entities": "베이지색 장갑판의 로봇 팔(찰리)과 젖은 회색 셔츠를 입고 피가 묻은 인간의 팔(현우)이 정확히 일치함.",
        "hard_violations": [],
        "physics": "현우의 팔이 파도에 밀려 뒤로 젖혀지며 추락하는 힘이 잘 느껴지며, 두 손은 허공에서 완전히 분리되어 있음."
       },
       {
        "label": "B",
        "direction": "로봇의 팔과 사람의 손이 서로를 향해 뻗어 있으며 맞닿아 있음.",
        "built_space": "비 내리는 밤의 선박 갑판 가장자리. 금속 난간과 젖은 바닥, 덮개 등의 요소가 레퍼런스와 일치함.",
        "entities": "베이지색 장갑판의 로봇 팔과 회색 셔츠를 입고 상처 난 인간의 팔이 나타남.",
        "hard_violations": [
         "두 손이 완전히 분리되어 허공에 있어야 한다는 지시를 위반하고 손가락이 닿아 있음."
        ],
        "physics": "강력한 파도에 의해 분리되는 순간임에도 불구하고 사람의 손이 여전히 로봇 손가락에 걸쳐 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "장소와 외형은 잘 이어지지만 두 손이 여전히 겹쳐 보여, 손아귀에서 거칠게 빠져나와 간격이 생긴 결정적 순간을 구현하지 못했습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "대각선 선측 구도와 야간 폭풍을 유지하면서, 파도에 휩쓸리는 현우의 손과 열린 금속 손 사이의 명확한 간격으로 놓치는 순간을 정확히 보여줍니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 금속 손은 왼쪽에서 오른쪽 현우의 손을 향하고, 현우의 손가락은 왼쪽 금속 손 쪽으로 뻗어 있습니다. 두 손의 윤곽이 겹쳐 분리된 틈이 확인되지 않으며, 현우가 오른쪽 선외로 끌려가는 동작보다 다시 손을 잡으려는 모양에 가깝습니다. 얼굴과 시선은 프레임 밖입니다.",
        "built_space": "왼쪽에 젖은 갑판과 굵은 수직 기둥 하나, 그 옆에 감긴 밧줄이 있고, 녹슨 선측 벽과 상부 관형 난간 한 줄이 오른쪽 아래로 이어집니다. 난간의 수직 지지대는 물보라에 일부 가린 것을 포함해 세 개가 보입니다. 찰리의 팔은 갑판 쪽에서, 현우의 팔은 난간 바깥 오른쪽에서 들어와 이전 장소의 배치를 유지합니다. 불가능한 반사나 중복 설비는 보이지 않습니다.",
        "entities": "보이는 인물 부분은 찰리의 기계 팔과 손, 현우의 맨손과 팔뿐입니다. 찰리의 긁힌 샌드 베이지 장갑판과 검은 관절, 현우의 젖고 찢어진 회색 소매 및 상처 난 피부가 이전 사진과 맞습니다. 얼굴이 없어 현우의 정확한 나이와 한국계 미국인 정체성은 판별할 수 없지만, 손과 팔은 젊은 남성 설정에 어긋나지 않습니다. 바다와 파도, 비가 있으며 추가 인물이나 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "양쪽 손은 각각 화면 밖으로 이어지는 팔에 연결되어 있어 따로 떠 있지 않습니다. 파도가 선측을 넘어 갑판으로 쏟아지는 흐름도 자연스럽습니다. 다만 현우의 손바닥이 금속 손가락에 닿거나 겹쳐 보여, 파도의 힘으로 접촉이 완전히 끊어진 상태는 입증되지 않습니다. 프레임 밖 몸의 지지는 판단할 수 없습니다."
       },
       {
        "label": "B",
        "direction": "찰리의 열린 금속 손가락은 오른쪽 현우의 손을 향하고, 현우의 굽힌 손가락은 왼쪽 찰리 쪽을 향합니다. 현우의 팔은 오른쪽 선외로 이어지며 두 손 사이에 뚜렷한 빈틈이 있어 서로 멀어지는 관계가 읽힙니다. 파도는 선외에서 난간을 넘어 왼쪽 갑판으로 부서지고, 물보라는 손 사이 간격을 가리지 않습니다. 얼굴과 시선은 보이지 않습니다.",
        "built_space": "왼쪽 갑판에 굵은 수직 기둥 하나와 감긴 밧줄이 있고, 녹슨 선측 벽 위의 관형 난간 한 줄이 중앙에서 오른쪽 아래로 뻗습니다. 수직 난간 지지대 세 개가 뚜렷하게 보입니다. 찰리는 갑판 안쪽, 현우는 난간 바깥쪽이라는 팔의 진입 위치가 유지됩니다. 이전 사진의 젖은 금속 재질과 대각선 선측 시점이 이어지며, 중복 설비나 불가능한 반사는 없습니다.",
        "entities": "찰리의 육중한 샌드 베이지 금속 팔과 관절식 손, 현우의 상처 난 맨손과 젖고 찢어진 회색 소매가 보입니다. 모두 이전 사진에서 보이는 해당 인물의 재질과 의상에 부합합니다. 현우의 얼굴과 머리는 잘려 있어 민족적 정체성과 정확한 연령은 직접 확인할 수 없고, 보이는 팔은 젊은 남성 설정과 양립합니다. 큰 파도와 폭우가 있으며 추가 인물이나 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "두 손 모두 손목과 팔에 정상적으로 연결되어 있습니다. 찰리의 손은 펼쳐져 접촉을 잃었고, 현우의 손은 팔을 따라 선외로 물러난 위치여서 파도에 의해 손아귀에서 빠져나온 동작이 가능합니다. 파도가 난간을 넘어 갑판으로 떨어지는 중력 방향도 자연스럽습니다. 전신은 보이지 않으므로 공중에 떠 있는 몸이나 지지 없는 물체로 판단할 요소가 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "장소와 외형은 잘 이어지지만 두 손이 여전히 겹쳐 보여, 손아귀에서 거칠게 빠져나와 간격이 생긴 결정적 순간을 구현하지 못했습니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "대각선 선측 구도와 야간 폭풍을 유지하면서, 파도에 휩쓸리는 현우의 손과 열린 금속 손 사이의 명확한 간격으로 놓치는 순간을 정확히 보여줍니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 금속 손은 왼쪽에서 오른쪽 현우의 손을 향하고, 현우의 손가락은 왼쪽 금속 손 쪽으로 뻗어 있습니다. 두 손의 윤곽이 겹쳐 분리된 틈이 확인되지 않으며, 현우가 오른쪽 선외로 끌려가는 동작보다 다시 손을 잡으려는 모양에 가깝습니다. 얼굴과 시선은 프레임 밖입니다.",
        "built_space": "왼쪽에 젖은 갑판과 굵은 수직 기둥 하나, 그 옆에 감긴 밧줄이 있고, 녹슨 선측 벽과 상부 관형 난간 한 줄이 오른쪽 아래로 이어집니다. 난간의 수직 지지대는 물보라에 일부 가린 것을 포함해 세 개가 보입니다. 찰리의 팔은 갑판 쪽에서, 현우의 팔은 난간 바깥 오른쪽에서 들어와 이전 장소의 배치를 유지합니다. 불가능한 반사나 중복 설비는 보이지 않습니다.",
        "entities": "보이는 인물 부분은 찰리의 기계 팔과 손, 현우의 맨손과 팔뿐입니다. 찰리의 긁힌 샌드 베이지 장갑판과 검은 관절, 현우의 젖고 찢어진 회색 소매 및 상처 난 피부가 이전 사진과 맞습니다. 얼굴이 없어 현우의 정확한 나이와 한국계 미국인 정체성은 판별할 수 없지만, 손과 팔은 젊은 남성 설정에 어긋나지 않습니다. 바다와 파도, 비가 있으며 추가 인물이나 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "양쪽 손은 각각 화면 밖으로 이어지는 팔에 연결되어 있어 따로 떠 있지 않습니다. 파도가 선측을 넘어 갑판으로 쏟아지는 흐름도 자연스럽습니다. 다만 현우의 손바닥이 금속 손가락에 닿거나 겹쳐 보여, 파도의 힘으로 접촉이 완전히 끊어진 상태는 입증되지 않습니다. 프레임 밖 몸의 지지는 판단할 수 없습니다."
       },
       {
        "label": "A",
        "direction": "찰리의 열린 금속 손가락은 오른쪽 현우의 손을 향하고, 현우의 굽힌 손가락은 왼쪽 찰리 쪽을 향합니다. 현우의 팔은 오른쪽 선외로 이어지며 두 손 사이에 뚜렷한 빈틈이 있어 서로 멀어지는 관계가 읽힙니다. 파도는 선외에서 난간을 넘어 왼쪽 갑판으로 부서지고, 물보라는 손 사이 간격을 가리지 않습니다. 얼굴과 시선은 보이지 않습니다.",
        "built_space": "왼쪽 갑판에 굵은 수직 기둥 하나와 감긴 밧줄이 있고, 녹슨 선측 벽 위의 관형 난간 한 줄이 중앙에서 오른쪽 아래로 뻗습니다. 수직 난간 지지대 세 개가 뚜렷하게 보입니다. 찰리는 갑판 안쪽, 현우는 난간 바깥쪽이라는 팔의 진입 위치가 유지됩니다. 이전 사진의 젖은 금속 재질과 대각선 선측 시점이 이어지며, 중복 설비나 불가능한 반사는 없습니다.",
        "entities": "찰리의 육중한 샌드 베이지 금속 팔과 관절식 손, 현우의 상처 난 맨손과 젖고 찢어진 회색 소매가 보입니다. 모두 이전 사진에서 보이는 해당 인물의 재질과 의상에 부합합니다. 현우의 얼굴과 머리는 잘려 있어 민족적 정체성과 정확한 연령은 직접 확인할 수 없고, 보이는 팔은 젊은 남성 설정과 양립합니다. 큰 파도와 폭우가 있으며 추가 인물이나 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "두 손 모두 손목과 팔에 정상적으로 연결되어 있습니다. 찰리의 손은 펼쳐져 접촉을 잃었고, 현우의 손은 팔을 따라 선외로 물러난 위치여서 파도에 의해 손아귀에서 빠져나온 동작이 가능합니다. 파도가 난간을 넘어 갑판으로 떨어지는 중력 방향도 자연스럽습니다. 전신은 보이지 않으므로 공중에 떠 있는 몸이나 지지 없는 물체로 판단할 요소가 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.984
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.734
   },
   "violations": {
    "B": [
     "[gemini-pro] 두 손이 완전히 분리되어 허공에 있어야 한다는 지시를 위반하고 손가락이 닿아 있음."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 734
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "파도에 의해 두 손이 거칠게 뜯겨져 나간 분리 상태를 프롬프트의 지시대로 정확하게 연출했습니다."
   },
   {
    "label": "B",
    "score": 734,
    "verdict_ko": "손이 이미 뜯겨져 나가 분리되었다는 핵심 지시와 달리, 두 손가락이 여전히 맞닿아 있어 연출 의도에서 벗어났습니다.  ★위반: [gemini-pro] 두 손이 완전히 분리되어 허공에 있어야 한다는 지시를 위반하고 손가락이 닿아 있음."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9_sel.png",
    "asset_id": "6e0342c8-0da3-4a8c-8b19-c3217ac5881e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-e777-7878-92b7-0f94bc4f599a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S80sh9"
  }
 },
 "S80sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:50:42.579855+00:00",
  "fingerprint": "e82c3a4a815d71c133743c8be6f039b7bbf11c1c18eda9fa18a753b97a0a92de",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S80sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S80sh15_sel.png",
  "source_sha256": "0f7835892c6987d5e0dfe17612a174487caf34143f62886e52ede4148dfa4cfb",
  "file": "S80sh15_cine.png",
  "staged_sha256": "889be16c137e07c8525b10c4248fd70387a57f11682fb044afb8e1670a3dfb13",
  "latency_ms": 9096
 },
 "S80sh20::signage": {
  "fp": "4905fe22411e7df2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S80sh20::bgfirst_bg": {
  "input_fingerprint": "50ca4ed39ec1fd85",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh20__bgfirst_bg.png",
  "asset_id": "e9739767-c664-4b57-931d-69841d4b4382",
  "input_asset_ids": [
   "fee007fc-1173-489b-af5a-181096534991",
   "dac9a97a-00da-4766-aa32-41c0c8abe71c"
  ]
 },
 "S80sh20": {
  "input_fingerprint": "86b6c0f1b2c4df71",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is unconscious and sinking beneath the sea beside Charlie, his body limp and unsupported in the water. The source does not establish the direction of his head, the orientation of his torso, or the arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy shipboard cargo sinks through increasingly dark water. Charlie is submerged and descending with his preexisting body damage still present. 현우: He is unconscious and sinking deeper underwater with his body limp.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is unconscious and sinking beneath the sea beside Charlie, his body limp and unsupported in the water. The source does not establish the direction of his head, the orientation of his torso, or the arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy shipboard cargo sinks through increasingly dark water. Charlie is submerged and descending with his preexisting body damage still present. 현우: He is unconscious and sinking deeper underwater with his body limp.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is unconscious and sinking beneath the sea beside Charlie, his body limp and unsupported in the water. The source does not establish the direction of his head, the orientation of his torso, or the arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy shipboard cargo sinks through increasingly dark water. Charlie is submerged and descending with his preexisting body damage still present. 현우: He is unconscious and sinking deeper underwater with his body limp.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh20__bgfirst_bg.png",
     "asset_id": "e9739767-c664-4b57-931d-69841d4b4382",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S80sh20.png",
     "asset_id": "fee007fc-1173-489b-af5a-181096534991",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1293785>",
     "asset_id": "038f90ae-579f-46c4-8fd6-1c2c6d014b77",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B02.png",
     "asset_id": "dac9a97a-00da-4766-aa32-41c0c8abe71c",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1293785>",
     "asset_id": "038f90ae-579f-46c4-8fd6-1c2c6d014b77",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1120327>",
     "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 캐릭터가 수면의 빛을 등지고 아래로 가라앉고 있으며, 몸의 방향은 깊은 바닥을 향함.",
    "built_space": "위에서 빛이 스며들고 아래로 갈수록 칠흑같이 어두워지는 심해 공간이며, 명시된 가라앉는 선박 화물은 보이지 않음.",
    "entities": "현우는 레퍼런스와 다른 짧은 반팔 티셔츠 차림임. 찰리는 숄을 두르지 않았고 요구된 몸체의 파손 흔적도 묘사되지 않음.",
    "hard_violations": [],
    "physics": "두 사람 모두 물속에 떠서 가라앉고 있으며, 의식이 없는 상태에 맞게 팔다리가 저항 없이 축 늘어져 있음."
   },
   {
    "label": "B",
    "direction": "두 캐릭터가 수평에 가깝게 나란히 부유하며 깊은 곳으로 가라앉고 있음.",
    "built_space": "빛이 점차 사라지는 심해 배경이며, 두 캐릭터 주변으로 직육면체 형태의 화물 상자들이 함께 가라앉고 있음. 화면 상하단에 레터박스가 존재함.",
    "entities": "현우는 긴팔 셔츠를 입고 있으나 레퍼런스와 약간 다른 레이어드 형태임. 찰리는 숄이 없지만 가슴과 배 부위에 기계적인 파손 흔적이 뚜렷하게 묘사됨.",
    "hard_violations": [],
    "physics": "두 캐릭터 모두 부력에 의해 물속에 떠서 가라앉는 중이며, 힘이 빠진 채 축 늘어진 몸 상태를 자연스럽게 유지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "넓은 와이드 샷 구도 안에서 주변으로 가라앉는 화물들과 찰리의 파손된 몸체 등 지시된 상황과 디테일을 훨씬 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "수중 배경과 축 늘어진 실루엣은 잘 표현되었으나, 지시된 가라앉는 화물이 누락되었고 현우의 의상이 레퍼런스와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 캐릭터가 수면의 빛을 등지고 아래로 가라앉고 있으며, 몸의 방향은 깊은 바닥을 향함.",
        "built_space": "위에서 빛이 스며들고 아래로 갈수록 칠흑같이 어두워지는 심해 공간이며, 명시된 가라앉는 선박 화물은 보이지 않음.",
        "entities": "현우는 레퍼런스와 다른 짧은 반팔 티셔츠 차림임. 찰리는 숄을 두르지 않았고 요구된 몸체의 파손 흔적도 묘사되지 않음.",
        "hard_violations": [],
        "physics": "두 사람 모두 물속에 떠서 가라앉고 있으며, 의식이 없는 상태에 맞게 팔다리가 저항 없이 축 늘어져 있음."
       },
       {
        "label": "B",
        "direction": "두 캐릭터가 수평에 가깝게 나란히 부유하며 깊은 곳으로 가라앉고 있음.",
        "built_space": "빛이 점차 사라지는 심해 배경이며, 두 캐릭터 주변으로 직육면체 형태의 화물 상자들이 함께 가라앉고 있음. 화면 상하단에 레터박스가 존재함.",
        "entities": "현우는 긴팔 셔츠를 입고 있으나 레퍼런스와 약간 다른 레이어드 형태임. 찰리는 숄이 없지만 가슴과 배 부위에 기계적인 파손 흔적이 뚜렷하게 묘사됨.",
        "hard_violations": [],
        "physics": "두 캐릭터 모두 부력에 의해 물속에 떠서 가라앉는 중이며, 힘이 빠진 채 축 늘어진 몸 상태를 자연스럽게 유지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "넓은 와이드 샷 구도 안에서 주변으로 가라앉는 화물들과 찰리의 파손된 몸체 등 지시된 상황과 디테일을 훨씬 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "수중 배경과 축 늘어진 실루엣은 잘 표현되었으나, 지시된 가라앉는 화물이 누락되었고 현우의 의상이 레퍼런스와 다릅니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 캐릭터가 수면의 빛을 등지고 아래로 가라앉고 있으며, 몸의 방향은 깊은 바닥을 향함.",
        "built_space": "위에서 빛이 스며들고 아래로 갈수록 칠흑같이 어두워지는 심해 공간이며, 명시된 가라앉는 선박 화물은 보이지 않음.",
        "entities": "현우는 레퍼런스와 다른 짧은 반팔 티셔츠 차림임. 찰리는 숄을 두르지 않았고 요구된 몸체의 파손 흔적도 묘사되지 않음.",
        "hard_violations": [],
        "physics": "두 사람 모두 물속에 떠서 가라앉고 있으며, 의식이 없는 상태에 맞게 팔다리가 저항 없이 축 늘어져 있음."
       },
       {
        "label": "B",
        "direction": "두 캐릭터가 수평에 가깝게 나란히 부유하며 깊은 곳으로 가라앉고 있음.",
        "built_space": "빛이 점차 사라지는 심해 배경이며, 두 캐릭터 주변으로 직육면체 형태의 화물 상자들이 함께 가라앉고 있음. 화면 상하단에 레터박스가 존재함.",
        "entities": "현우는 긴팔 셔츠를 입고 있으나 레퍼런스와 약간 다른 레이어드 형태임. 찰리는 숄이 없지만 가슴과 배 부위에 기계적인 파손 흔적이 뚜렷하게 묘사됨.",
        "hard_violations": [],
        "physics": "두 캐릭터 모두 부력에 의해 물속에 떠서 가라앉는 중이며, 힘이 빠진 채 축 늘어진 몸 상태를 자연스럽게 유지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "수면 없는 어두운 심해의 와이드 구도와 화물 사이로 나란히 가라앉는 배치가 우세하지만, 찰리의 체형·얼굴·두른 천은 참조와 다르다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "축 처진 자세는 잘 보이지만, 밝고 선명한 수면과 화면을 크게 채운 앙각 구도가 심해의 단절 없는 여백을 훼손하며 의상과 찰리의 비례도 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽에서 머리를 왼쪽 아래로 떨어뜨리고 있으며 특정 대상을 바라보지 않는다. 오른쪽 찰리는 얼굴을 카메라 쪽으로 약간 숙이고 있다. 두 몸은 발이 아래쪽을 향한 채 나란히 놓여 하강 장면으로 읽히며, 능동적인 수영이나 추진 동작은 없다. 무기나 겨누는 물건은 없다.",
        "built_space": "구조물과 고정 설비는 보이지 않는다. 두 인물 주변에는 작은 상자형 화물 여러 개가 서로 떨어져 있고, 수면·해저·배의 바닥은 드러나지 않는다. 인물 위아래와 좌우에 어두운 물의 여백이 남아 있어 지정된 심해 공간과 와이드 구도에 가깝다. 상하 검은 띠는 있으나 분할 화면은 아니다.",
        "entities": "현우와 찰리 두 인물만 보인다. 현우는 헝클어진 검은 머리의 젊고 마른 동아시아계 남성으로 보이며, 어두운 긴소매 상의와 바지를 입었다. 얼굴의 정확한 일치와 나이는 어둠 때문에 확정하기 어렵고, 상의는 참조의 단추 셔츠와 차이가 있다. 찰리는 베이지색 각진 장갑과 흰 마스크형 얼굴을 지녔고 가슴과 허벅지에 손상처럼 보이는 부분이 있다. 다만 참조보다 인간형 비례가 강하고 얼굴 형상도 다르며 줄무늬 천은 없다. 주변 화물은 보이고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 몸은 공중이 아니라 물속에 있으며 부력과 물의 저항을 받으며 가라앉는 상태로 해석할 수 있다. 현우는 목과 한쪽 팔이 처지고 다른 손은 배 부근에 놓여 있으며, 다리는 서로 다른 정도로 굽어 있다. 찰리도 손과 발을 아래로 늘어뜨렸으나 몸통은 현우보다 곧다. 화물 역시 물속에서 침강하는 배치이므로 손이나 바닥의 지지가 없어도 물리적 모순은 아니다."
       },
       {
        "label": "B",
        "direction": "현우의 머리는 아래쪽으로 숙여져 있고 찰리의 얼굴도 아래와 화면 오른쪽을 향한다. 둘 다 특정 대상을 응시하지 않는다. 카메라가 발 아래에서 올려다보므로 발바닥이 관객 쪽을 향하고, 몸은 머리보다 발이 아래인 자세로 나란히 내려오는 것으로 읽힌다. 조준 물체나 능동적인 수영 동작은 없다.",
        "built_space": "고정 설비나 해저는 없지만 화면 상단에 밝고 물결치는 수면이 뚜렷하게 보인다. 이는 수면 없이 이어지는 심해의 어두운 배경이라는 지시와 다르다. 두 인물은 화면 높이 대부분을 차지해 주변 물의 여백이 좁다. 명확한 대형 화물이나 상자는 식별되지 않고 작은 부유 입자가 주변을 채운다.",
        "entities": "현우와 찰리 두 인물만 보인다. 현우는 검은 머리의 젊은 동아시아계 남성으로 보이지만, 참조의 긴소매 단추 셔츠 대신 반소매 티셔츠를 입고 있다. 찰리는 베이지 장갑과 흰 기계식 마스크를 지녔으나 다리가 길어 참조의 짧은 다리와 고릴라형 체형에서 벗어난다. 참조의 줄무늬 천은 없으며 기존 손상도 뚜렷하게 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "물속이라는 환경이 명확하며 두 몸은 부력과 유체 저항을 받는다. 현우의 양팔과 손목, 머리가 이완되어 내려오고 무릎도 약간 굽어 있어 의식 없는 침강 자세로 가능하다. 찰리의 팔과 손도 아래로 떨어져 있다. 발밑 지지대가 없다는 사실은 수중 침강에서 불가능한 부유를 뜻하지 않는다. 다만 정지 화면만으로 실제 하강 속도나 움직임은 확인할 수 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "수면 없는 어두운 심해의 와이드 구도와 화물 사이로 나란히 가라앉는 배치가 우세하지만, 찰리의 체형·얼굴·두른 천은 참조와 다르다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "축 처진 자세는 잘 보이지만, 밝고 선명한 수면과 화면을 크게 채운 앙각 구도가 심해의 단절 없는 여백을 훼손하며 의상과 찰리의 비례도 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽에서 머리를 왼쪽 아래로 떨어뜨리고 있으며 특정 대상을 바라보지 않는다. 오른쪽 찰리는 얼굴을 카메라 쪽으로 약간 숙이고 있다. 두 몸은 발이 아래쪽을 향한 채 나란히 놓여 하강 장면으로 읽히며, 능동적인 수영이나 추진 동작은 없다. 무기나 겨누는 물건은 없다.",
        "built_space": "구조물과 고정 설비는 보이지 않는다. 두 인물 주변에는 작은 상자형 화물 여러 개가 서로 떨어져 있고, 수면·해저·배의 바닥은 드러나지 않는다. 인물 위아래와 좌우에 어두운 물의 여백이 남아 있어 지정된 심해 공간과 와이드 구도에 가깝다. 상하 검은 띠는 있으나 분할 화면은 아니다.",
        "entities": "현우와 찰리 두 인물만 보인다. 현우는 헝클어진 검은 머리의 젊고 마른 동아시아계 남성으로 보이며, 어두운 긴소매 상의와 바지를 입었다. 얼굴의 정확한 일치와 나이는 어둠 때문에 확정하기 어렵고, 상의는 참조의 단추 셔츠와 차이가 있다. 찰리는 베이지색 각진 장갑과 흰 마스크형 얼굴을 지녔고 가슴과 허벅지에 손상처럼 보이는 부분이 있다. 다만 참조보다 인간형 비례가 강하고 얼굴 형상도 다르며 줄무늬 천은 없다. 주변 화물은 보이고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 몸은 공중이 아니라 물속에 있으며 부력과 물의 저항을 받으며 가라앉는 상태로 해석할 수 있다. 현우는 목과 한쪽 팔이 처지고 다른 손은 배 부근에 놓여 있으며, 다리는 서로 다른 정도로 굽어 있다. 찰리도 손과 발을 아래로 늘어뜨렸으나 몸통은 현우보다 곧다. 화물 역시 물속에서 침강하는 배치이므로 손이나 바닥의 지지가 없어도 물리적 모순은 아니다."
       },
       {
        "label": "A",
        "direction": "현우의 머리는 아래쪽으로 숙여져 있고 찰리의 얼굴도 아래와 화면 오른쪽을 향한다. 둘 다 특정 대상을 응시하지 않는다. 카메라가 발 아래에서 올려다보므로 발바닥이 관객 쪽을 향하고, 몸은 머리보다 발이 아래인 자세로 나란히 내려오는 것으로 읽힌다. 조준 물체나 능동적인 수영 동작은 없다.",
        "built_space": "고정 설비나 해저는 없지만 화면 상단에 밝고 물결치는 수면이 뚜렷하게 보인다. 이는 수면 없이 이어지는 심해의 어두운 배경이라는 지시와 다르다. 두 인물은 화면 높이 대부분을 차지해 주변 물의 여백이 좁다. 명확한 대형 화물이나 상자는 식별되지 않고 작은 부유 입자가 주변을 채운다.",
        "entities": "현우와 찰리 두 인물만 보인다. 현우는 검은 머리의 젊은 동아시아계 남성으로 보이지만, 참조의 긴소매 단추 셔츠 대신 반소매 티셔츠를 입고 있다. 찰리는 베이지 장갑과 흰 기계식 마스크를 지녔으나 다리가 길어 참조의 짧은 다리와 고릴라형 체형에서 벗어난다. 참조의 줄무늬 천은 없으며 기존 손상도 뚜렷하게 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "물속이라는 환경이 명확하며 두 몸은 부력과 유체 저항을 받는다. 현우의 양팔과 손목, 머리가 이완되어 내려오고 무릎도 약간 굽어 있어 의식 없는 침강 자세로 가능하다. 찰리의 팔과 손도 아래로 떨어져 있다. 발밑 지지대가 없다는 사실은 수중 침강에서 불가능한 부유를 뜻하지 않는다. 다만 정지 화면만으로 실제 하강 속도나 움직임은 확인할 수 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.125,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.125,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1125
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "넓은 와이드 샷 구도 안에서 주변으로 가라앉는 화물들과 찰리의 파손된 몸체 등 지시된 상황과 디테일을 훨씬 충실하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1125,
    "verdict_ko": "수중 배경과 축 늘어진 실루엣은 잘 표현되었으나, 지시된 가라앉는 화물이 누락되었고 현우의 의상이 레퍼런스와 다릅니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B02.png",
    "asset_id": "dac9a97a-00da-4766-aa32-41c0c8abe71c",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1293785>",
    "asset_id": "038f90ae-579f-46c4-8fd6-1c2c6d014b77",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1120327>",
    "asset_id": "23c61b20-dd4d-48bf-b945-fbdf76c4338f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-e928-73e4-83c3-9c6f97201a6d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh20__bgfirst_bg.png",
   "bg_asset_id": "e9739767-c664-4b57-931d-69841d4b4382",
   "bg_record_key": "S80sh20::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S80sh20::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:51:50.097558+00:00",
  "fingerprint": "c1d4a4750161b2411a171a264cb1e1aa5156a8069af6bece4f8a5226f8465618",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S80sh20_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S80sh20_sel.png",
  "source_sha256": "ec7e9c7a831b8ff3434689c7d85874c0b42f172bb3cdd0952f7693e52c7a0cfd",
  "file": "S80sh20_cine.png",
  "staged_sha256": "993c82c337da368caa0b4eda4aa63dfb975fdde12105a3e86f6c13af45863af3",
  "latency_ms": 10149
 },
 "S81sh5::signage": {
  "fp": "b9ac00a4e5c29977",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::8a4521a0f910a4bf": {
  "subjects": [],
  "subject_text": "제주도 연구소 임시보호소 병실\n흰색 침대와 밝은 형광등이 있는 병실. 창밖으로 에메랄드빛 바다가 넓게 펼쳐지고 실내에는 통화 장치가 마련되어 있다.",
  "identity": "canonical",
  "scope_id": "L266",
  "scope_role": "location_interior",
  "scope_sha": "5e6010fbf1088b41"
 },
 "S81sh5::bgfirst_bg": {
  "input_fingerprint": "3905670d577532fa",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5__bgfirst_bg.png",
  "asset_id": "0ebb2820-cefc-4db9-9c50-9e74728ec4d6",
  "input_asset_ids": [
   "87280629-3bc9-4dc2-9c89-4967995f760f",
   "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a"
  ]
 },
 "S81sh5": {
  "input_fingerprint": "5f62a8e69171019e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A white bed is lit by strong fluorescent lighting in the temporary shelter, with an emerald-blue sea visible through the window. The storm-dark shipboard setting has ended. 서지민: She wears research clothing and has her researcher access card available.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A white bed is lit by strong fluorescent lighting in the temporary shelter, with an emerald-blue sea visible through the window. The storm-dark shipboard setting has ended. 서지민: She wears research clothing and has her researcher access card available.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A white bed is lit by strong fluorescent lighting in the temporary shelter, with an emerald-blue sea visible through the window. The storm-dark shipboard setting has ended. 서지민: She wears research clothing and has her researcher access card available.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5__bgfirst_bg.png",
     "asset_id": "0ebb2820-cefc-4db9-9c50-9e74728ec4d6",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S81sh5.png",
     "asset_id": "87280629-3bc9-4dc2-9c89-4967995f760f",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:930556>",
     "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
     "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:930556>",
     "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 렌즈를 벗어나 방 안쪽(환자의 침상 방향)을 다정하게 향함.",
    "built_space": "방 내부에서 비스듬히 바라본 구도이며, 좌측 하단에 침대 모서리가 걸쳐 있고 우측에 열린 문과 문틀이 배치됨.",
    "entities": "서지민의 앳된 얼굴, 검은 머리, 흰색 연구복, 사원증 및 창밖의 바다 풍경이 모두 레퍼런스와 일치함.",
    "hard_violations": [
     "[gpt-high] 연구자 출입증과 별개인 휴대 기기를 손에 추가해, 지정되지 않은 소품을 만들어 냈다."
    ],
    "physics": "두 발로 바닥을 딛고 한 손은 문틀에, 다른 한 손은 자연스럽게 기기를 들고 안정적인 자세로 서 있음."
   },
   {
    "label": "B",
    "direction": "시선이 카메라 렌즈를 정면으로 응시함.",
    "built_space": "지시된 프레이밍과 달리 침대 전체가 프레임 중앙부를 차지하며, 문가가 정면 구도로 배치됨.",
    "entities": "서지민의 외모, 흰색 연구복, 사원증, 창밖 바다 풍경은 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "바닥에 두 발을 대고 양팔을 내린 채 다소 경직된 정자세로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 프레이밍(왼쪽 하단에 잘린 침대 모서리, 비스듬한 구도의 문가)을 완벽하게 구현했으며, 다정한 미소와 시선 처리가 돋보입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "침대가 화면 중앙에 온전히 노출되고 정면 구도로 카메라를 응시하여 샷 텍스트의 핵심 프레이밍 지시를 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 렌즈를 벗어나 방 안쪽(환자의 침상 방향)을 다정하게 향함.",
        "built_space": "방 내부에서 비스듬히 바라본 구도이며, 좌측 하단에 침대 모서리가 걸쳐 있고 우측에 열린 문과 문틀이 배치됨.",
        "entities": "서지민의 앳된 얼굴, 검은 머리, 흰색 연구복, 사원증 및 창밖의 바다 풍경이 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "두 발로 바닥을 딛고 한 손은 문틀에, 다른 한 손은 자연스럽게 기기를 들고 안정적인 자세로 서 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈를 정면으로 응시함.",
        "built_space": "지시된 프레이밍과 달리 침대 전체가 프레임 중앙부를 차지하며, 문가가 정면 구도로 배치됨.",
        "entities": "서지민의 외모, 흰색 연구복, 사원증, 창밖 바다 풍경은 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "바닥에 두 발을 대고 양팔을 내린 채 다소 경직된 정자세로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 프레이밍(왼쪽 하단에 잘린 침대 모서리, 비스듬한 구도의 문가)을 완벽하게 구현했으며, 다정한 미소와 시선 처리가 돋보입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "침대가 화면 중앙에 온전히 노출되고 정면 구도로 카메라를 응시하여 샷 텍스트의 핵심 프레이밍 지시를 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 렌즈를 벗어나 방 안쪽(환자의 침상 방향)을 다정하게 향함.",
        "built_space": "방 내부에서 비스듬히 바라본 구도이며, 좌측 하단에 침대 모서리가 걸쳐 있고 우측에 열린 문과 문틀이 배치됨.",
        "entities": "서지민의 앳된 얼굴, 검은 머리, 흰색 연구복, 사원증 및 창밖의 바다 풍경이 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "두 발로 바닥을 딛고 한 손은 문틀에, 다른 한 손은 자연스럽게 기기를 들고 안정적인 자세로 서 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈를 정면으로 응시함.",
        "built_space": "지시된 프레이밍과 달리 침대 전체가 프레임 중앙부를 차지하며, 문가가 정면 구도로 배치됨.",
        "entities": "서지민의 외모, 흰색 연구복, 사원증, 창밖 바다 풍경은 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "바닥에 두 발을 대고 양팔을 내린 채 다소 경직된 정자세로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "문턱에서 미소 짓는 서지민의 전신과 연구복은 충실하지만, 침대 대부분을 보여 주어 왼쪽 아래 작은 모서리만 남기라는 구도를 어긴다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "병실 안에서 문가를 비스듬히 보는 시점과 침상 쪽 미소는 더 정확하지만, 지정되지 않은 휴대 기기를 추가했으며 침대 전경도 지나치게 크다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "서지민의 몸과 얼굴은 문밖에서 병실 안쪽 카메라 방향을 향하고 있으며 미소 짓는다. 시선은 왼쪽 침대보다는 렌즈 가까이에 걸려 있어 현우의 침상 쪽을 본다는 관계는 약하다. 겨누거나 조작하는 물건은 없다.",
        "built_space": "왼쪽에 바다를 향한 큰 3분할 창 하나와 흰 침대 하나, 오른쪽에 서지민이 선 열린 출입구 하나가 보인다. 전경 양옆에도 출입구 테두리가 보여 침상 곁보다는 다른 문 입구에서 방을 바라보는 인상이 강하다. 천장에는 실내 선형 조명 하나와 문 너머 조명 하나, 직사각형 환기 설비 두 개가 보인다. 참고의 밝은 벽과 회색 바닥은 유지되지만, 창 오른쪽 벽의 화면 대신 출입구가 강조되어 장소 배치의 일치가 약하다. 침대는 작은 모서리가 아니라 매트리스 측면과 하부 프레임까지 크게 노출된다.",
        "entities": "인물은 검은 어깨 길이 머리의 앳된 동아시아계 여성 한 명으로, 한국인 20세 서지민이라는 설정과 대체로 부합한다. 흰 연구복, 허리 장비, 장갑, 흰 신발은 인물 참고와 가깝고 목에 출입증이 있다. 참고의 머리 위 고글은 없다. 흰 침대와 낮의 푸른 바다가 있으며 현우나 다른 사람은 나타나지 않는다. 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 신발이 문턱 부근 바닥에 닿아 몸을 지탱한다. 양팔은 몸 옆으로 내려와 있고 출입증은 목걸이에 매달려 있다. 침대는 바닥에 닿은 바퀴와 프레임으로 지지되며, 바닥의 흐릿한 반사도 인물과 창의 위치에 맞는다. 지지 없이 떠 있는 몸이나 물건은 없다."
       },
       {
        "label": "B",
        "direction": "서지민은 얼굴과 시선을 화면 왼쪽 병실 내부의 침상 방향으로 돌리고 다정하게 웃는다. 카메라를 정면으로 응시하지 않아 보이지 않는 현우를 향한 반응이 더 잘 읽힌다. 가슴 앞의 직사각형 기기는 사용하지 않고 들고 있으며 화면인지 뒷면인지는 확정하기 어렵다.",
        "built_space": "왼쪽에 바다 창 하나, 가운데 벽에 상단 카메라가 달린 화면 하나와 세로 제어판 하나, 오른쪽에 열린 출입구 하나가 보인다. 문 너머에는 복도와 옆 출입구 일부, 여러 천장 조명이 있다. 참고의 창·벽걸이 화면·밝은 벽·회색 바닥 관계를 A보다 잘 유지한다. 서지민은 문턱에 서 있고 카메라는 병실 안에서 출입구를 비스듬히 본다. 다만 왼쪽 아래 침대는 작은 모서리를 넘어 넓은 전경을 차지한다.",
        "entities": "검은 어깨 길이 머리와 앳된 얼굴의 동아시아계 여성 한 명이 보여 서지민의 기본 외형과 부합한다. 흰 연구복과 신발, 목걸이형 출입증은 있지만 참고의 고글·장갑·허리 장비는 보이지 않는다. 출입증과 별개로 한 손에 휴대 기기로 보이는 직사각형 물건을 들고 있는데, 이는 지시된 소품이 아니다. 흰 침대와 낮의 푸른 바다가 있고 다른 사람은 없다. 출입증의 작은 표기는 판독되지 않는다.",
        "hard_violations": [
         "연구자 출입증과 별개인 휴대 기기를 손에 추가해, 지정되지 않은 소품을 만들어 냈다."
        ],
        "physics": "앞쪽 발바닥과 뒤쪽 발이 문턱 바닥에 닿아 체중을 받는다. 한 손은 문 가장자리에 자연스럽게 접촉하고 다른 손은 직사각형 기기를 가슴 앞에서 붙잡는다. 출입증은 목걸이에 매달리고 침구는 매트리스 위에 놓여 있다. 다리를 살짝 교차한 자세는 서서 기대는 동작으로 가능하며 지지 없는 부유는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "문턱에서 미소 짓는 서지민의 전신과 연구복은 충실하지만, 침대 대부분을 보여 주어 왼쪽 아래 작은 모서리만 남기라는 구도를 어긴다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "병실 안에서 문가를 비스듬히 보는 시점과 침상 쪽 미소는 더 정확하지만, 지정되지 않은 휴대 기기를 추가했으며 침대 전경도 지나치게 크다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "서지민의 몸과 얼굴은 문밖에서 병실 안쪽 카메라 방향을 향하고 있으며 미소 짓는다. 시선은 왼쪽 침대보다는 렌즈 가까이에 걸려 있어 현우의 침상 쪽을 본다는 관계는 약하다. 겨누거나 조작하는 물건은 없다.",
        "built_space": "왼쪽에 바다를 향한 큰 3분할 창 하나와 흰 침대 하나, 오른쪽에 서지민이 선 열린 출입구 하나가 보인다. 전경 양옆에도 출입구 테두리가 보여 침상 곁보다는 다른 문 입구에서 방을 바라보는 인상이 강하다. 천장에는 실내 선형 조명 하나와 문 너머 조명 하나, 직사각형 환기 설비 두 개가 보인다. 참고의 밝은 벽과 회색 바닥은 유지되지만, 창 오른쪽 벽의 화면 대신 출입구가 강조되어 장소 배치의 일치가 약하다. 침대는 작은 모서리가 아니라 매트리스 측면과 하부 프레임까지 크게 노출된다.",
        "entities": "인물은 검은 어깨 길이 머리의 앳된 동아시아계 여성 한 명으로, 한국인 20세 서지민이라는 설정과 대체로 부합한다. 흰 연구복, 허리 장비, 장갑, 흰 신발은 인물 참고와 가깝고 목에 출입증이 있다. 참고의 머리 위 고글은 없다. 흰 침대와 낮의 푸른 바다가 있으며 현우나 다른 사람은 나타나지 않는다. 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 신발이 문턱 부근 바닥에 닿아 몸을 지탱한다. 양팔은 몸 옆으로 내려와 있고 출입증은 목걸이에 매달려 있다. 침대는 바닥에 닿은 바퀴와 프레임으로 지지되며, 바닥의 흐릿한 반사도 인물과 창의 위치에 맞는다. 지지 없이 떠 있는 몸이나 물건은 없다."
       },
       {
        "label": "A",
        "direction": "서지민은 얼굴과 시선을 화면 왼쪽 병실 내부의 침상 방향으로 돌리고 다정하게 웃는다. 카메라를 정면으로 응시하지 않아 보이지 않는 현우를 향한 반응이 더 잘 읽힌다. 가슴 앞의 직사각형 기기는 사용하지 않고 들고 있으며 화면인지 뒷면인지는 확정하기 어렵다.",
        "built_space": "왼쪽에 바다 창 하나, 가운데 벽에 상단 카메라가 달린 화면 하나와 세로 제어판 하나, 오른쪽에 열린 출입구 하나가 보인다. 문 너머에는 복도와 옆 출입구 일부, 여러 천장 조명이 있다. 참고의 창·벽걸이 화면·밝은 벽·회색 바닥 관계를 A보다 잘 유지한다. 서지민은 문턱에 서 있고 카메라는 병실 안에서 출입구를 비스듬히 본다. 다만 왼쪽 아래 침대는 작은 모서리를 넘어 넓은 전경을 차지한다.",
        "entities": "검은 어깨 길이 머리와 앳된 얼굴의 동아시아계 여성 한 명이 보여 서지민의 기본 외형과 부합한다. 흰 연구복과 신발, 목걸이형 출입증은 있지만 참고의 고글·장갑·허리 장비는 보이지 않는다. 출입증과 별개로 한 손에 휴대 기기로 보이는 직사각형 물건을 들고 있는데, 이는 지시된 소품이 아니다. 흰 침대와 낮의 푸른 바다가 있고 다른 사람은 없다. 출입증의 작은 표기는 판독되지 않는다.",
        "hard_violations": [
         "연구자 출입증과 별개인 휴대 기기를 손에 추가해, 지정되지 않은 소품을 만들어 냈다."
        ],
        "physics": "앞쪽 발바닥과 뒤쪽 발이 문턱 바닥에 닿아 체중을 받는다. 한 손은 문 가장자리에 자연스럽게 접촉하고 다른 손은 직사각형 기기를 가슴 앞에서 붙잡는다. 출입증은 목걸이에 매달리고 침구는 매트리스 위에 놓여 있다. 다리를 살짝 교차한 자세는 서서 기대는 동작으로 가능하며 지지 없는 부유는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.429
   },
   "violations": {
    "A": [
     "[gpt-high] 연구자 출입증과 별개인 휴대 기기를 손에 추가해, 지정되지 않은 소품을 만들어 냈다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1429
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "지정된 프레이밍(왼쪽 하단에 잘린 침대 모서리, 비스듬한 구도의 문가)을 완벽하게 구현했으며, 다정한 미소와 시선 처리가 돋보입니다.  ★위반: [gpt-high] 연구자 출입증과 별개인 휴대 기기를 손에 추가해, 지정되지 않은 소품을 만들어 냈다."
   },
   {
    "label": "B",
    "score": 1429,
    "verdict_ko": "침대가 화면 중앙에 온전히 노출되고 정면 구도로 카메라를 응시하여 샷 텍스트의 핵심 프레이밍 지시를 크게 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
    "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:930556>",
    "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-ec67-7f4d-aed7-a2a8dae8f06d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5__bgfirst_bg.png",
   "bg_asset_id": "0ebb2820-cefc-4db9-9c50-9e74728ec4d6",
   "bg_record_key": "S81sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S81sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:53:05.948762+00:00",
  "fingerprint": "5f589b6424fdb3834d12ba42a1dc32fad2a087ed84cf188879efe95e3e7b4d09",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S81sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S81sh5_sel.png",
  "source_sha256": "b2781e33e676d26f9590beb024434eb21d5ba2a0c6272c4957462668b491f850",
  "file": "S81sh5_cine.png",
  "staged_sha256": "80ae0cc07b63ca4cac4ef127028a00d964cf64596ec3e65fadcd12d7b7043e1e",
  "latency_ms": 9011
 },
 "S81sh13::signage": {
  "fp": "8d90e05dd031158a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S81sh13::bgfirst_bg": {
  "input_fingerprint": "947b3aa49300d01c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13__bgfirst_bg.png",
  "asset_id": "2596dc8d-81c7-4a2d-aa17-c63234623f09",
  "input_asset_ids": [
   "2d7ff4e0-834c-440a-92c5-40e9da371422",
   "4bcfae29-00ac-4c7b-a17f-ba83dae38ae3"
  ]
 },
 "S81sh13": {
  "input_fingerprint": "b57a293d95db4d25",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 현우: He is wearing the fresh white clothes provided during his recovery and has stopped in the research-facility corridor.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 현우: He is wearing the fresh white clothes provided during his recovery and has stopped in the research-facility corridor.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 현우: He is wearing the fresh white clothes provided during his recovery and has stopped in the research-facility corridor.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13__bgfirst_bg.png",
     "asset_id": "2596dc8d-81c7-4a2d-aa17-c63234623f09",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S81sh13.png",
     "asset_id": "2d7ff4e0-834c-440a-92c5-40e9da371422",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:826250>",
     "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B01.png",
     "asset_id": "4bcfae29-00ac-4c7b-a17f-ba83dae38ae3",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:826250>",
     "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선이 정면이 아닌 화면 우측 측면을 향해 있음.",
    "built_space": "레퍼런스와 일치하는 연구소 복도가 배경으로 자연스럽게 배치됨.",
    "entities": "현우의 얼굴과 헤어스타일은 레퍼런스와 일치하나, 흰 옷의 형태(단추 없음)가 약간 다름.",
    "hard_violations": [],
    "physics": "자연스럽게 서 있는 상태로 물리적 오류 없음."
   },
   {
    "label": "B",
    "direction": "두 눈을 동그랗게 뜨고 정면(카메라 렌즈 방향)을 빤히 주시하고 있음.",
    "built_space": "레퍼런스와 완벽히 일치하는 복도 배경이 깊이감 있게 구현됨.",
    "entities": "현우의 얼굴, 헤어스타일, 그리고 단추가 있는 흰 옷의 형태가 레퍼런스와 매우 잘 일치함.",
    "hard_violations": [
     "[gpt-high] 화면 하단 흰 상의에 파란 문자형 로고 일부가 선명하게 노출되어, 로고와 읽을 수 있는 표식을 금지한 조건을 위반한다."
    ],
    "physics": "정면을 향해 멈춰 선 자세가 안정적으로 표현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "정면을 빤히 주시하는 놀란 표정과 지정된 복도 배경을 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "앞을 주시해야 한다는 지시와 달리 시선이 측면을 향하고 있어 몰입도를 떨어뜨립니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선이 정면이 아닌 화면 우측 측면을 향해 있음.",
        "built_space": "레퍼런스와 일치하는 연구소 복도가 배경으로 자연스럽게 배치됨.",
        "entities": "현우의 얼굴과 헤어스타일은 레퍼런스와 일치하나, 흰 옷의 형태(단추 없음)가 약간 다름.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 상태로 물리적 오류 없음."
       },
       {
        "label": "B",
        "direction": "두 눈을 동그랗게 뜨고 정면(카메라 렌즈 방향)을 빤히 주시하고 있음.",
        "built_space": "레퍼런스와 완벽히 일치하는 복도 배경이 깊이감 있게 구현됨.",
        "entities": "현우의 얼굴, 헤어스타일, 그리고 단추가 있는 흰 옷의 형태가 레퍼런스와 매우 잘 일치함.",
        "hard_violations": [],
        "physics": "정면을 향해 멈춰 선 자세가 안정적으로 표현됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "정면을 빤히 주시하는 놀란 표정과 지정된 복도 배경을 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "앞을 주시해야 한다는 지시와 달리 시선이 측면을 향하고 있어 몰입도를 떨어뜨립니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선이 정면이 아닌 화면 우측 측면을 향해 있음.",
        "built_space": "레퍼런스와 일치하는 연구소 복도가 배경으로 자연스럽게 배치됨.",
        "entities": "현우의 얼굴과 헤어스타일은 레퍼런스와 일치하나, 흰 옷의 형태(단추 없음)가 약간 다름.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 상태로 물리적 오류 없음."
       },
       {
        "label": "B",
        "direction": "두 눈을 동그랗게 뜨고 정면(카메라 렌즈 방향)을 빤히 주시하고 있음.",
        "built_space": "레퍼런스와 완벽히 일치하는 복도 배경이 깊이감 있게 구현됨.",
        "entities": "현우의 얼굴, 헤어스타일, 그리고 단추가 있는 흰 옷의 형태가 레퍼런스와 매우 잘 일치함.",
        "hard_violations": [],
        "physics": "정면을 향해 멈춰 선 자세가 안정적으로 표현됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴을 더 크게 잡고 시선 앞에 여백을 남겨, 멈춰 서서 두 눈을 동그랗게 뜬 현우의 놀람을 더 충실히 구현했으나 복도 배경은 요구보다 선명하다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "정면을 응시하는 놀란 얼굴과 연구시설 복도는 잘 맞지만, 화면 하단 의복에 파란 로고 일부가 노출되어 명시적인 로고 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 거의 카메라 정면을 향하며 렌즈 부근의 화면 밖 대상을 응시한다. 구체적인 응시 대상은 보이지 않지만, 앞을 빤히 본다는 지시와 양립한다. 이동이나 손에 든 물체는 보이지 않는다.",
        "built_space": "현우가 복도 중앙에 서 있고 뒤로 열린 금속·유리 문틀 한 세트가 보인다. 왼쪽에는 장비 상자를 얹은 카트 한 대, 오른쪽에는 실험실 관찰창 한 개와 포개진 장비 상자 두 개가 보인다. 천장 조명과 회색 바닥, 벽의 금속 보호대가 장소 참고와 대체로 일치한다. 창 안쪽의 조명과 실험대는 유리를 통한 시야로 설명 가능하며 불가능한 반사는 보이지 않는다.",
        "entities": "앳된 동아시아계 남성 한 명이며 헝클어진 검은 머리와 얼굴 윤곽은 현우 참고와 대체로 맞는다. 외모만으로 한국계 미국인이라는 국적·배경까지 확인할 수는 없다. 흰 회복복의 둥근 목선과 단추가 보이고, 화면 아래에는 파란 로고 일부가 드러난다. 눈은 정상적인 홍채와 동공을 유지한 채 크게 떠져 있다. 다른 사람이나 휴대 소품은 없다.",
        "hard_violations": [
         "화면 하단 흰 상의에 파란 문자형 로고 일부가 선명하게 노출되어, 로고와 읽을 수 있는 표식을 금지한 조건을 위반한다."
        ],
        "physics": "머리가 목과 어깨에 자연스럽게 연결되고 상체는 수직으로 서 있다. 발과 바닥 접촉은 클로즈업 밖이라 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 장비 상자는 카트 선반이나 바닥 위에 놓여 있으며, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "얼굴은 화면 오른쪽으로 조금 돌아가 있고 두 눈도 같은 쪽의 화면 밖 대상을 집중해서 본다. 대상 자체는 보이지 않는다. 얼굴이 향한 앞쪽에 시선이 놓여 있어 앞을 빤히 주시하는 행동으로 읽히며, 오른쪽에 시선 여백이 남는다.",
        "built_space": "현우는 화면 왼쪽 전경의 복도에 서 있다. 뒤에는 열린 유리 양문 한 세트와 복도 끝 외창 한 개, 왼쪽 장비 카트 한 대, 먼 오른쪽 작은 금속 탁자 한 대가 보인다. 오른쪽 관찰창은 중경과 가까운 전경에 각각 한 개씩 보이는데, 가까운 창은 장소 참고에서 직접 확인되지 않아 공간 일치성이 다소 떨어진다. 금속 문틀, 회색 바닥, 천장 조명과 끝창의 낮 풍경은 참고와 부합한다. 배경은 부드럽게 흐려지기보다 비교적 또렷하다.",
        "entities": "젊은 동아시아계 남성 한 명만 등장한다. 검고 헝클어진 머리, 앳된 얼굴과 마른 체격은 현우 설정에 부합하나 얼굴은 참고보다 조금 둥글게 보인다. 흰 회복복을 입었지만 보이는 목선은 참고의 단추 달린 둥근 목선과 다소 다르다. 정상적인 눈을 크게 뜨고 입을 벌려 놀람을 연기한다. 판독 가능한 글자나 로고, 추가 인물, 휴대 소품은 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 자연스럽게 이어지고 정지한 상체 자세로 보인다. 하체가 잘려 있어 발의 지지와 멈추기 직전의 체중 이동은 판단할 수 없다. 보이는 카트와 탁자는 바닥에 서 있고 상자는 선반에 얹혀 있다. 지지 없는 신체나 물체, 불가능한 관절 배치는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴을 더 크게 잡고 시선 앞에 여백을 남겨, 멈춰 서서 두 눈을 동그랗게 뜬 현우의 놀람을 더 충실히 구현했으나 복도 배경은 요구보다 선명하다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "정면을 응시하는 놀란 얼굴과 연구시설 복도는 잘 맞지만, 화면 하단 의복에 파란 로고 일부가 노출되어 명시적인 로고 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 거의 카메라 정면을 향하며 렌즈 부근의 화면 밖 대상을 응시한다. 구체적인 응시 대상은 보이지 않지만, 앞을 빤히 본다는 지시와 양립한다. 이동이나 손에 든 물체는 보이지 않는다.",
        "built_space": "현우가 복도 중앙에 서 있고 뒤로 열린 금속·유리 문틀 한 세트가 보인다. 왼쪽에는 장비 상자를 얹은 카트 한 대, 오른쪽에는 실험실 관찰창 한 개와 포개진 장비 상자 두 개가 보인다. 천장 조명과 회색 바닥, 벽의 금속 보호대가 장소 참고와 대체로 일치한다. 창 안쪽의 조명과 실험대는 유리를 통한 시야로 설명 가능하며 불가능한 반사는 보이지 않는다.",
        "entities": "앳된 동아시아계 남성 한 명이며 헝클어진 검은 머리와 얼굴 윤곽은 현우 참고와 대체로 맞는다. 외모만으로 한국계 미국인이라는 국적·배경까지 확인할 수는 없다. 흰 회복복의 둥근 목선과 단추가 보이고, 화면 아래에는 파란 로고 일부가 드러난다. 눈은 정상적인 홍채와 동공을 유지한 채 크게 떠져 있다. 다른 사람이나 휴대 소품은 없다.",
        "hard_violations": [
         "화면 하단 흰 상의에 파란 문자형 로고 일부가 선명하게 노출되어, 로고와 읽을 수 있는 표식을 금지한 조건을 위반한다."
        ],
        "physics": "머리가 목과 어깨에 자연스럽게 연결되고 상체는 수직으로 서 있다. 발과 바닥 접촉은 클로즈업 밖이라 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 장비 상자는 카트 선반이나 바닥 위에 놓여 있으며, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "얼굴은 화면 오른쪽으로 조금 돌아가 있고 두 눈도 같은 쪽의 화면 밖 대상을 집중해서 본다. 대상 자체는 보이지 않는다. 얼굴이 향한 앞쪽에 시선이 놓여 있어 앞을 빤히 주시하는 행동으로 읽히며, 오른쪽에 시선 여백이 남는다.",
        "built_space": "현우는 화면 왼쪽 전경의 복도에 서 있다. 뒤에는 열린 유리 양문 한 세트와 복도 끝 외창 한 개, 왼쪽 장비 카트 한 대, 먼 오른쪽 작은 금속 탁자 한 대가 보인다. 오른쪽 관찰창은 중경과 가까운 전경에 각각 한 개씩 보이는데, 가까운 창은 장소 참고에서 직접 확인되지 않아 공간 일치성이 다소 떨어진다. 금속 문틀, 회색 바닥, 천장 조명과 끝창의 낮 풍경은 참고와 부합한다. 배경은 부드럽게 흐려지기보다 비교적 또렷하다.",
        "entities": "젊은 동아시아계 남성 한 명만 등장한다. 검고 헝클어진 머리, 앳된 얼굴과 마른 체격은 현우 설정에 부합하나 얼굴은 참고보다 조금 둥글게 보인다. 흰 회복복을 입었지만 보이는 목선은 참고의 단추 달린 둥근 목선과 다소 다르다. 정상적인 눈을 크게 뜨고 입을 벌려 놀람을 연기한다. 판독 가능한 글자나 로고, 추가 인물, 휴대 소품은 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 자연스럽게 이어지고 정지한 상체 자세로 보인다. 하체가 잘려 있어 발의 지지와 멈추기 직전의 체중 이동은 판단할 수 없다. 보이는 카트와 탁자는 바닥에 서 있고 상자는 선반에 얹혀 있다. 지지 없는 신체나 물체, 불가능한 관절 배치는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.625,
    "B": 1.625
   },
   "adjusted": {
    "A": 1.625,
    "B": 1.375
   },
   "violations": {
    "B": [
     "[gpt-high] 화면 하단 흰 상의에 파란 문자형 로고 일부가 선명하게 노출되어, 로고와 읽을 수 있는 표식을 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1375,
   "A": 1625
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1375,
    "verdict_ko": "정면을 빤히 주시하는 놀란 표정과 지정된 복도 배경을 완벽하게 구현했습니다.  ★위반: [gpt-high] 화면 하단 흰 상의에 파란 문자형 로고 일부가 선명하게 노출되어, 로고와 읽을 수 있는 표식을 금지한 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1625,
    "verdict_ko": "앞을 주시해야 한다는 지시와 달리 시선이 측면을 향하고 있어 몰입도를 떨어뜨립니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B01.png",
    "asset_id": "4bcfae29-00ac-4c7b-a17f-ba83dae38ae3",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-efa5-7ba6-b3a5-81c8da1dc1bf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13__bgfirst_bg.png",
   "bg_asset_id": "2596dc8d-81c7-4a2d-aa17-c63234623f09",
   "bg_record_key": "S81sh13::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S81sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:54:16.841640+00:00",
  "fingerprint": "a130751e741f7bd903cac009f5c4c10b8b5edb1f91294bc8ac08b4ded6d16bd4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S81sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S81sh13_sel.png",
  "source_sha256": "6fbb2234793e2b0ab2d4ed35e3b378e948ce6fd1e536bbe0fe670ad8b09a1aa9",
  "file": "S81sh13_cine.png",
  "staged_sha256": "f9cc7674c95e70cbc643e69ddbfc173e39cd9156012ce291297dc1efc49ba601",
  "latency_ms": 9372
 },
 "S81sh16::signage": {
  "fp": "27e27e766769817c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S81sh16": {
  "input_fingerprint": "d99584bcc9de83e7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두꺼운 자동문이 좌우로 스르륵 열리고 있는 틈새 너머로 두 눈을 뜬 현우의 상체.\n\nLOCATION (lock): At the access-controlled doorway between the research corridor and main center, under the facility's interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Left separating door panel in the middle-left of the frame, foreground; Right separating door panel in the middle-right of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Automatic door panels (Thick panels sliding apart to either side) — Their room-facing surfaces and inner edges flank the view of 현우 across the threshold; used as Moving lateral frame that progressively reveals his upper body; Corridor beyond the threshold (Visible behind 현우 through the opening) — Seen from inside the room looking back toward the corridor; used as Maintains spatial continuity across the reverse-position cut.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained, setting-appropriate ambient illumination across the threshold without inventing a contrast between the two spaces.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The access-controlled door is opening after a researcher card has been presented to the wall reader. 현우: He remains dressed in white recovery clothes at the entrance to the research center.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두꺼운 자동문이 좌우로 스르륵 열리고 있는 틈새 너머로 두 눈을 뜬 현우의 상체.\n\nLOCATION (lock): At the access-controlled doorway between the research corridor and main center, under the facility's interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Left separating door panel in the middle-left of the frame, foreground; Right separating door panel in the middle-right of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Automatic door panels (Thick panels sliding apart to either side) — Their room-facing surfaces and inner edges flank the view of 현우 across the threshold; used as Moving lateral frame that progressively reveals his upper body; Corridor beyond the threshold (Visible behind 현우 through the opening) — Seen from inside the room looking back toward the corridor; used as Maintains spatial continuity across the reverse-position cut.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained, setting-appropriate ambient illumination across the threshold without inventing a contrast between the two spaces.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The access-controlled door is opening after a researcher card has been presented to the wall reader. 현우: He remains dressed in white recovery clothes at the entrance to the research center.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두꺼운 자동문이 좌우로 스르륵 열리고 있는 틈새 너머로 두 눈을 뜬 현우의 상체.\n\nLOCATION (lock): At the access-controlled doorway between the research corridor and main center, under the facility's interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Left separating door panel in the middle-left of the frame, foreground; Right separating door panel in the middle-right of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Automatic door panels (Thick panels sliding apart to either side) — Their room-facing surfaces and inner edges flank the view of 현우 across the threshold; used as Moving lateral frame that progressively reveals his upper body; Corridor beyond the threshold (Visible behind 현우 through the opening) — Seen from inside the room looking back toward the corridor; used as Maintains spatial continuity across the reverse-position cut.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained, setting-appropriate ambient illumination across the threshold without inventing a contrast between the two spaces.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The access-controlled door is opening after a researcher card has been presented to the wall reader. 현우: He remains dressed in white recovery clothes at the entrance to the research center.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우가 정면의 카메라 방향을 응시하고 있음.",
    "built_space": "프레임 전경에 금속 재질의 문이 배치되어 있으나, 좌우 패널 사이에 수평 방향의 금속 막대가 가로질러 하나로 연결되어 있음. 배경은 레퍼런스와 일치하는 연구소 복도임.",
    "entities": "현우의 외모, 헝클어진 검은 머리, 흰색 환자복 모두 레퍼런스 및 프롬프트와 잘 일치함.",
    "hard_violations": [
     "[gemini-pro] physically impossible staging (좌우로 분리되어 열리는 자동문 패널 사이를 가로질러 하나의 구조물로 연결하는 수평 막대가 존재하여 물리적으로 불가능함)"
    ],
    "physics": "공중에 뜬 물체는 없으나, 좌우로 분리되어야 할 두 문 패널이 수평 금속 막대로 물리적으로 연결되어 있어 구조적인 모순이 있음."
   },
   {
    "label": "B",
    "direction": "현우가 정면의 카메라 방향을 응시하고 있음.",
    "built_space": "전경에 좌우로 완전히 분리된 두꺼운 금속 문 패널이 위치하며, 중앙의 수직 틈새를 통해 현우가 보임. 좌측 패널 상단 부분이 열려 있는 형태이나 좌우로 미닫이 작동이 가능한 구조임. 배경은 레퍼런스의 복도와 일치함.",
    "entities": "현우의 앳된 외모, 검은 머리, 흰색 환자복이 레퍼런스 및 지시사항과 정확히 일치함.",
    "hard_violations": [],
    "physics": "문 패널이 물리적으로 분리되어 있어 정상적으로 열리고 닫힐 수 있으며, 지지되지 않거나 불가능한 자세를 취한 객체가 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "좌우로 열리는 자동문의 틈새를 통해 두 눈을 뜬 현우의 상체를 보여주는 숏의 구도와 공간적 맥락을 물리적 모순 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "열리는 자동문의 양측 패널을 수평 막대로 연결해버려 프롬프트에서 지시된 좌우 분리 동작이 물리적으로 불가능한 구조를 만들었으므로 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우가 정면의 카메라 방향을 응시하고 있음.",
        "built_space": "프레임 전경에 금속 재질의 문이 배치되어 있으나, 좌우 패널 사이에 수평 방향의 금속 막대가 가로질러 하나로 연결되어 있음. 배경은 레퍼런스와 일치하는 연구소 복도임.",
        "entities": "현우의 외모, 헝클어진 검은 머리, 흰색 환자복 모두 레퍼런스 및 프롬프트와 잘 일치함.",
        "hard_violations": [
         "physically impossible staging (좌우로 분리되어 열리는 자동문 패널 사이를 가로질러 하나의 구조물로 연결하는 수평 막대가 존재하여 물리적으로 불가능함)"
        ],
        "physics": "공중에 뜬 물체는 없으나, 좌우로 분리되어야 할 두 문 패널이 수평 금속 막대로 물리적으로 연결되어 있어 구조적인 모순이 있음."
       },
       {
        "label": "B",
        "direction": "현우가 정면의 카메라 방향을 응시하고 있음.",
        "built_space": "전경에 좌우로 완전히 분리된 두꺼운 금속 문 패널이 위치하며, 중앙의 수직 틈새를 통해 현우가 보임. 좌측 패널 상단 부분이 열려 있는 형태이나 좌우로 미닫이 작동이 가능한 구조임. 배경은 레퍼런스의 복도와 일치함.",
        "entities": "현우의 앳된 외모, 검은 머리, 흰색 환자복이 레퍼런스 및 지시사항과 정확히 일치함.",
        "hard_violations": [],
        "physics": "문 패널이 물리적으로 분리되어 있어 정상적으로 열리고 닫힐 수 있으며, 지지되지 않거나 불가능한 자세를 취한 객체가 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "좌우로 열리는 자동문의 틈새를 통해 두 눈을 뜬 현우의 상체를 보여주는 숏의 구도와 공간적 맥락을 물리적 모순 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "열리는 자동문의 양측 패널을 수평 막대로 연결해버려 프롬프트에서 지시된 좌우 분리 동작이 물리적으로 불가능한 구조를 만들었으므로 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우가 정면의 카메라 방향을 응시하고 있음.",
        "built_space": "프레임 전경에 금속 재질의 문이 배치되어 있으나, 좌우 패널 사이에 수평 방향의 금속 막대가 가로질러 하나로 연결되어 있음. 배경은 레퍼런스와 일치하는 연구소 복도임.",
        "entities": "현우의 외모, 헝클어진 검은 머리, 흰색 환자복 모두 레퍼런스 및 프롬프트와 잘 일치함.",
        "hard_violations": [
         "physically impossible staging (좌우로 분리되어 열리는 자동문 패널 사이를 가로질러 하나의 구조물로 연결하는 수평 막대가 존재하여 물리적으로 불가능함)"
        ],
        "physics": "공중에 뜬 물체는 없으나, 좌우로 분리되어야 할 두 문 패널이 수평 금속 막대로 물리적으로 연결되어 있어 구조적인 모순이 있음."
       },
       {
        "label": "B",
        "direction": "현우가 정면의 카메라 방향을 응시하고 있음.",
        "built_space": "전경에 좌우로 완전히 분리된 두꺼운 금속 문 패널이 위치하며, 중앙의 수직 틈새를 통해 현우가 보임. 좌측 패널 상단 부분이 열려 있는 형태이나 좌우로 미닫이 작동이 가능한 구조임. 배경은 레퍼런스의 복도와 일치함.",
        "entities": "현우의 앳된 외모, 검은 머리, 흰색 환자복이 레퍼런스 및 지시사항과 정확히 일치함.",
        "hard_violations": [],
        "physics": "문 패널이 물리적으로 분리되어 있어 정상적으로 열리고 닫힐 수 있으며, 지지되지 않거나 불가능한 자세를 취한 객체가 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "문틈 너머 눈을 뜬 현우와 복도는 맞지만, 좌우 문짝의 구조가 비대칭이고 왼쪽 불투명 면이 상체를 많이 가려 두꺼운 양개 자동문이 열리는 구도가 덜 명확하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "전경의 두꺼운 좌우 문짝, 그 틈의 현우 상체, 뒤쪽 복도를 미디엄 숏으로 명확히 연결하며 얼굴·회복복·시설 조명의 연속성도 더 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸을 출입구 안쪽 카메라 방향으로 세우고, 두 눈을 뜬 채 화면 오른쪽의 실내 쪽을 바라본다. 시선이 향하는 구체적인 대상은 보이지 않으며, 지문도 특정 대상을 지정하지 않는다. 두 문 사이에는 세로 틈이 있지만 정지 화면만으로 개방과 폐쇄의 이동 방향을 확정할 수는 없다.",
        "built_space": "전경에 왼쪽 문짝 하나와 오른쪽 문짝 하나로 읽히는 구조가 있다. 왼쪽은 상부 유리와 넓은 불투명 하부판을 갖지만, 오른쪽은 두꺼운 세로 부재와 넓은 투명 영역 위주여서 좌우 문짝의 대응이 불명확하다. 현우는 그 뒤 문턱 부근에 있고 복도가 뒤로 이어진다. 복도에는 먼 문틀 한 조, 오른쪽 실험실 창 두 구역, 사각 천장등과 금속 하부 마감이 보여 이전 사진의 공간과 대체로 연결된다. 불가능한 인물 반사는 보이지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 있다. 앳된 얼굴, 헝클어진 검은 머리, 정상적인 홍채와 동공을 가진 열린 두 눈, 흰색 계열 회복복이 보인다. 얼굴과 목둘레 형태는 이전 사진에 대체로 부합한다. 한국계 미국인이라는 국적·배경은 외관만으로 확인할 수 없다. 연구원 카드나 판독기를 사용하는 손은 보이지 않지만 카드 제시 이후의 장면이므로 누락으로 볼 필요는 없다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "현우의 상체는 수직이고 옷은 중력에 따라 아래로 늘어진다. 발과 바닥 접촉은 프레임 밖이라 직접 확인되지 않지만 공중에 뜬 자세를 나타내는 징후는 없다. 문짝은 출입구의 수직 프레임 안에 설치된 구조로 보이며, 구동 레일은 화면 밖이다. 손에 든 물체나 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 두 눈을 뜨고 문틈을 통해 거의 카메라 정면의 실내를 바라본다. 별도의 시선 대상은 화면에 없으며 지문의 방향 요구와 충돌하지 않는다. 좌우 문짝은 중앙에서 떨어져 각각 양옆으로 물러난 배치여서 양방향 개방 순간으로 자연스럽게 읽히지만, 실제 이동 방향은 정지 화면만으로 확정되지 않는다.",
        "built_space": "전경에 두꺼운 금속 세로 가장자리와 유리, 같은 높이의 가로 금속대를 갖춘 문짝이 좌우 하나씩 있다. 두 가장자리가 화면 중간 왼쪽과 중간 오른쪽에서 현우를 감싸며, 현우는 문 뒤 중앙에 서 있다. 뒤쪽에는 먼 문틀 한 조, 오른쪽 실험실 창 두 구역, 왼쪽 장비 선반 한 대와 수납장 구역, 사각 천장등이 보인다. 금속 창틀·밝은 벽·회색 바닥은 이전 사진과 잘 이어진다. 왼쪽 수납장 구역은 참고보다 두드러지지만, 고정 설비의 명백한 중복이나 불가능한 반사는 확인되지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명이며 다른 사람은 없다. 검은 헝클어진 머리, 앳된 얼굴과 마른 체격이 참고 인물에 가깝다. 두 눈은 정상적인 인간 눈으로 열려 있고, 흰색 단추 회복복과 목둘레가 이전 장면의 차림을 따른다. 정확한 나이와 한국계 미국인 배경은 영상만으로 확정할 수 없다. 자동문 두 짝과 뒤쪽 연구시설 복도가 보이며 카드 제시 장면 자체는 포함되지 않는다. 판독 가능한 글자·로고·자막은 없다.",
        "hard_violations": [],
        "physics": "현우는 문 뒤에 수직으로 서 있고 어깨와 팔, 옷 주름이 자연스럽게 아래로 이어진다. 발은 미디엄 숏 밖에 있어 접촉점을 볼 수 없으나 부유나 비정상적인 신체 지지는 나타나지 않는다. 좌우 문짝의 금속대와 유리는 각 프레임에 연결되어 있으며 출입구에 설치된 슬라이딩 구조로 읽힌다. 들고 있는 물체나 지지 없이 떠 있는 요소는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "문틈 너머 눈을 뜬 현우와 복도는 맞지만, 좌우 문짝의 구조가 비대칭이고 왼쪽 불투명 면이 상체를 많이 가려 두꺼운 양개 자동문이 열리는 구도가 덜 명확하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "전경의 두꺼운 좌우 문짝, 그 틈의 현우 상체, 뒤쪽 복도를 미디엄 숏으로 명확히 연결하며 얼굴·회복복·시설 조명의 연속성도 더 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸을 출입구 안쪽 카메라 방향으로 세우고, 두 눈을 뜬 채 화면 오른쪽의 실내 쪽을 바라본다. 시선이 향하는 구체적인 대상은 보이지 않으며, 지문도 특정 대상을 지정하지 않는다. 두 문 사이에는 세로 틈이 있지만 정지 화면만으로 개방과 폐쇄의 이동 방향을 확정할 수는 없다.",
        "built_space": "전경에 왼쪽 문짝 하나와 오른쪽 문짝 하나로 읽히는 구조가 있다. 왼쪽은 상부 유리와 넓은 불투명 하부판을 갖지만, 오른쪽은 두꺼운 세로 부재와 넓은 투명 영역 위주여서 좌우 문짝의 대응이 불명확하다. 현우는 그 뒤 문턱 부근에 있고 복도가 뒤로 이어진다. 복도에는 먼 문틀 한 조, 오른쪽 실험실 창 두 구역, 사각 천장등과 금속 하부 마감이 보여 이전 사진의 공간과 대체로 연결된다. 불가능한 인물 반사는 보이지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 있다. 앳된 얼굴, 헝클어진 검은 머리, 정상적인 홍채와 동공을 가진 열린 두 눈, 흰색 계열 회복복이 보인다. 얼굴과 목둘레 형태는 이전 사진에 대체로 부합한다. 한국계 미국인이라는 국적·배경은 외관만으로 확인할 수 없다. 연구원 카드나 판독기를 사용하는 손은 보이지 않지만 카드 제시 이후의 장면이므로 누락으로 볼 필요는 없다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "현우의 상체는 수직이고 옷은 중력에 따라 아래로 늘어진다. 발과 바닥 접촉은 프레임 밖이라 직접 확인되지 않지만 공중에 뜬 자세를 나타내는 징후는 없다. 문짝은 출입구의 수직 프레임 안에 설치된 구조로 보이며, 구동 레일은 화면 밖이다. 손에 든 물체나 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 두 눈을 뜨고 문틈을 통해 거의 카메라 정면의 실내를 바라본다. 별도의 시선 대상은 화면에 없으며 지문의 방향 요구와 충돌하지 않는다. 좌우 문짝은 중앙에서 떨어져 각각 양옆으로 물러난 배치여서 양방향 개방 순간으로 자연스럽게 읽히지만, 실제 이동 방향은 정지 화면만으로 확정되지 않는다.",
        "built_space": "전경에 두꺼운 금속 세로 가장자리와 유리, 같은 높이의 가로 금속대를 갖춘 문짝이 좌우 하나씩 있다. 두 가장자리가 화면 중간 왼쪽과 중간 오른쪽에서 현우를 감싸며, 현우는 문 뒤 중앙에 서 있다. 뒤쪽에는 먼 문틀 한 조, 오른쪽 실험실 창 두 구역, 왼쪽 장비 선반 한 대와 수납장 구역, 사각 천장등이 보인다. 금속 창틀·밝은 벽·회색 바닥은 이전 사진과 잘 이어진다. 왼쪽 수납장 구역은 참고보다 두드러지지만, 고정 설비의 명백한 중복이나 불가능한 반사는 확인되지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명이며 다른 사람은 없다. 검은 헝클어진 머리, 앳된 얼굴과 마른 체격이 참고 인물에 가깝다. 두 눈은 정상적인 인간 눈으로 열려 있고, 흰색 단추 회복복과 목둘레가 이전 장면의 차림을 따른다. 정확한 나이와 한국계 미국인 배경은 영상만으로 확정할 수 없다. 자동문 두 짝과 뒤쪽 연구시설 복도가 보이며 카드 제시 장면 자체는 포함되지 않는다. 판독 가능한 글자·로고·자막은 없다.",
        "hard_violations": [],
        "physics": "현우는 문 뒤에 수직으로 서 있고 어깨와 팔, 옷 주름이 자연스럽게 아래로 이어진다. 발은 미디엄 숏 밖에 있어 접촉점을 볼 수 없으나 부유나 비정상적인 신체 지지는 나타나지 않는다. 좌우 문짝의 금속대와 유리는 각 프레임에 연결되어 있으며 출입구에 설치된 슬라이딩 구조로 읽힌다. 들고 있는 물체나 지지 없이 떠 있는 요소는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 1.778
   },
   "adjusted": {
    "A": 0.75,
    "B": 1.778
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible staging (좌우로 분리되어 열리는 자동문 패널 사이를 가로질러 하나의 구조물로 연결하는 수평 막대가 존재하여 물리적으로 불가능함)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1778,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1778,
    "verdict_ko": "좌우로 열리는 자동문의 틈새를 통해 두 눈을 뜬 현우의 상체를 보여주는 숏의 구도와 공간적 맥락을 물리적 모순 없이 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "열리는 자동문의 양측 패널을 수평 막대로 연결해버려 프롬프트에서 지시된 좌우 분리 동작이 물리적으로 불가능한 구조를 만들었으므로 실격입니다.  ★위반: [gemini-pro] physically impossible staging (좌우로 분리되어 열리는 자동문 패널 사이를 가로질러 하나의 구조물로 연결하는 수평 막대가 존재하여 물리적으로 불가능함)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13_sel.png",
    "asset_id": "4c4989cc-330c-4174-b09e-70e018c5938e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-f2ea-76a0-805d-cbae6bb97f73",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S81sh13"
  }
 },
 "S81sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:55:32.251917+00:00",
  "fingerprint": "2249b5fb9379e761ca91076332065b9563249673659c0315d608752aee58fb1d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S81sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S81sh16_sel.png",
  "source_sha256": "c0a1c6ad0205e28efd744ea82351a92ae504fab00e7933cb6239decad1d23101",
  "file": "S81sh16_cine.png",
  "staged_sha256": "21d14a2161d2e270355549b70e9b022405b4edcc35188b6fb1d7803efc38267b",
  "latency_ms": 10517
 },
 "S82sh14::signage": {
  "fp": "1b297fe9484dc4df",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S82sh14": {
  "input_fingerprint": "4c7e57441a25d7bc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies inactive on a stainless-steel bed inside a circular glass enclosure, with hardware damaged by prolonged immersion and multiple wires already attached to the body. Robotic arms scan the body amid computers and research equipment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies inactive on a stainless-steel bed inside a circular glass enclosure, with hardware damaged by prolonged immersion and multiple wires already attached to the body. Robotic arms scan the body amid computers and research equipment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies inactive on a stainless-steel bed inside a circular glass enclosure, with hardware damaged by prolonged immersion and multiple wires already attached to the body. Robotic arms scan the body amid computers and research equipment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14__bgfirst_bg.png",
     "asset_id": "09c1e32d-ad58-4bd4-a51a-d8b7432ecb22",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S82sh14.png",
     "asset_id": "90735080-f327-4c19-9101-6d98f10c3031",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1461211>",
     "asset_id": "24d1c172-ad7f-4d9b-9990-330c659e9ca6",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
     "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1461211>",
     "asset_id": "24d1c172-ad7f-4d9b-9990-330c659e9ca6",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "스캐닝 로봇 팔들은 찰리의 머리와 몸통을 향해 뻗어 있으며, 주변의 연구원들은 각자의 모니터를 향해 시선을 두고 있다.",
    "built_space": "원형 유리 벽으로 둘러싸인 공간 중앙에 스테인리스 침대가 카메라와 평행하게 놓여 있다. 외곽에는 다수의 컴퓨터 작업 공간과 연구원들이 배치되어 있으며, 시점은 제공된 공간 레퍼런스 사진과 동일한 눈높이 뷰를 취하고 있다.",
    "entities": "찰리는 베이지색 장갑과 흰색 마스크 등 레퍼런스의 외형과 일치한다. 스테인리스 침대, 스캐닝 기계 팔, 연구원들은 잘 나타나 있으나, 찰리의 몸에 연결되었어야 할 다수의 전선은 보이지 않는다.",
    "hard_violations": [
     "[gpt-high] 고정 장소 사진에는 독립 받침대형 스캐너가 두 대인데, 왼쪽 앞에 세 번째 스캐너와 받침대를 추가하여 시설 수를 변경했다."
    ],
    "physics": "찰리는 침대 위에 누워 있으며, 왼쪽 팔이 침대 가장자리 바깥으로 축 늘어져 중력에 완전히 순응하는 지지 상태를 보여준다."
   },
   {
    "label": "B",
    "direction": "로봇 팔들이 찰리의 흉부와 하체를 향해 레이저와 스캐너를 겨냥하고 있으며, 연구원들은 컴퓨터 화면을 응시하고 있다.",
    "built_space": "원형 유리 벽 내부를 내려다보는 하이 앵글 뷰. 스테인리스 침대가 대각선으로 배치되어 있고, 주변을 따라 작업 공간과 연구원들이 적절한 비율로 자리 잡고 있다.",
    "entities": "찰리의 외형은 레퍼런스와 일치하며, 침대, 로봇 팔, 연구원들이 모두 존재한다. 찰리의 몸 위로 다수의 굵은 전선들이 얽혀 연결되어 있다.",
    "hard_violations": [
     "[gemini-pro] 의식 없는 신체의 모든 부위가 받침대에 닿아 중력에 순응해야 한다는 규칙을 어기고, 양손(특히 멀리 있는 왼손)의 손가락과 손바닥이 허공을 향해 뻣뻣하게 들려 있음."
    ],
    "physics": "찰리의 몸통은 침대에 닿아 있으나, 양손이 지지대에 완전히 닿지 않고 뻣뻣하게 위를 향해 굳어 있어 '미동 없이 축 늘어진' 상태의 물리적 묘사에 실패했다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "'축 늘어져 누운'이라는 지문에 맞게 팔이 늘어진 자세를 완벽히 구현했으나, 몸에 연결된 다수의 전선 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지문에 명시된 전선은 잘 묘사되었으나, '축 늘어져' 있어야 할 양손이 바닥에 닿지 않고 허공에 뻣뻣하게 들려 있어 무의식 상태의 물리 법칙을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "스캐닝 로봇 팔들은 찰리의 머리와 몸통을 향해 뻗어 있으며, 주변의 연구원들은 각자의 모니터를 향해 시선을 두고 있다.",
        "built_space": "원형 유리 벽으로 둘러싸인 공간 중앙에 스테인리스 침대가 카메라와 평행하게 놓여 있다. 외곽에는 다수의 컴퓨터 작업 공간과 연구원들이 배치되어 있으며, 시점은 제공된 공간 레퍼런스 사진과 동일한 눈높이 뷰를 취하고 있다.",
        "entities": "찰리는 베이지색 장갑과 흰색 마스크 등 레퍼런스의 외형과 일치한다. 스테인리스 침대, 스캐닝 기계 팔, 연구원들은 잘 나타나 있으나, 찰리의 몸에 연결되었어야 할 다수의 전선은 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리는 침대 위에 누워 있으며, 왼쪽 팔이 침대 가장자리 바깥으로 축 늘어져 중력에 완전히 순응하는 지지 상태를 보여준다."
       },
       {
        "label": "B",
        "direction": "로봇 팔들이 찰리의 흉부와 하체를 향해 레이저와 스캐너를 겨냥하고 있으며, 연구원들은 컴퓨터 화면을 응시하고 있다.",
        "built_space": "원형 유리 벽 내부를 내려다보는 하이 앵글 뷰. 스테인리스 침대가 대각선으로 배치되어 있고, 주변을 따라 작업 공간과 연구원들이 적절한 비율로 자리 잡고 있다.",
        "entities": "찰리의 외형은 레퍼런스와 일치하며, 침대, 로봇 팔, 연구원들이 모두 존재한다. 찰리의 몸 위로 다수의 굵은 전선들이 얽혀 연결되어 있다.",
        "hard_violations": [
         "의식 없는 신체의 모든 부위가 받침대에 닿아 중력에 순응해야 한다는 규칙을 어기고, 양손(특히 멀리 있는 왼손)의 손가락과 손바닥이 허공을 향해 뻣뻣하게 들려 있음."
        ],
        "physics": "찰리의 몸통은 침대에 닿아 있으나, 양손이 지지대에 완전히 닿지 않고 뻣뻣하게 위를 향해 굳어 있어 '미동 없이 축 늘어진' 상태의 물리적 묘사에 실패했다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "'축 늘어져 누운'이라는 지문에 맞게 팔이 늘어진 자세를 완벽히 구현했으나, 몸에 연결된 다수의 전선 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지문에 명시된 전선은 잘 묘사되었으나, '축 늘어져' 있어야 할 양손이 바닥에 닿지 않고 허공에 뻣뻣하게 들려 있어 무의식 상태의 물리 법칙을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "스캐닝 로봇 팔들은 찰리의 머리와 몸통을 향해 뻗어 있으며, 주변의 연구원들은 각자의 모니터를 향해 시선을 두고 있다.",
        "built_space": "원형 유리 벽으로 둘러싸인 공간 중앙에 스테인리스 침대가 카메라와 평행하게 놓여 있다. 외곽에는 다수의 컴퓨터 작업 공간과 연구원들이 배치되어 있으며, 시점은 제공된 공간 레퍼런스 사진과 동일한 눈높이 뷰를 취하고 있다.",
        "entities": "찰리는 베이지색 장갑과 흰색 마스크 등 레퍼런스의 외형과 일치한다. 스테인리스 침대, 스캐닝 기계 팔, 연구원들은 잘 나타나 있으나, 찰리의 몸에 연결되었어야 할 다수의 전선은 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리는 침대 위에 누워 있으며, 왼쪽 팔이 침대 가장자리 바깥으로 축 늘어져 중력에 완전히 순응하는 지지 상태를 보여준다."
       },
       {
        "label": "B",
        "direction": "로봇 팔들이 찰리의 흉부와 하체를 향해 레이저와 스캐너를 겨냥하고 있으며, 연구원들은 컴퓨터 화면을 응시하고 있다.",
        "built_space": "원형 유리 벽 내부를 내려다보는 하이 앵글 뷰. 스테인리스 침대가 대각선으로 배치되어 있고, 주변을 따라 작업 공간과 연구원들이 적절한 비율로 자리 잡고 있다.",
        "entities": "찰리의 외형은 레퍼런스와 일치하며, 침대, 로봇 팔, 연구원들이 모두 존재한다. 찰리의 몸 위로 다수의 굵은 전선들이 얽혀 연결되어 있다.",
        "hard_violations": [
         "의식 없는 신체의 모든 부위가 받침대에 닿아 중력에 순응해야 한다는 규칙을 어기고, 양손(특히 멀리 있는 왼손)의 손가락과 손바닥이 허공을 향해 뻣뻣하게 들려 있음."
        ],
        "physics": "찰리의 몸통은 침대에 닿아 있으나, 양손이 지지대에 완전히 닿지 않고 뻣뻣하게 위를 향해 굳어 있어 '미동 없이 축 늘어진' 상태의 물리적 묘사에 실패했다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "높은 시점에서 대각선으로 놓인 침대와 찰리의 전신을 유리 너머로 보여주며, 장소 기준의 스캐너 두 대도 유지하지만 팔의 육중한 비율과 침수 손상은 다소 약하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "유리 격벽과 찰리의 정지 상태는 구현했으나 기준 장소에 없는 세 번째 스캐너를 추가했고, 높은 대각선 조망보다 정면에 가까운 구도로 바뀌었다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 흰 얼굴은 천장 쪽을 향하고 시선이나 능동적인 동작은 드러나지 않는다. 왼쪽 스캐너 끝은 머리·상체 쪽으로, 오른쪽 스캐너 끝은 몸통 쪽으로 향하지만 끝단들이 몸에 가까이 접근하지는 않는다. 주변 연구원들은 각자의 모니터나 작업면을 바라본다.",
        "built_space": "원형 유리 격벽 하나 안에 중앙 스테인리스 침대 한 대와 독립 받침대형 스캐너 두 대가 있다. 침대의 긴 축은 화면 왼쪽 아래에서 오른쪽 위로 뻗으며, 높은 외부 시점에서 가까운 유리와 세로 이음부를 통과해 전신이 보인다. 격벽 바깥에는 약 일곱 명의 연구원이 서로 다른 작업대에 배치되어 있다. 금속 바닥, 원형 배수구, 주변 작업대와 벽면 화면은 장소 사진의 구성을 대체로 유지한다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 한 개체이며 샌드 베이지 각진 장갑, 흰 마스크형 얼굴, 검은 관절, 원형 가슴 부품이 기준과 맞는다. 다만 기준의 특히 긴 팔과 짧은 다리로 이루어진 고릴라형 비율보다 다소 일반적인 인간형에 가깝다. 장갑에 마모가 있고 여러 케이블이 몸 주변과 침대에 연결되어 있으나 장기 침수에 따른 손상은 뚜렷하지 않다. 연구원들은 배경 지시에서 요구한 백의 차림의 성인들로 보이며, 작은 크기 때문에 개별 나이·민족성은 확정하기 어렵다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 몸통과 다리는 침대 위에 놓이고 양팔과 손도 몸 옆의 침대 면에 내려앉아 있다. 발은 침대 끝부분에 걸쳐 지지된다. 공중에 따로 들고 있는 손이나 물건은 보이지 않는다. 침대는 중앙 금속 받침으로, 스캐너들은 바닥에 닿은 받침대와 관절로 지지된다. 케이블은 기계와 침대 사이에서 처지거나 바닥에 놓인다. 연구원들은 의자에 앉거나 바닥에 서 있다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 위쪽을 향한다. 뒤쪽 스캐너는 머리·어깨 쪽으로 내려오고, 왼쪽 앞 스캐너는 늘어진 팔 쪽을 향한다. 오른쪽 스캐너도 몸 안쪽을 향하나 끝단은 몸에서 떨어져 있다. 배경과 오른쪽 전경의 연구원들은 각자의 컴퓨터를 향하며, 전경 모니터의 화면 방향도 사용자 위치와 맞는다.",
        "built_space": "원형 유리 격벽 안에 침대 한 대가 있고, 스캐너 받침대는 뒤쪽 하나, 오른쪽 하나, 왼쪽 앞 하나로 총 세 대다. 장소 사진의 두 대 구성에 한 대가 추가되었다. 카메라는 유리 바깥에서 거의 정면으로 침대 발치를 바라보며 약간 내려다보지만, 요구된 높은 대각선 조망은 약하다. 외곽과 오른쪽 전경에는 약 여덟 명이 별도 컴퓨터를 사용한다. 원형 천장 조명, 금속 바닥과 배수구, 문틀과 주변 작업대는 기준 장소를 잘 따른다. 명백한 불가능 반사는 보이지 않는다.",
        "entities": "찰리 한 개체의 베이지 장갑, 흰 얼굴, 검은 기계 관절과 가슴 원형 부품이 보인다. 큰 손과 긴 팔은 기준의 육중한 체형을 비교적 잘 드러낸다. 여러 선이 몸 주변으로 연결되고 장갑에는 얼룩과 마모가 있지만 침수 손상의 원인은 시각적으로 확정할 수 없다. 백의 차림 성인 연구원들이 보이며, 오른쪽 전경 인물은 다른 연구원보다 크게 잡혀 시선을 분산한다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [
         "고정 장소 사진에는 독립 받침대형 스캐너가 두 대인데, 왼쪽 앞에 세 번째 스캐너와 받침대를 추가하여 시설 수를 변경했다."
        ],
        "physics": "찰리의 몸통과 골반은 침대에 기대어 있고 다리는 침대 위에 놓여 있다. 왼쪽 화면의 팔은 침대 밖으로 중력 방향으로 늘어지며 어깨에 연결되어 있으므로 공중 부유로 보이지 않는다. 반대쪽 팔도 몸과 침대 가장자리 쪽으로 내려와 있다. 침대와 세 스캐너는 각각 바닥 받침이 있으며 케이블도 연결점 사이에서 자연스럽게 처진다. 연구원들의 앉은 자세는 의자와 작업대의 위치에 맞는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "높은 시점에서 대각선으로 놓인 침대와 찰리의 전신을 유리 너머로 보여주며, 장소 기준의 스캐너 두 대도 유지하지만 팔의 육중한 비율과 침수 손상은 다소 약하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "유리 격벽과 찰리의 정지 상태는 구현했으나 기준 장소에 없는 세 번째 스캐너를 추가했고, 높은 대각선 조망보다 정면에 가까운 구도로 바뀌었다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 흰 얼굴은 천장 쪽을 향하고 시선이나 능동적인 동작은 드러나지 않는다. 왼쪽 스캐너 끝은 머리·상체 쪽으로, 오른쪽 스캐너 끝은 몸통 쪽으로 향하지만 끝단들이 몸에 가까이 접근하지는 않는다. 주변 연구원들은 각자의 모니터나 작업면을 바라본다.",
        "built_space": "원형 유리 격벽 하나 안에 중앙 스테인리스 침대 한 대와 독립 받침대형 스캐너 두 대가 있다. 침대의 긴 축은 화면 왼쪽 아래에서 오른쪽 위로 뻗으며, 높은 외부 시점에서 가까운 유리와 세로 이음부를 통과해 전신이 보인다. 격벽 바깥에는 약 일곱 명의 연구원이 서로 다른 작업대에 배치되어 있다. 금속 바닥, 원형 배수구, 주변 작업대와 벽면 화면은 장소 사진의 구성을 대체로 유지한다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 한 개체이며 샌드 베이지 각진 장갑, 흰 마스크형 얼굴, 검은 관절, 원형 가슴 부품이 기준과 맞는다. 다만 기준의 특히 긴 팔과 짧은 다리로 이루어진 고릴라형 비율보다 다소 일반적인 인간형에 가깝다. 장갑에 마모가 있고 여러 케이블이 몸 주변과 침대에 연결되어 있으나 장기 침수에 따른 손상은 뚜렷하지 않다. 연구원들은 배경 지시에서 요구한 백의 차림의 성인들로 보이며, 작은 크기 때문에 개별 나이·민족성은 확정하기 어렵다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 몸통과 다리는 침대 위에 놓이고 양팔과 손도 몸 옆의 침대 면에 내려앉아 있다. 발은 침대 끝부분에 걸쳐 지지된다. 공중에 따로 들고 있는 손이나 물건은 보이지 않는다. 침대는 중앙 금속 받침으로, 스캐너들은 바닥에 닿은 받침대와 관절로 지지된다. 케이블은 기계와 침대 사이에서 처지거나 바닥에 놓인다. 연구원들은 의자에 앉거나 바닥에 서 있다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 위쪽을 향한다. 뒤쪽 스캐너는 머리·어깨 쪽으로 내려오고, 왼쪽 앞 스캐너는 늘어진 팔 쪽을 향한다. 오른쪽 스캐너도 몸 안쪽을 향하나 끝단은 몸에서 떨어져 있다. 배경과 오른쪽 전경의 연구원들은 각자의 컴퓨터를 향하며, 전경 모니터의 화면 방향도 사용자 위치와 맞는다.",
        "built_space": "원형 유리 격벽 안에 침대 한 대가 있고, 스캐너 받침대는 뒤쪽 하나, 오른쪽 하나, 왼쪽 앞 하나로 총 세 대다. 장소 사진의 두 대 구성에 한 대가 추가되었다. 카메라는 유리 바깥에서 거의 정면으로 침대 발치를 바라보며 약간 내려다보지만, 요구된 높은 대각선 조망은 약하다. 외곽과 오른쪽 전경에는 약 여덟 명이 별도 컴퓨터를 사용한다. 원형 천장 조명, 금속 바닥과 배수구, 문틀과 주변 작업대는 기준 장소를 잘 따른다. 명백한 불가능 반사는 보이지 않는다.",
        "entities": "찰리 한 개체의 베이지 장갑, 흰 얼굴, 검은 기계 관절과 가슴 원형 부품이 보인다. 큰 손과 긴 팔은 기준의 육중한 체형을 비교적 잘 드러낸다. 여러 선이 몸 주변으로 연결되고 장갑에는 얼룩과 마모가 있지만 침수 손상의 원인은 시각적으로 확정할 수 없다. 백의 차림 성인 연구원들이 보이며, 오른쪽 전경 인물은 다른 연구원보다 크게 잡혀 시선을 분산한다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [
         "고정 장소 사진에는 독립 받침대형 스캐너가 두 대인데, 왼쪽 앞에 세 번째 스캐너와 받침대를 추가하여 시설 수를 변경했다."
        ],
        "physics": "찰리의 몸통과 골반은 침대에 기대어 있고 다리는 침대 위에 놓여 있다. 왼쪽 화면의 팔은 침대 밖으로 중력 방향으로 늘어지며 어깨에 연결되어 있으므로 공중 부유로 보이지 않는다. 반대쪽 팔도 몸과 침대 가장자리 쪽으로 내려와 있다. 침대와 세 스캐너는 각각 바닥 받침이 있으며 케이블도 연결점 사이에서 자연스럽게 처진다. 연구원들의 앉은 자세는 의자와 작업대의 위치에 맞는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.125,
    "B": 1.321
   },
   "violations": {
    "B": [
     "[gemini-pro] 의식 없는 신체의 모든 부위가 받침대에 닿아 중력에 순응해야 한다는 규칙을 어기고, 양손(특히 멀리 있는 왼손)의 손가락과 손바닥이 허공을 향해 뻣뻣하게 들려 있음."
    ],
    "A": [
     "[gpt-high] 고정 장소 사진에는 독립 받침대형 스캐너가 두 대인데, 왼쪽 앞에 세 번째 스캐너와 받침대를 추가하여 시설 수를 변경했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1125,
   "B": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1125,
    "verdict_ko": "'축 늘어져 누운'이라는 지문에 맞게 팔이 늘어진 자세를 완벽히 구현했으나, 몸에 연결된 다수의 전선 묘사가 누락되었습니다.  ★위반: [gpt-high] 고정 장소 사진에는 독립 받침대형 스캐너가 두 대인데, 왼쪽 앞에 세 번째 스캐너와 받침대를 추가하여 시설 수를 변경했다."
   },
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "지문에 명시된 전선은 잘 묘사되었으나, '축 늘어져' 있어야 할 양손이 바닥에 닿지 않고 허공에 뻣뻣하게 들려 있어 무의식 상태의 물리 법칙을 위반했습니다.  ★위반: [gemini-pro] 의식 없는 신체의 모든 부위가 받침대에 닿아 중력에 순응해야 한다는 규칙을 어기고, 양손(특히 멀리 있는 왼손)의 손가락과 손바닥이 허공을 향해 뻣뻣하게 들려 있음."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
    "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1461211>",
    "asset_id": "24d1c172-ad7f-4d9b-9990-330c659e9ca6",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-f490-7774-9f8f-e765a55f7708",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14__bgfirst_bg.png",
   "bg_asset_id": "09c1e32d-ad58-4bd4-a51a-d8b7432ecb22",
   "bg_record_key": "S82sh14::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S82sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:57:04.331858+00:00",
  "fingerprint": "18c5efca2e7bd5d8b1f273f83634a5fc77c88fc192906b9768c9698a97aec0ba",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S82sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S82sh14_sel.png",
  "source_sha256": "a4cd29df89452e9b57571d401eec5efe21ab4686782a148de9220f0ffd225254",
  "file": "S82sh14_cine.png",
  "staged_sha256": "8d6c50148c43e9f28efd1cfa062adbf3a284b96a4c1bf454e635d13487506d54",
  "latency_ms": 10174
 },
 "S82sh17::signage": {
  "fp": "6c736eeca5a08bea",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S82sh17": {
  "input_fingerprint": "027ed8b07417bc28",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 차가운 유리 벽에 양손을 바짝 대고 찰리를 뚫어져라 내려다보는 현우의 절박한 얼굴.\n\nLOCATION (lock): At the observation side of the circular glass laboratory enclosure, with cool light over the examination bed beyond. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Glass barrier between 현우 and 찰리 in the middle-center of the frame, midground; Bed supporting 찰리 beyond the glass in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Separates 현우's pressed hands from 찰리) — Seen obliquely along its curvature, with 찰리 and part of the bed visible through it; used as Physical separation within the intimate composition; Stainless-steel bed (Partially visible beneath 찰리) — A small portion of the near side appears beyond the glass at lower right; used as Anchors the downward eyeline and the body's depth.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the laboratory's ambient tonal balance, allowing clear transmission through the glass rather than emphasizing reflected light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, circular glass enclosure, scanning equipment, and cool laboratory lighting. Exclude furnishings from the separate shelter bedroom.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains inactive on the stainless-steel bed within the circular glass enclosure, with immersion-damaged hardware and multiple attached wires. The surrounding robotic arms continue the scanning procedure. 현우: He remains in white recovery clothes, close to the laboratory's glass enclosure and visibly worried.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 차가운 유리 벽에 양손을 바짝 대고 찰리를 뚫어져라 내려다보는 현우의 절박한 얼굴.\n\nLOCATION (lock): At the observation side of the circular glass laboratory enclosure, with cool light over the examination bed beyond. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Glass barrier between 현우 and 찰리 in the middle-center of the frame, midground; Bed supporting 찰리 beyond the glass in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Separates 현우's pressed hands from 찰리) — Seen obliquely along its curvature, with 찰리 and part of the bed visible through it; used as Physical separation within the intimate composition; Stainless-steel bed (Partially visible beneath 찰리) — A small portion of the near side appears beyond the glass at lower right; used as Anchors the downward eyeline and the body's depth.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the laboratory's ambient tonal balance, allowing clear transmission through the glass rather than emphasizing reflected light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, circular glass enclosure, scanning equipment, and cool laboratory lighting. Exclude furnishings from the separate shelter bedroom.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains inactive on the stainless-steel bed within the circular glass enclosure, with immersion-damaged hardware and multiple attached wires. The surrounding robotic arms continue the scanning procedure. 현우: He remains in white recovery clothes, close to the laboratory's glass enclosure and visibly worried.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 차가운 유리 벽에 양손을 바짝 대고 찰리를 뚫어져라 내려다보는 현우의 절박한 얼굴.\n\nLOCATION (lock): At the observation side of the circular glass laboratory enclosure, with cool light over the examination bed beyond. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Glass barrier between 현우 and 찰리 in the middle-center of the frame, midground; Bed supporting 찰리 beyond the glass in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Separates 현우's pressed hands from 찰리) — Seen obliquely along its curvature, with 찰리 and part of the bed visible through it; used as Physical separation within the intimate composition; Stainless-steel bed (Partially visible beneath 찰리) — A small portion of the near side appears beyond the glass at lower right; used as Anchors the downward eyeline and the body's depth.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the laboratory's ambient tonal balance, allowing clear transmission through the glass rather than emphasizing reflected light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, circular glass enclosure, scanning equipment, and cool laboratory lighting. Exclude furnishings from the separate shelter bedroom.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains inactive on the stainless-steel bed within the circular glass enclosure, with immersion-damaged hardware and multiple attached wires. The surrounding robotic arms continue the scanning procedure. 현우: He remains in white recovery clothes, close to the laboratory's glass enclosure and visibly worried.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우가 아래를 내려다봄.",
    "built_space": "유리벽 너머 넓은 연구실 공간이 보이며, 이전 샷의 넓은 구도를 그대로 모방함.",
    "entities": "현우와 찰리 외에 다수의 연구원들이 배경에 등장함.",
    "hard_violations": [
     "[gemini-pro] extra bodies (배경의 다수 연구원들)",
     "[gemini-pro] physically impossible staging (하반신 없이 공중에 떠 있는 현우, 전면에 중복 배치된 거대한 찰리 신체)",
     "[gpt-high] 제외하라고 지정된 이전 장면의 연구원 일곱 명이 그대로 등장합니다.",
     "[gpt-high] 오른쪽 전경의 찰리 외에 중앙 침대에도 찰리의 손·팔과 발이 남아 중복 신체가 발생합니다.",
     "[gpt-high] 원래 중앙 침대 외에 전경 찰리 아래 별도의 금속 받침이 추가되어 고정 침대의 단일 배치를 깨뜨립니다.",
     "[gpt-high] 현우의 상체가 중앙 침대와 겹친 채 하단에서 끊겨, 유리벽 바깥에 서 있는 사람으로 성립하지 않는 합성형 배치입니다."
    ],
    "physics": "현우의 하반신이 없고 허공에 떠 있어 어떤 지지대도 확인할 수 없음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 유리창 너머 아래쪽의 찰리를 향함.",
    "built_space": "유리벽이 화면 중경에 배치되고, 우측 하단에 침대가 있음.",
    "entities": "현우와 찰리가 등장하나, 배경에 금지된 인물 1명이 희미하게 보임.",
    "hard_violations": [
     "[gemini-pro] extra bodies (배경의 인물)",
     "[gemini-pro] physically impossible staging (미니어처 크기로 심각하게 축소된 찰리와 침대)",
     "[gpt-high] 현우와 찰리 외에는 등장시키지 말라는 지시와 달리, 배경 작업대에 연구원 한 명이 보입니다."
    ],
    "physics": "현우의 손은 유리에 지탱되며, 찰리는 침대에 누워 있으나 전체적인 크기 비율이 현실과 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "요구된 클로즈업 구도를 따랐으나, 배경에 금지된 인물이 추가되었고 찰리의 크기가 미니어처처럼 축소되는 치명적 오류가 있음."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 클로즈업을 무시하고 넓은 화각을 그렸으며, 공중에 뜬 현우와 다수의 금지된 배경 인물들로 인해 사용할 수 없음."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 유리창 너머 아래쪽의 찰리를 향함.",
        "built_space": "유리벽이 화면 중경에 배치되고, 우측 하단에 침대가 있음.",
        "entities": "현우와 찰리가 등장하나, 배경에 금지된 인물 1명이 희미하게 보임.",
        "hard_violations": [
         "extra bodies (배경의 인물)",
         "physically impossible staging (미니어처 크기로 심각하게 축소된 찰리와 침대)"
        ],
        "physics": "현우의 손은 유리에 지탱되며, 찰리는 침대에 누워 있으나 전체적인 크기 비율이 현실과 맞지 않음."
       },
       {
        "label": "A",
        "direction": "현우가 아래를 내려다봄.",
        "built_space": "유리벽 너머 넓은 연구실 공간이 보이며, 이전 샷의 넓은 구도를 그대로 모방함.",
        "entities": "현우와 찰리 외에 다수의 연구원들이 배경에 등장함.",
        "hard_violations": [
         "extra bodies (배경의 다수 연구원들)",
         "physically impossible staging (하반신 없이 공중에 떠 있는 현우, 전면에 중복 배치된 거대한 찰리 신체)"
        ],
        "physics": "현우의 하반신이 없고 허공에 떠 있어 어떤 지지대도 확인할 수 없음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "요구된 클로즈업 구도를 따랐으나, 배경에 금지된 인물이 추가되었고 찰리의 크기가 미니어처처럼 축소되는 치명적 오류가 있음."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 클로즈업을 무시하고 넓은 화각을 그렸으며, 공중에 뜬 현우와 다수의 금지된 배경 인물들로 인해 사용할 수 없음."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 유리창 너머 아래쪽의 찰리를 향함.",
        "built_space": "유리벽이 화면 중경에 배치되고, 우측 하단에 침대가 있음.",
        "entities": "현우와 찰리가 등장하나, 배경에 금지된 인물 1명이 희미하게 보임.",
        "hard_violations": [
         "extra bodies (배경의 인물)",
         "physically impossible staging (미니어처 크기로 심각하게 축소된 찰리와 침대)"
        ],
        "physics": "현우의 손은 유리에 지탱되며, 찰리는 침대에 누워 있으나 전체적인 크기 비율이 현실과 맞지 않음."
       },
       {
        "label": "A",
        "direction": "현우가 아래를 내려다봄.",
        "built_space": "유리벽 너머 넓은 연구실 공간이 보이며, 이전 샷의 넓은 구도를 그대로 모방함.",
        "entities": "현우와 찰리 외에 다수의 연구원들이 배경에 등장함.",
        "hard_violations": [
         "extra bodies (배경의 다수 연구원들)",
         "physically impossible staging (하반신 없이 공중에 떠 있는 현우, 전면에 중복 배치된 거대한 찰리 신체)"
        ],
        "physics": "현우의 하반신이 없고 허공에 떠 있어 어떤 지지대도 확인할 수 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 양손의 유리 접촉, 찰리를 향한 하향 시선은 더 충실하지만, 금지된 연구원이 남아 있고 찰리와 침대가 작은 배경 부분보다 크게 노출됩니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "이전 와이드숏에 현우와 별도의 찰리를 합성한 듯한 배치로, 연구원 다수와 중복 신체·침대가 남고 현우의 몸도 침대에서 부자연스럽게 끊깁니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈은 화면 오른쪽 아래 침대에 누운 찰리 쪽을 향합니다. 양손은 손바닥을 유리 쪽으로 펼쳤고 얼굴도 유리 가까이 기울였습니다. 스캐너 끝은 찰리가 있는 침대 쪽을 향합니다.",
        "built_space": "곡선 유리벽 하나와 여러 세로 이음부, 스테인리스 침대 하나가 보입니다. 스캐너는 중앙 오른쪽에 한 대, 오른쪽 가장자리에 다른 한 대의 끝부분이 보입니다. 현우는 유리 너머, 찰리는 침대 위에 배치되어 분리 관계는 읽힙니다. 다만 침대와 찰리 거의 전신이 오른쪽 아래를 크게 차지해 요청한 작은 배경 노출과 다릅니다. 뒤쪽 작업대에는 제외되어야 할 연구원 한 명이 보입니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 검은 머리와 흰 회복복, 참고 이미지와 유사한 얼굴을 갖췄습니다. 표정은 걱정스럽지만 절박함은 비교적 약합니다. 찰리는 샌드 베이지 장갑, 흰 마스크형 얼굴, 긴 기계 팔을 지닌 참고 이미지의 로봇으로 식별됩니다. 연결 전선과 금속 침대, 스캔 장비도 보입니다. 허용되지 않은 흰옷의 연구원이 추가로 남아 있습니다.",
        "hard_violations": [
         "현우와 찰리 외에는 등장시키지 말라는 지시와 달리, 배경 작업대에 연구원 한 명이 보입니다."
        ],
        "physics": "찰리의 몸통과 팔다리는 침대 표면에 놓여 있으며 떠 있는 신체 부분은 보이지 않습니다. 전선은 몸과 장비에서 내려와 침대와 바닥에 걸쳐 있습니다. 현우의 팔은 어깨와 연결되어 있고 양손을 유리에 대는 자세는 가능합니다. 하체는 클로즈업 밖에 있어 발의 지지 상태는 확인할 수 없습니다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 아래 전경의 찰리 쪽을 내려다보며 양손을 앞으로 펼칩니다. 그러나 손바닥들이 동일한 유리면에 밀착되었다는 접촉 표현은 불명확합니다. 오른쪽 스캐너는 왼쪽의 원래 침대 쪽을 향해, 새로 크게 배치된 전경 찰리와 스캔 대상의 위치가 어긋납니다.",
        "built_space": "이전 이미지의 원형 유리실과 높은 와이드 시점이 거의 그대로 남아 있습니다. 중앙에 원래 침대 하나가 있고 오른쪽 전경에는 별도의 찰리를 받치는 금속 구조가 추가되어 침대가 중복된 배치로 읽힙니다. 스캐너는 오른쪽 한 대와 현우 뒤에 가려진 왼쪽 장비 일부가 보입니다. 현우는 관찰 구역 바깥이 아니라 중앙 침대와 겹쳐 배치되어 있습니다. 둘레 작업대에는 연구원 일곱 명이 그대로 남아 있습니다.",
        "entities": "현우의 검은 머리, 젊은 동아시아계 남성 얼굴과 흰 회복복은 참고 이미지와 대체로 맞습니다. 오른쪽 전경 찰리도 베이지 장갑과 흰 기계 얼굴을 갖췄지만, 중앙 침대에는 별도의 찰리 손·팔과 발 부분이 남아 신체가 중복됩니다. 허용되지 않은 연구원 일곱 명도 등장합니다. 현우 얼굴 중심의 클로즈업이 아니라 실험실 전체를 보여주는 구도입니다.",
        "hard_violations": [
         "제외하라고 지정된 이전 장면의 연구원 일곱 명이 그대로 등장합니다.",
         "오른쪽 전경의 찰리 외에 중앙 침대에도 찰리의 손·팔과 발이 남아 중복 신체가 발생합니다.",
         "원래 중앙 침대 외에 전경 찰리 아래 별도의 금속 받침이 추가되어 고정 침대의 단일 배치를 깨뜨립니다.",
         "현우의 상체가 중앙 침대와 겹친 채 하단에서 끊겨, 유리벽 바깥에 서 있는 사람으로 성립하지 않는 합성형 배치입니다."
        ],
        "physics": "전경 찰리의 몸통 아래에는 금속 받침이 보이지만, 중앙 침대에 남은 기계 손과 발은 그 몸통과 연결되지 않습니다. 현우는 팔과 손 자체는 해부학적으로 연결되어 있으나, 상체가 침대 위에서 잘린 듯 끝나고 하체나 앉은 접촉면이 없어 몸의 지지가 성립하지 않습니다. 이는 정상적인 화면 밖 하체 생략과 달리 프레임 내부에서 몸이 소실되는 형태입니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 양손의 유리 접촉, 찰리를 향한 하향 시선은 더 충실하지만, 금지된 연구원이 남아 있고 찰리와 침대가 작은 배경 부분보다 크게 노출됩니다."
       },
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "이전 와이드숏에 현우와 별도의 찰리를 합성한 듯한 배치로, 연구원 다수와 중복 신체·침대가 남고 현우의 몸도 침대에서 부자연스럽게 끊깁니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈은 화면 오른쪽 아래 침대에 누운 찰리 쪽을 향합니다. 양손은 손바닥을 유리 쪽으로 펼쳤고 얼굴도 유리 가까이 기울였습니다. 스캐너 끝은 찰리가 있는 침대 쪽을 향합니다.",
        "built_space": "곡선 유리벽 하나와 여러 세로 이음부, 스테인리스 침대 하나가 보입니다. 스캐너는 중앙 오른쪽에 한 대, 오른쪽 가장자리에 다른 한 대의 끝부분이 보입니다. 현우는 유리 너머, 찰리는 침대 위에 배치되어 분리 관계는 읽힙니다. 다만 침대와 찰리 거의 전신이 오른쪽 아래를 크게 차지해 요청한 작은 배경 노출과 다릅니다. 뒤쪽 작업대에는 제외되어야 할 연구원 한 명이 보입니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 검은 머리와 흰 회복복, 참고 이미지와 유사한 얼굴을 갖췄습니다. 표정은 걱정스럽지만 절박함은 비교적 약합니다. 찰리는 샌드 베이지 장갑, 흰 마스크형 얼굴, 긴 기계 팔을 지닌 참고 이미지의 로봇으로 식별됩니다. 연결 전선과 금속 침대, 스캔 장비도 보입니다. 허용되지 않은 흰옷의 연구원이 추가로 남아 있습니다.",
        "hard_violations": [
         "현우와 찰리 외에는 등장시키지 말라는 지시와 달리, 배경 작업대에 연구원 한 명이 보입니다."
        ],
        "physics": "찰리의 몸통과 팔다리는 침대 표면에 놓여 있으며 떠 있는 신체 부분은 보이지 않습니다. 전선은 몸과 장비에서 내려와 침대와 바닥에 걸쳐 있습니다. 현우의 팔은 어깨와 연결되어 있고 양손을 유리에 대는 자세는 가능합니다. 하체는 클로즈업 밖에 있어 발의 지지 상태는 확인할 수 없습니다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 아래 전경의 찰리 쪽을 내려다보며 양손을 앞으로 펼칩니다. 그러나 손바닥들이 동일한 유리면에 밀착되었다는 접촉 표현은 불명확합니다. 오른쪽 스캐너는 왼쪽의 원래 침대 쪽을 향해, 새로 크게 배치된 전경 찰리와 스캔 대상의 위치가 어긋납니다.",
        "built_space": "이전 이미지의 원형 유리실과 높은 와이드 시점이 거의 그대로 남아 있습니다. 중앙에 원래 침대 하나가 있고 오른쪽 전경에는 별도의 찰리를 받치는 금속 구조가 추가되어 침대가 중복된 배치로 읽힙니다. 스캐너는 오른쪽 한 대와 현우 뒤에 가려진 왼쪽 장비 일부가 보입니다. 현우는 관찰 구역 바깥이 아니라 중앙 침대와 겹쳐 배치되어 있습니다. 둘레 작업대에는 연구원 일곱 명이 그대로 남아 있습니다.",
        "entities": "현우의 검은 머리, 젊은 동아시아계 남성 얼굴과 흰 회복복은 참고 이미지와 대체로 맞습니다. 오른쪽 전경 찰리도 베이지 장갑과 흰 기계 얼굴을 갖췄지만, 중앙 침대에는 별도의 찰리 손·팔과 발 부분이 남아 신체가 중복됩니다. 허용되지 않은 연구원 일곱 명도 등장합니다. 현우 얼굴 중심의 클로즈업이 아니라 실험실 전체를 보여주는 구도입니다.",
        "hard_violations": [
         "제외하라고 지정된 이전 장면의 연구원 일곱 명이 그대로 등장합니다.",
         "오른쪽 전경의 찰리 외에 중앙 침대에도 찰리의 손·팔과 발이 남아 중복 신체가 발생합니다.",
         "원래 중앙 침대 외에 전경 찰리 아래 별도의 금속 받침이 추가되어 고정 침대의 단일 배치를 깨뜨립니다.",
         "현우의 상체가 중앙 침대와 겹친 채 하단에서 끊겨, 유리벽 바깥에 서 있는 사람으로 성립하지 않는 합성형 배치입니다."
        ],
        "physics": "전경 찰리의 몸통 아래에는 금속 받침이 보이지만, 중앙 침대에 남은 기계 손과 발은 그 몸통과 연결되지 않습니다. 현우는 팔과 손 자체는 해부학적으로 연결되어 있으나, 상체가 침대 위에서 잘린 듯 끝나고 하체나 앉은 접촉면이 없어 몸의 지지가 성립하지 않습니다. 이는 정상적인 화면 밖 하체 생략과 달리 프레임 내부에서 몸이 소실되는 형태입니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.75,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.5,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gemini-pro] extra bodies (배경의 인물)",
     "[gemini-pro] physically impossible staging (미니어처 크기로 심각하게 축소된 찰리와 침대)",
     "[gpt-high] 현우와 찰리 외에는 등장시키지 말라는 지시와 달리, 배경 작업대에 연구원 한 명이 보입니다."
    ],
    "A": [
     "[gemini-pro] extra bodies (배경의 다수 연구원들)",
     "[gemini-pro] physically impossible staging (하반신 없이 공중에 떠 있는 현우, 전면에 중복 배치된 거대한 찰리 신체)",
     "[gpt-high] 제외하라고 지정된 이전 장면의 연구원 일곱 명이 그대로 등장합니다.",
     "[gpt-high] 오른쪽 전경의 찰리 외에 중앙 침대에도 찰리의 손·팔과 발이 남아 중복 신체가 발생합니다.",
     "[gpt-high] 원래 중앙 침대 외에 전경 찰리 아래 별도의 금속 받침이 추가되어 고정 침대의 단일 배치를 깨뜨립니다.",
     "[gpt-high] 현우의 상체가 중앙 침대와 겹친 채 하단에서 끊겨, 유리벽 바깥에 서 있는 사람으로 성립하지 않는 합성형 배치입니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "요구된 클로즈업 구도를 따랐으나, 배경에 금지된 인물이 추가되었고 찰리의 크기가 미니어처처럼 축소되는 치명적 오류가 있음.  ★위반: [gemini-pro] extra bodies (배경의 인물) / [gemini-pro] physically impossible staging (미니어처 크기로 심각하게 축소된 찰리와 침대) / [gpt-high] 현우와 찰리 외에는 등장시키지 말라는 지시와 달리, 배경 작업대에 연구원 한 명이 보입니다."
   },
   {
    "label": "A",
    "score": 500,
    "verdict_ko": "지정된 클로즈업을 무시하고 넓은 화각을 그렸으며, 공중에 뜬 현우와 다수의 금지된 배경 인물들로 인해 사용할 수 없음.  ★위반: [gemini-pro] extra bodies (배경의 다수 연구원들) / [gemini-pro] physically impossible staging (하반신 없이 공중에 떠 있는 현우, 전면에 중복 배치된 거대한 찰리 신체) / [gpt-high] 제외하라고 지정된 이전 장면의 연구원 일곱 명이 그대로 등장합니다. / [gpt-high] 오른쪽 전경의 찰리 외에 중앙 침대에도 찰리의 손·팔과 발이 남아 중복 신체가 발생합니다. / [gpt-high] 원래 중앙 침대 외에 전경 찰리 아래 별도의 금속 받침이 추가되어 고정 침대의 단일 배치를 깨뜨립니다. / [gpt-high] 현우의 상체가 중앙 침대와 겹친 채 하단에서 끊겨, 유리벽 바깥에 서 있는 사람으로 성립하지 않는 합성형 배치입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14_sel.png",
    "asset_id": "c78997e0-2898-4ac1-b041-54b98246752d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1461211>",
    "asset_id": "24d1c172-ad7f-4d9b-9990-330c659e9ca6",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-f7e2-77f1-909b-41de96958350",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh14"
  },
  "locked_char_refs_kept_body_identity": [
   "찰리(C06)"
  ],
  "staged_characters_added": [
   "C06"
  ]
 },
 "S82sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T19:58:15.438692+00:00",
  "fingerprint": "7d030c4fc99d74b2a47916768b0828ada5ea0aa1645c0137f6eed458f1d627a1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S82sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S82sh17_sel.png",
  "source_sha256": "248c2c9239866f0caef77a38040c6b073b8b67e28fe81d8a89f8b9b18a7ea514",
  "file": "S82sh17_cine.png",
  "staged_sha256": "9c6f81bb7959c7d5926b45370b2f5234c3a2414378110e0a2899f128b821ad18",
  "latency_ms": 10005
 },
 "S82sh26::signage": {
  "fp": "a4c961976af9b64e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S82sh26": {
  "input_fingerprint": "142c475ed450a635",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 유리 벽 앞에 나란히 서서 찰리를 내려다보는 현우와 지소영의 굳은 뒷모습 풀샷.\n\nLOCATION (lock): In the observation area outside the laboratory's circular glass wall, bathed in the blue light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Bed visible between the observers beyond the glass in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Remains between the observers and 찰리) — The bed and 찰리 are seen through the section in front of the two observers; used as Sustains the separation across the widened composition; Stainless-steel bed (Occupied by the motionless 찰리) — Seen beyond the interval between 현우 and 지소영; used as Shared eyeline destination, kept small enough to preserve laboratory space; Scanning arms (Continue scanning 찰리) — Visible around the bed beyond the observers; used as Ongoing treatment contrasts with the observers' stillness.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The described blue laboratory illumination remains subdued, preserving separation between the two backs, the glass, and 찰리 beyond.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same glass enclosure, stainless-steel bed, surrounding research equipment, and cool lighting. Exclude shelter beds and domestic furnishings from the temporary accommodation.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains on the stainless-steel bed behind the circular glass wall, inactive and connected to multiple wires during treatment. Robotic scanning arms and the surrounding computer stations remain in place. 현우: He is wearing white recovery clothes and stands at the glass wall with a worried expression. 지소영: She stands at the glass wall wearing her neat research coat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 유리 벽 앞에 나란히 서서 찰리를 내려다보는 현우와 지소영의 굳은 뒷모습 풀샷.\n\nLOCATION (lock): In the observation area outside the laboratory's circular glass wall, bathed in the blue light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Bed visible between the observers beyond the glass in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Remains between the observers and 찰리) — The bed and 찰리 are seen through the section in front of the two observers; used as Sustains the separation across the widened composition; Stainless-steel bed (Occupied by the motionless 찰리) — Seen beyond the interval between 현우 and 지소영; used as Shared eyeline destination, kept small enough to preserve laboratory space; Scanning arms (Continue scanning 찰리) — Visible around the bed beyond the observers; used as Ongoing treatment contrasts with the observers' stillness.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The described blue laboratory illumination remains subdued, preserving separation between the two backs, the glass, and 찰리 beyond.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same glass enclosure, stainless-steel bed, surrounding research equipment, and cool lighting. Exclude shelter beds and domestic furnishings from the temporary accommodation.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains on the stainless-steel bed behind the circular glass wall, inactive and connected to multiple wires during treatment. Robotic scanning arms and the surrounding computer stations remain in place. 현우: He is wearing white recovery clothes and stands at the glass wall with a worried expression. 지소영: She stands at the glass wall wearing her neat research coat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 유리 벽 앞에 나란히 서서 찰리를 내려다보는 현우와 지소영의 굳은 뒷모습 풀샷.\n\nLOCATION (lock): In the observation area outside the laboratory's circular glass wall, bathed in the blue light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Bed visible between the observers beyond the glass in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Remains between the observers and 찰리) — The bed and 찰리 are seen through the section in front of the two observers; used as Sustains the separation across the widened composition; Stainless-steel bed (Occupied by the motionless 찰리) — Seen beyond the interval between 현우 and 지소영; used as Shared eyeline destination, kept small enough to preserve laboratory space; Scanning arms (Continue scanning 찰리) — Visible around the bed beyond the observers; used as Ongoing treatment contrasts with the observers' stillness.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The described blue laboratory illumination remains subdued, preserving separation between the two backs, the glass, and 찰리 beyond.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same glass enclosure, stainless-steel bed, surrounding research equipment, and cool lighting. Exclude shelter beds and domestic furnishings from the temporary accommodation.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains on the stainless-steel bed behind the circular glass wall, inactive and connected to multiple wires during treatment. Robotic scanning arms and the surrounding computer stations remain in place. 현우: He is wearing white recovery clothes and stands at the glass wall with a worried expression. 지소영: She stands at the glass wall wearing her neat research coat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 지소영은 카메라를 등지고 서서 유리 벽 너머 침대 위의 찰리를 향해 시선을 고정하고 있습니다. 스캐닝 로봇 팔들은 명확하게 찰리를 향해 뻗어 있습니다.",
    "built_space": "카메라가 원형 유리 벽 바깥에 위치하며, 지침대로 유리 벽을 사이에 두고 안쪽의 스테인리스 침대가 두 관찰자 사이 중앙 배경에 정확히 배치되었습니다.",
    "entities": "현우(헝클어진 검은 머리, 백색 옷), 지소영(정돈된 검은 머리, 연구원 코트), 찰리(샌드 베이지 장갑판, 눕혀진 상태) 모두 지정된 레퍼런스의 외형적 특징과 일치합니다.",
    "hard_violations": [],
    "physics": "현우와 지소영은 바닥에 두 발을 딛고 서 있으며, 찰리는 침대 표면에 완전히 지지되어 누워 있습니다. 공중에 떠 있거나 지지되지 않은 사물은 없습니다."
   },
   {
    "label": "B",
    "direction": "두 인물이 카메라를 등진 채 유리 안쪽의 찰리를 바라보고 있으며, 로봇 팔 역시 찰리를 겨냥하고 있습니다.",
    "built_space": "유리 벽을 사이에 둔 공간 분리와 두 인물 사이로 침대가 보이는 중앙 배치는 지침에 맞게 구현되었습니다.",
    "entities": "주요 세 인물의 외형은 레퍼런스를 잘 따랐으나, 우측 유리 너머 배경에 샷 텍스트에 없는 추가 인물이 존재합니다.",
    "hard_violations": [
     "[gemini-pro] invented people or objects (우측 배경 책상에 프롬프트가 명시적으로 금지한 이전 샷의 추가 인물이 포함됨)",
     "[gpt-high] 오른쪽 배경 컴퓨터 앞에 앉은 추가 연구원이 있어, 지정된 인물 외에는 등장시키지 말라는 조건을 위반한다."
    ],
    "physics": "모든 인물과 장비가 바닥이나 침대 등 적절한 표면에 위치하며, 물리적으로 불가능한 포즈나 떠 있는 사물은 발견되지 않습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 프레이밍과 앵글을 완벽하게 구현했으며, 이전 샷의 인물 배제 지침을 정확히 준수하여 핵심 피사체에 집중한 결과물입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "프레이밍과 전반적인 분위기는 양호하나, 우측 배경에 금지된 엑스트라 인물을 포함하여 치명적인 지침 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 지소영은 카메라를 등지고 서서 유리 벽 너머 침대 위의 찰리를 향해 시선을 고정하고 있습니다. 스캐닝 로봇 팔들은 명확하게 찰리를 향해 뻗어 있습니다.",
        "built_space": "카메라가 원형 유리 벽 바깥에 위치하며, 지침대로 유리 벽을 사이에 두고 안쪽의 스테인리스 침대가 두 관찰자 사이 중앙 배경에 정확히 배치되었습니다.",
        "entities": "현우(헝클어진 검은 머리, 백색 옷), 지소영(정돈된 검은 머리, 연구원 코트), 찰리(샌드 베이지 장갑판, 눕혀진 상태) 모두 지정된 레퍼런스의 외형적 특징과 일치합니다.",
        "hard_violations": [],
        "physics": "현우와 지소영은 바닥에 두 발을 딛고 서 있으며, 찰리는 침대 표면에 완전히 지지되어 누워 있습니다. 공중에 떠 있거나 지지되지 않은 사물은 없습니다."
       },
       {
        "label": "B",
        "direction": "두 인물이 카메라를 등진 채 유리 안쪽의 찰리를 바라보고 있으며, 로봇 팔 역시 찰리를 겨냥하고 있습니다.",
        "built_space": "유리 벽을 사이에 둔 공간 분리와 두 인물 사이로 침대가 보이는 중앙 배치는 지침에 맞게 구현되었습니다.",
        "entities": "주요 세 인물의 외형은 레퍼런스를 잘 따랐으나, 우측 유리 너머 배경에 샷 텍스트에 없는 추가 인물이 존재합니다.",
        "hard_violations": [
         "invented people or objects (우측 배경 책상에 프롬프트가 명시적으로 금지한 이전 샷의 추가 인물이 포함됨)"
        ],
        "physics": "모든 인물과 장비가 바닥이나 침대 등 적절한 표면에 위치하며, 물리적으로 불가능한 포즈나 떠 있는 사물은 발견되지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 프레이밍과 앵글을 완벽하게 구현했으며, 이전 샷의 인물 배제 지침을 정확히 준수하여 핵심 피사체에 집중한 결과물입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "프레이밍과 전반적인 분위기는 양호하나, 우측 배경에 금지된 엑스트라 인물을 포함하여 치명적인 지침 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 지소영은 카메라를 등지고 서서 유리 벽 너머 침대 위의 찰리를 향해 시선을 고정하고 있습니다. 스캐닝 로봇 팔들은 명확하게 찰리를 향해 뻗어 있습니다.",
        "built_space": "카메라가 원형 유리 벽 바깥에 위치하며, 지침대로 유리 벽을 사이에 두고 안쪽의 스테인리스 침대가 두 관찰자 사이 중앙 배경에 정확히 배치되었습니다.",
        "entities": "현우(헝클어진 검은 머리, 백색 옷), 지소영(정돈된 검은 머리, 연구원 코트), 찰리(샌드 베이지 장갑판, 눕혀진 상태) 모두 지정된 레퍼런스의 외형적 특징과 일치합니다.",
        "hard_violations": [],
        "physics": "현우와 지소영은 바닥에 두 발을 딛고 서 있으며, 찰리는 침대 표면에 완전히 지지되어 누워 있습니다. 공중에 떠 있거나 지지되지 않은 사물은 없습니다."
       },
       {
        "label": "B",
        "direction": "두 인물이 카메라를 등진 채 유리 안쪽의 찰리를 바라보고 있으며, 로봇 팔 역시 찰리를 겨냥하고 있습니다.",
        "built_space": "유리 벽을 사이에 둔 공간 분리와 두 인물 사이로 침대가 보이는 중앙 배치는 지침에 맞게 구현되었습니다.",
        "entities": "주요 세 인물의 외형은 레퍼런스를 잘 따랐으나, 우측 유리 너머 배경에 샷 텍스트에 없는 추가 인물이 존재합니다.",
        "hard_violations": [
         "invented people or objects (우측 배경 책상에 프롬프트가 명시적으로 금지한 이전 샷의 추가 인물이 포함됨)"
        ],
        "physics": "모든 인물과 장비가 바닥이나 침대 등 적절한 표면에 위치하며, 물리적으로 불가능한 포즈나 떠 있는 사물은 발견되지 않습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "오른쪽 배경에 금지된 추가 연구원이 보이며, 두 관찰자를 허벅지 부근에서 잘라 굳은 뒷모습 풀샷도 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "추가 인물 없이 두 사람 사이 중앙에 찰리와 검사대를 배치하고 유리벽의 분리를 살렸지만, 발이 잘려 두 후보 모두 엄밀한 풀샷에는 미달한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 지소영은 카메라에 등을 보이고 유리 안쪽의 찰리 쪽을 향한다. 눈은 보이지 않아 정확한 하향 시선은 확인할 수 없다. 왼쪽 관절형 검사 팔 끝은 찰리의 몸통과 다리 쪽으로 내려가며, 뒤쪽 원형 장치도 검사대를 향한다. 오른쪽 배경의 추가 인물은 컴퓨터 작업대 쪽을 보고 있다.",
        "built_space": "원형 유리 enclosure 한 개와 원형 하부 금속 테두리, 중앙의 금속 검사대 한 개가 보인다. 현우는 유리 밖 왼쪽, 지소영은 오른쪽에 서고, 둘 사이로 찰리가 가로로 누운 검사대가 보인다. 검사대 주변에는 왼쪽 관절형 팔, 뒤쪽 원형 검사 장치, 오른쪽에 일부 가려진 장치 끝이 보인다. 양쪽 배경에 컴퓨터 작업대가 있으며 오른쪽에는 앉은 연구원 한 명이 있다. 유리의 희미한 반사에서 명백한 광학적 모순은 확인되지 않는다.",
        "entities": "현우의 검은 머리와 흰 회복복은 대체로 맞지만 뒷모습이라 얼굴과 정확한 나이는 확인할 수 없다. 지소영은 흰 연구 가운을 입었으나 머리가 참고의 짧고 정돈된 머리보다 긴 단발이다. 찰리는 베이지색 장갑판, 흰 마스크형 머리, 육중한 팔을 가진 기계형 몸체로 보이며 검사대에 연결된 여러 전선도 있다. 오른쪽 작업대의 흰옷 인물은 허용된 세 인물 외의 추가 인물이다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "오른쪽 배경 컴퓨터 앞에 앉은 추가 연구원이 있어, 지정된 인물 외에는 등장시키지 말라는 조건을 위반한다."
        ],
        "physics": "찰리는 등과 몸통, 다리를 검사대에 얹고 누워 있으며 팔도 침상 가장자리에 놓여 있다. 전선은 몸체에서 침상 아래와 바닥으로 늘어진다. 검사 장치들은 관절과 지지대에 연결되어 있다. 현우의 뒤로 모은 손은 서로 맞닿아 있고 지소영의 팔은 아래로 내려져 있다. 두 관찰자의 발은 화면 밖이지만 몸이 공중에 떠 있다는 증거는 없다. 추가 연구원은 작업 의자에 앉아 있다."
       },
       {
        "label": "B",
        "direction": "두 관찰자는 등을 보인 채 중앙 검사대의 찰리를 향한다. 지소영의 머리는 약간 왼쪽 아래, 즉 찰리 쪽으로 돌아가 있고 현우도 검사대 방향을 향한다. 눈 자체는 보이지 않는다. 좌우 검사 팔의 끝은 찰리의 상체 위쪽을 향해 안쪽으로 기울어져 있어 검사 대상이 명확하다.",
        "built_space": "원형 유리벽 한 개와 원형 금속 바닥 테두리, 중앙의 스테인리스 검사대 한 개, 검사대 좌우의 로봇 검사 팔 두 개가 보인다. 현우는 유리 밖 왼쪽, 지소영은 오른쪽에 서 있으며 검사대는 둘 사이의 중경에서 발끝 쪽을 카메라로 향한다. 배경에는 컴퓨터 작업대와 장비장이 있고 추가 사람은 보이지 않는다. 유리가 관찰자와 찰리 사이에 유지되며, 참고의 금속·유리 연구실 재료와 푸른 조명을 대체로 따른다. 두 사람의 발은 하단에서 잘려 완전한 전신 구도는 아니다.",
        "entities": "현우는 검은 헝클어진 머리와 흰 회복복을 착용한 젊은 남성의 뒷모습으로 보인다. 지소영은 참고에 가까운 짧고 정돈된 검은 머리와 깔끔한 흰 연구 가운을 착용한다. 두 사람의 얼굴이 가려져 정확한 얼굴 일치와 나이·민족성은 판별할 수 없다. 찰리의 베이지 장갑판, 두꺼운 팔, 짧은 다리와 흰 머리 부분이 보이며, 얼굴 무늬는 거리와 각도상 확인하기 어렵다. 금속 검사대, 여러 연결 전선, 검사 팔과 컴퓨터 설비가 모두 존재하고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 몸통과 머리, 팔과 다리는 검사대 위에 놓여 있고 발바닥이 카메라 쪽으로 보인다. 의식 없는 몸이 능동적으로 사지를 들어 올린 모습은 아니다. 검사 팔은 관절식 구조와 받침대에 연결되고, 전선은 침상에서 바닥으로 내려와 바닥에 놓인다. 두 관찰자는 수직으로 서서 팔을 자연스럽게 내리고 있다. 발의 접지는 프레임 밖이지만 부유나 불가능한 자세의 증거는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "오른쪽 배경에 금지된 추가 연구원이 보이며, 두 관찰자를 허벅지 부근에서 잘라 굳은 뒷모습 풀샷도 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "추가 인물 없이 두 사람 사이 중앙에 찰리와 검사대를 배치하고 유리벽의 분리를 살렸지만, 발이 잘려 두 후보 모두 엄밀한 풀샷에는 미달한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 지소영은 카메라에 등을 보이고 유리 안쪽의 찰리 쪽을 향한다. 눈은 보이지 않아 정확한 하향 시선은 확인할 수 없다. 왼쪽 관절형 검사 팔 끝은 찰리의 몸통과 다리 쪽으로 내려가며, 뒤쪽 원형 장치도 검사대를 향한다. 오른쪽 배경의 추가 인물은 컴퓨터 작업대 쪽을 보고 있다.",
        "built_space": "원형 유리 enclosure 한 개와 원형 하부 금속 테두리, 중앙의 금속 검사대 한 개가 보인다. 현우는 유리 밖 왼쪽, 지소영은 오른쪽에 서고, 둘 사이로 찰리가 가로로 누운 검사대가 보인다. 검사대 주변에는 왼쪽 관절형 팔, 뒤쪽 원형 검사 장치, 오른쪽에 일부 가려진 장치 끝이 보인다. 양쪽 배경에 컴퓨터 작업대가 있으며 오른쪽에는 앉은 연구원 한 명이 있다. 유리의 희미한 반사에서 명백한 광학적 모순은 확인되지 않는다.",
        "entities": "현우의 검은 머리와 흰 회복복은 대체로 맞지만 뒷모습이라 얼굴과 정확한 나이는 확인할 수 없다. 지소영은 흰 연구 가운을 입었으나 머리가 참고의 짧고 정돈된 머리보다 긴 단발이다. 찰리는 베이지색 장갑판, 흰 마스크형 머리, 육중한 팔을 가진 기계형 몸체로 보이며 검사대에 연결된 여러 전선도 있다. 오른쪽 작업대의 흰옷 인물은 허용된 세 인물 외의 추가 인물이다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "오른쪽 배경 컴퓨터 앞에 앉은 추가 연구원이 있어, 지정된 인물 외에는 등장시키지 말라는 조건을 위반한다."
        ],
        "physics": "찰리는 등과 몸통, 다리를 검사대에 얹고 누워 있으며 팔도 침상 가장자리에 놓여 있다. 전선은 몸체에서 침상 아래와 바닥으로 늘어진다. 검사 장치들은 관절과 지지대에 연결되어 있다. 현우의 뒤로 모은 손은 서로 맞닿아 있고 지소영의 팔은 아래로 내려져 있다. 두 관찰자의 발은 화면 밖이지만 몸이 공중에 떠 있다는 증거는 없다. 추가 연구원은 작업 의자에 앉아 있다."
       },
       {
        "label": "A",
        "direction": "두 관찰자는 등을 보인 채 중앙 검사대의 찰리를 향한다. 지소영의 머리는 약간 왼쪽 아래, 즉 찰리 쪽으로 돌아가 있고 현우도 검사대 방향을 향한다. 눈 자체는 보이지 않는다. 좌우 검사 팔의 끝은 찰리의 상체 위쪽을 향해 안쪽으로 기울어져 있어 검사 대상이 명확하다.",
        "built_space": "원형 유리벽 한 개와 원형 금속 바닥 테두리, 중앙의 스테인리스 검사대 한 개, 검사대 좌우의 로봇 검사 팔 두 개가 보인다. 현우는 유리 밖 왼쪽, 지소영은 오른쪽에 서 있으며 검사대는 둘 사이의 중경에서 발끝 쪽을 카메라로 향한다. 배경에는 컴퓨터 작업대와 장비장이 있고 추가 사람은 보이지 않는다. 유리가 관찰자와 찰리 사이에 유지되며, 참고의 금속·유리 연구실 재료와 푸른 조명을 대체로 따른다. 두 사람의 발은 하단에서 잘려 완전한 전신 구도는 아니다.",
        "entities": "현우는 검은 헝클어진 머리와 흰 회복복을 착용한 젊은 남성의 뒷모습으로 보인다. 지소영은 참고에 가까운 짧고 정돈된 검은 머리와 깔끔한 흰 연구 가운을 착용한다. 두 사람의 얼굴이 가려져 정확한 얼굴 일치와 나이·민족성은 판별할 수 없다. 찰리의 베이지 장갑판, 두꺼운 팔, 짧은 다리와 흰 머리 부분이 보이며, 얼굴 무늬는 거리와 각도상 확인하기 어렵다. 금속 검사대, 여러 연결 전선, 검사 팔과 컴퓨터 설비가 모두 존재하고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 몸통과 머리, 팔과 다리는 검사대 위에 놓여 있고 발바닥이 카메라 쪽으로 보인다. 의식 없는 몸이 능동적으로 사지를 들어 올린 모습은 아니다. 검사 팔은 관절식 구조와 받침대에 연결되고, 전선은 침상에서 바닥으로 내려와 바닥에 놓인다. 두 관찰자는 수직으로 서서 팔을 자연스럽게 내리고 있다. 발의 접지는 프레임 밖이지만 부유나 불가능한 자세의 증거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] invented people or objects (우측 배경 책상에 프롬프트가 명시적으로 금지한 이전 샷의 추가 인물이 포함됨)",
     "[gpt-high] 오른쪽 배경 컴퓨터 앞에 앉은 추가 연구원이 있어, 지정된 인물 외에는 등장시키지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 프레이밍과 앵글을 완벽하게 구현했으며, 이전 샷의 인물 배제 지침을 정확히 준수하여 핵심 피사체에 집중한 결과물입니다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "프레이밍과 전반적인 분위기는 양호하나, 우측 배경에 금지된 엑스트라 인물을 포함하여 치명적인 지침 위반이 발생했습니다.  ★위반: [gemini-pro] invented people or objects (우측 배경 책상에 프롬프트가 명시적으로 금지한 이전 샷의 추가 인물이 포함됨) / [gpt-high] 오른쪽 배경 컴퓨터 앞에 앉은 추가 연구원이 있어, 지정된 인물 외에는 등장시키지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh17_sel.png",
    "asset_id": "5f6cedd5-5ed8-43a0-b3be-f3021b305f45",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:703308>",
    "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1461211>",
    "asset_id": "24d1c172-ad7f-4d9b-9990-330c659e9ca6",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-f998-717b-a287-8a6b1d19a134",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh17"
  },
  "locked_char_refs_kept_body_identity": [
   "찰리(C06)"
  ],
  "staged_characters_added": [
   "C06"
  ]
 },
 "S82sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:00:13.005182+00:00",
  "fingerprint": "bdeee6714e2c7069e0db75234864c8f5a6f310a8ddda6a6c30c123e0a3666aeb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S82sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S82sh26_sel.png",
  "source_sha256": "ee0aec161399f0c5b79971ca532b4f8b1d1cc2b562999de1dd6ca4ffd94f21d7",
  "file": "S82sh26_cine.png",
  "staged_sha256": "0fee7a2b16fc5f0e51569d33beb7946f4806de8b16bbc947ab47d0a5258c0bcd",
  "latency_ms": 12455
 },
 "S83sh5::signage": {
  "fp": "a888e4ab561146ab",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S83sh5": {
  "input_fingerprint": "ca0e75bb15e37fe2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call is now in video mode, showing Amber after her completed operation and Raul appearing in the live image.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call is now in video mode, showing Amber after her completed operation and Raul appearing in the live image.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call is now in video mode, showing Amber after her completed operation and Raul appearing in the live image.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh5__bgfirst_bg.png",
     "asset_id": "3bc12182-b8a5-4b6b-bbfb-1fc47da3fb74",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S83sh5.png",
     "asset_id": "481275cf-ea10-437e-8452-befc99e03a7f",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
     "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
    "built_space": "클로즈업 지시를 무시하고 병실 전체(침대, 창문 포함)를 보여주며, 벽에 걸린 모니터는 공간의 원근법을 무시하고 가로로 기형적으로 길게 늘어나 있음.",
    "entities": "화면 속에 앰버와 라울이 지정된 인상착의로 등장함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
    ],
    "physics": "모니터가 벽에 걸려 있으나 형태 왜곡으로 인해 지지 구조가 비현실적으로 보임."
   },
   {
    "label": "B",
    "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
    "built_space": "벽걸이 모니터와 장착 패널을 가깝게 잡은 클로즈업 구도이며, 카메라와 모니터의 배치가 자연스러움.",
    "entities": "화면 속에 앰버와 라울이 지정된 인상착의(금발 머리, 꽁지머리 등)로 정확히 등장함.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)",
     "[gpt-high] 오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "모니터가 벽면 패널에 안정적으로 부착되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 프레이밍과 화면 속 인물들의 표정을 정확히 구현했으나, 우측 벽면에 금지된 텍스트가 생성된 점이 감점 요인입니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시된 클로즈업을 무시하고 방 전체를 보여주는 와이드 샷으로 렌더링했으며, 모니터의 가로 비율과 원근이 비현실적으로 왜곡되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "벽걸이 모니터와 장착 패널을 가깝게 잡은 클로즈업 구도이며, 카메라와 모니터의 배치가 자연스러움.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의(금발 머리, 꽁지머리 등)로 정확히 등장함.",
        "hard_violations": [
         "읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)"
        ],
        "physics": "모니터가 벽면 패널에 안정적으로 부착되어 있음."
       },
       {
        "label": "A",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "클로즈업 지시를 무시하고 병실 전체(침대, 창문 포함)를 보여주며, 벽에 걸린 모니터는 공간의 원근법을 무시하고 가로로 기형적으로 길게 늘어나 있음.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의로 등장함.",
        "hard_violations": [
         "물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
        ],
        "physics": "모니터가 벽에 걸려 있으나 형태 왜곡으로 인해 지지 구조가 비현실적으로 보임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 프레이밍과 화면 속 인물들의 표정을 정확히 구현했으나, 우측 벽면에 금지된 텍스트가 생성된 점이 감점 요인입니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시된 클로즈업을 무시하고 방 전체를 보여주는 와이드 샷으로 렌더링했으며, 모니터의 가로 비율과 원근이 비현실적으로 왜곡되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "벽걸이 모니터와 장착 패널을 가깝게 잡은 클로즈업 구도이며, 카메라와 모니터의 배치가 자연스러움.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의(금발 머리, 꽁지머리 등)로 정확히 등장함.",
        "hard_violations": [
         "읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)"
        ],
        "physics": "모니터가 벽면 패널에 안정적으로 부착되어 있음."
       },
       {
        "label": "A",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "클로즈업 지시를 무시하고 병실 전체(침대, 창문 포함)를 보여주며, 벽에 걸린 모니터는 공간의 원근법을 무시하고 가로로 기형적으로 길게 늘어나 있음.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의로 등장함.",
        "hard_violations": [
         "물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
        ],
        "physics": "모니터가 벽에 걸려 있으나 형태 왜곡으로 인해 지지 구조가 비현실적으로 보임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "비스듬한 모니터의 전체 테두리와 웃는 두 얼굴은 잘 담았지만, 오른쪽 안내문과 단자의 읽을 수 있는 글자가 문자 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "주간 임시보호실의 모니터 안에서 두 아이가 활짝 웃는 순간을 충실히 구현했으나, 화면 상단 테두리가 일부 잘리고 방의 비중이 다소 크다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 라울 모두 원격 통화 카메라 쪽을 바라보며 이를 드러내고 웃는다. 모니터의 영상 표시 면은 촬영 카메라에 비스듬히 보이며, 위쪽 웹캠은 방 안쪽을 향한다.",
        "built_space": "흰 벽에 모니터 한 대, 상단 웹캠 한 개, 뒤쪽 금속 설치판 한 개가 보인다. 설치판 오른쪽에는 연결 단자부 두 구획이 있고, 더 오른쪽에는 안내문과 벽 제어판 일부가 보인다. 참고 장소의 벽걸이 장비 구성을 대체로 유지하며 모니터 테두리 전체가 프레임 안에 있다. 화면의 옅은 창빛 반사는 맞은편 쪽 창에서 생길 수 있는 배치다.",
        "entities": "화면 속 인물은 두 명뿐이다. 앰버는 약 10세의 금발 여자아이로 큰 눈, 둥근 얼굴과 남색 상의가 참고와 대체로 맞는다. 라울은 약 10세의 갈색 피부 남자아이로 뒤로 묶은 검은 머리와 남색 상의가 참고와 맞는다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물 모두 제시된 외형에 부합한다. 원격 배경은 장소를 특정하기 어려운 밝은 면이다. 오른쪽 안내문의 한글과 단자의 영문 표기가 노출되어 있다.",
        "hard_violations": [
         "오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "모니터는 벽의 설치판에 부착되어 있고 웹캠은 상단 받침에 놓여 있다. 두 아이의 머리는 목과 어깨에 자연스럽게 연결되어 있으며, 화면 밖 하체의 지지는 이 얼굴 클로즈업에서 판단할 필요가 없다. 공중에 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "두 아이는 원격 통화 카메라를 향해 시선을 두고 입을 벌려 활짝 웃는다. 모니터 앞면은 촬영 카메라 쪽에 비스듬히 노출되고 상단 웹캠은 임시보호실 안쪽을 향해 있어 통화 장비의 사용 방향이 자연스럽다.",
        "built_space": "오른쪽 벽의 모니터 한 대와 금속 설치판 한 개, 웹캠 한 개, 오른쪽 세로 제어부 한 개가 보인다. 왼쪽에는 바다가 보이는 창 한 구획과 병상 한 대의 발치 일부가 있으며, 천장 조명과 밝은 벽·회색 바닥도 참고 장소와 부합한다. 얼굴을 둘러싼 모니터의 좌우·아래 테두리는 보이지만 상단 테두리 일부는 프레임 밖으로 잘린다. 방이 차지하는 면적은 얼굴 클로즈업 요구에 비해 다소 크다.",
        "entities": "등장인물은 모니터 속 앰버와 라울 두 명뿐이다. 앰버의 어린 여자아이 외형, 금발, 큰 눈, 둥근 볼과 남색 상의가 참고에 부합한다. 라울의 어린 남자아이 외형, 갈색 피부, 뒤로 묶은 검은 머리와 남색 상의도 부합한다. 두 사람의 혼혈 배경은 영상만으로 확정할 수 없다. 통화 화면 안의 밝은 배경은 특정 장소를 드러내지 않으며, 읽을 수 있는 글자나 자막은 보이지 않는다.",
        "hard_violations": [],
        "physics": "모니터는 벽 부착 설치판에 지지되고 웹캠은 모니터 위 받침에 고정되어 있다. 병상은 보이는 바퀴로 바닥에 지지된다. 두 아이의 얼굴과 목·어깨 연결은 자연스럽고, 웃는 표정에도 해부학적 이상이 없다. 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "비스듬한 모니터의 전체 테두리와 웃는 두 얼굴은 잘 담았지만, 오른쪽 안내문과 단자의 읽을 수 있는 글자가 문자 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "주간 임시보호실의 모니터 안에서 두 아이가 활짝 웃는 순간을 충실히 구현했으나, 화면 상단 테두리가 일부 잘리고 방의 비중이 다소 크다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버와 라울 모두 원격 통화 카메라 쪽을 바라보며 이를 드러내고 웃는다. 모니터의 영상 표시 면은 촬영 카메라에 비스듬히 보이며, 위쪽 웹캠은 방 안쪽을 향한다.",
        "built_space": "흰 벽에 모니터 한 대, 상단 웹캠 한 개, 뒤쪽 금속 설치판 한 개가 보인다. 설치판 오른쪽에는 연결 단자부 두 구획이 있고, 더 오른쪽에는 안내문과 벽 제어판 일부가 보인다. 참고 장소의 벽걸이 장비 구성을 대체로 유지하며 모니터 테두리 전체가 프레임 안에 있다. 화면의 옅은 창빛 반사는 맞은편 쪽 창에서 생길 수 있는 배치다.",
        "entities": "화면 속 인물은 두 명뿐이다. 앰버는 약 10세의 금발 여자아이로 큰 눈, 둥근 얼굴과 남색 상의가 참고와 대체로 맞는다. 라울은 약 10세의 갈색 피부 남자아이로 뒤로 묶은 검은 머리와 남색 상의가 참고와 맞는다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물 모두 제시된 외형에 부합한다. 원격 배경은 장소를 특정하기 어려운 밝은 면이다. 오른쪽 안내문의 한글과 단자의 영문 표기가 노출되어 있다.",
        "hard_violations": [
         "오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "모니터는 벽의 설치판에 부착되어 있고 웹캠은 상단 받침에 놓여 있다. 두 아이의 머리는 목과 어깨에 자연스럽게 연결되어 있으며, 화면 밖 하체의 지지는 이 얼굴 클로즈업에서 판단할 필요가 없다. 공중에 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "두 아이는 원격 통화 카메라를 향해 시선을 두고 입을 벌려 활짝 웃는다. 모니터 앞면은 촬영 카메라 쪽에 비스듬히 노출되고 상단 웹캠은 임시보호실 안쪽을 향해 있어 통화 장비의 사용 방향이 자연스럽다.",
        "built_space": "오른쪽 벽의 모니터 한 대와 금속 설치판 한 개, 웹캠 한 개, 오른쪽 세로 제어부 한 개가 보인다. 왼쪽에는 바다가 보이는 창 한 구획과 병상 한 대의 발치 일부가 있으며, 천장 조명과 밝은 벽·회색 바닥도 참고 장소와 부합한다. 얼굴을 둘러싼 모니터의 좌우·아래 테두리는 보이지만 상단 테두리 일부는 프레임 밖으로 잘린다. 방이 차지하는 면적은 얼굴 클로즈업 요구에 비해 다소 크다.",
        "entities": "등장인물은 모니터 속 앰버와 라울 두 명뿐이다. 앰버의 어린 여자아이 외형, 금발, 큰 눈, 둥근 볼과 남색 상의가 참고에 부합한다. 라울의 어린 남자아이 외형, 갈색 피부, 뒤로 묶은 검은 머리와 남색 상의도 부합한다. 두 사람의 혼혈 배경은 영상만으로 확정할 수 없다. 통화 화면 안의 밝은 배경은 특정 장소를 드러내지 않으며, 읽을 수 있는 글자나 자막은 보이지 않는다.",
        "hard_violations": [],
        "physics": "모니터는 벽 부착 설치판에 지지되고 웹캠은 모니터 위 받침에 고정되어 있다. 병상은 보이는 바퀴로 바닥에 지지된다. 두 아이의 얼굴과 목·어깨 연결은 자연스럽고, 웃는 표정에도 해부학적 이상이 없다. 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)",
     "[gpt-high] 오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "A": [
     "[gemini-pro] 물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1417
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "지시된 클로즈업 프레이밍과 화면 속 인물들의 표정을 정확히 구현했으나, 우측 벽면에 금지된 텍스트가 생성된 점이 감점 요인입니다.  ★위반: [gemini-pro] 읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨) / [gpt-high] 오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "지시된 클로즈업을 무시하고 방 전체를 보여주는 와이드 샷으로 렌더링했으며, 모니터의 가로 비율과 원근이 비현실적으로 왜곡되었습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
    "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-fb52-746b-9e7e-a189f32e73c6",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh5__bgfirst_bg.png",
   "bg_asset_id": "3bc12182-b8a5-4b6b-bbfb-1fc47da3fb74",
   "bg_record_key": "S83sh5::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S83sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:11:24.432660+00:00",
  "fingerprint": "7f9b2e861309273c6311f848628f7db5bf0d6d3c5ae866e4844d216df6d0b3ab",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S83sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S83sh5_sel.png",
  "source_sha256": "0c2f8d4f9b0c979c52f20a162c8b24bfc4ac5243fba9ef328df643010c38d6cd",
  "file": "S83sh5_cine.png",
  "staged_sha256": "dbd165e65b1bdeba27be181923f93a2cc382dea81a367a8d288cf38c08093b90",
  "latency_ms": 10968
 },
 "S83sh6::signage": {
  "fp": "b21dae096eda3b23",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S83sh6": {
  "input_fingerprint": "2af467627d874651",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 화면을 뚫어져라 응시한 채 눈시울이 붉어진 현우의 얼굴.\n\nLOCATION (lock): At the video-call station inside the temporary-care bedroom, lit by the monitor and the room's established fluorescent lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet ambient illumination and gentle facial contrast keep the reddened eyes visible without imposing an unmotivated monitor glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same shelter-room surfaces, bed, and daytime lighting. Exclude the doorway figure from the reference and the research laboratory's glass enclosure.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The video call remains connected; Amber is back in control of the phone after Raul's appearance. 현우: He remains at the telephone in his shelter room, watching the video display and becoming choked with emotion.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 화면을 뚫어져라 응시한 채 눈시울이 붉어진 현우의 얼굴.\n\nLOCATION (lock): At the video-call station inside the temporary-care bedroom, lit by the monitor and the room's established fluorescent lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet ambient illumination and gentle facial contrast keep the reddened eyes visible without imposing an unmotivated monitor glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same shelter-room surfaces, bed, and daytime lighting. Exclude the doorway figure from the reference and the research laboratory's glass enclosure.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The video call remains connected; Amber is back in control of the phone after Raul's appearance. 현우: He remains at the telephone in his shelter room, watching the video display and becoming choked with emotion.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 화면을 뚫어져라 응시한 채 눈시울이 붉어진 현우의 얼굴.\n\nLOCATION (lock): At the video-call station inside the temporary-care bedroom, lit by the monitor and the room's established fluorescent lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet ambient illumination and gentle facial contrast keep the reddened eyes visible without imposing an unmotivated monitor glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same shelter-room surfaces, bed, and daytime lighting. Exclude the doorway figure from the reference and the research laboratory's glass enclosure.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The video call remains connected; Amber is back in control of the phone after Raul's appearance. 현우: He remains at the telephone in his shelter room, watching the video display and becoming choked with emotion.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "피사체는 카메라 렌즈(정면)를 응시하고 있으며, 배경 좌측에 피사체를 등진 형태의 또 다른 모니터가 놓여 있습니다.",
    "built_space": "배경에 침대와 창문이 보이나, 창문이 침대 뒤쪽 정면 벽에 위치하여 레퍼런스의 측면 창문 구조와 모순됩니다.",
    "entities": "현우의 인물 특징(머리 모양, 얼굴)은 일치하나, 눈의 흰자위 전체가 부자연스러운 붉은색으로 변형되었습니다.",
    "hard_violations": [
     "[gemini-pro] 해부학적으로 불가능하게 변형된 붉은 안구 (지침 위반)",
     "[gemini-pro] 레퍼런스의 공간 구조와 모순되는 창문 및 가구 배치",
     "[gpt-high] 왼쪽 모니터에 현우를 닮은 남성의 인물상을 추가로 노출해, 현우만 등장해야 한다는 인물 제한을 어겼다."
    ],
    "physics": "피사체의 상체가 화면에 보이며, 지면과의 접촉면은 보이지 않으나 의자에 앉아 있는 자세로 지지되고 있습니다."
   },
   {
    "label": "B",
    "direction": "피사체는 프레임 좌측에 위치한 영상 통화 모니터를 향해 시선을 두고 화면을 응시하고 있습니다.",
    "built_space": "화면 좌측에 창문이 있고 피사체 뒤쪽으로 침대의 머리맡 부분이 배치되어 레퍼런스의 공간 구조 및 방향과 정확히 일치합니다.",
    "entities": "현우의 외모와 헝클어진 머리가 잘 묘사되었으며, 눈시울이 자연스럽게 붉어진 상태가 감정적으로 잘 표현되었습니다.",
    "hard_violations": [
     "[gpt-high] 통화 화면에 이 숏의 가시 인물로 허용되지 않은 여성을 큰 영상과 작은 영상으로 노출했다."
    ],
    "physics": "피사체의 몸통은 프레임 하단 밖의 보이지 않는 의자에 안정적으로 지지되어 앉아 있는 자세를 유지하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 방 구조를 정확히 반영했으며, 자연스럽게 눈시울이 붉어진 표정과 모니터를 응시하는 시선 연출이 지침에 부합합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "방의 창문과 침대 배치가 레퍼런스와 어긋나며, 안구의 흰자위가 비정상적으로 붉게 변형되어 해부학적 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "피사체는 카메라 렌즈(정면)를 응시하고 있으며, 배경 좌측에 피사체를 등진 형태의 또 다른 모니터가 놓여 있습니다.",
        "built_space": "배경에 침대와 창문이 보이나, 창문이 침대 뒤쪽 정면 벽에 위치하여 레퍼런스의 측면 창문 구조와 모순됩니다.",
        "entities": "현우의 인물 특징(머리 모양, 얼굴)은 일치하나, 눈의 흰자위 전체가 부자연스러운 붉은색으로 변형되었습니다.",
        "hard_violations": [
         "해부학적으로 불가능하게 변형된 붉은 안구 (지침 위반)",
         "레퍼런스의 공간 구조와 모순되는 창문 및 가구 배치"
        ],
        "physics": "피사체의 상체가 화면에 보이며, 지면과의 접촉면은 보이지 않으나 의자에 앉아 있는 자세로 지지되고 있습니다."
       },
       {
        "label": "B",
        "direction": "피사체는 프레임 좌측에 위치한 영상 통화 모니터를 향해 시선을 두고 화면을 응시하고 있습니다.",
        "built_space": "화면 좌측에 창문이 있고 피사체 뒤쪽으로 침대의 머리맡 부분이 배치되어 레퍼런스의 공간 구조 및 방향과 정확히 일치합니다.",
        "entities": "현우의 외모와 헝클어진 머리가 잘 묘사되었으며, 눈시울이 자연스럽게 붉어진 상태가 감정적으로 잘 표현되었습니다.",
        "hard_violations": [],
        "physics": "피사체의 몸통은 프레임 하단 밖의 보이지 않는 의자에 안정적으로 지지되어 앉아 있는 자세를 유지하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 방 구조를 정확히 반영했으며, 자연스럽게 눈시울이 붉어진 표정과 모니터를 응시하는 시선 연출이 지침에 부합합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "방의 창문과 침대 배치가 레퍼런스와 어긋나며, 안구의 흰자위가 비정상적으로 붉게 변형되어 해부학적 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "피사체는 카메라 렌즈(정면)를 응시하고 있으며, 배경 좌측에 피사체를 등진 형태의 또 다른 모니터가 놓여 있습니다.",
        "built_space": "배경에 침대와 창문이 보이나, 창문이 침대 뒤쪽 정면 벽에 위치하여 레퍼런스의 측면 창문 구조와 모순됩니다.",
        "entities": "현우의 인물 특징(머리 모양, 얼굴)은 일치하나, 눈의 흰자위 전체가 부자연스러운 붉은색으로 변형되었습니다.",
        "hard_violations": [
         "해부학적으로 불가능하게 변형된 붉은 안구 (지침 위반)",
         "레퍼런스의 공간 구조와 모순되는 창문 및 가구 배치"
        ],
        "physics": "피사체의 상체가 화면에 보이며, 지면과의 접촉면은 보이지 않으나 의자에 앉아 있는 자세로 지지되고 있습니다."
       },
       {
        "label": "B",
        "direction": "피사체는 프레임 좌측에 위치한 영상 통화 모니터를 향해 시선을 두고 화면을 응시하고 있습니다.",
        "built_space": "화면 좌측에 창문이 있고 피사체 뒤쪽으로 침대의 머리맡 부분이 배치되어 레퍼런스의 공간 구조 및 방향과 정확히 일치합니다.",
        "entities": "현우의 외모와 헝클어진 머리가 잘 묘사되었으며, 눈시울이 자연스럽게 붉어진 상태가 감정적으로 잘 표현되었습니다.",
        "hard_violations": [],
        "physics": "피사체의 몸통은 프레임 하단 밖의 보이지 않는 의자에 안정적으로 지지되어 앉아 있는 자세를 유지하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 붉어진 눈시울은 충실하지만, 시선이 왼쪽 통화 화면에 정확히 닿지 않고 허용되지 않은 여성 영상도 노출된다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "전경 화면을 응시하는 방향과 침실의 공간 연속성이 더 명확하지만, 구도가 다소 넓고 옆 모니터에 현우를 닮은 추가 인물을 보여 완전한 성공은 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 동공은 왼쪽을 향하지만, 시선은 보이는 모니터의 상단 너머로 향하는 듯하다. 통화 상대 영상은 그보다 아래에 있어 화면을 뚫어져라 응시한다는 관계가 명확하지 않다. 모니터의 표시 면은 카메라에 크게 드러나며 현우 쪽으로 충분히 돌아가 있는지도 불명확하다.",
        "built_space": "왼쪽 창 일부, 아래쪽 침대 난간 하나, 왼쪽 모니터 하나, 뒤쪽 벽면 설비 띠와 오른쪽 위 조명 일부가 보인다. 흰 벽과 낮빛은 참고 장소에 부합하지만, 벽면 설비는 참고에서 확인되지 않는다. 얼굴 중심의 좁은 구도이므로 침대 전체나 출입문이 안 보이는 것은 문제가 아니다.",
        "entities": "헝클어진 검은 머리와 흰 상의를 가진 젊은 동아시아계 남성이 보이며 현우의 참고 얼굴과 대체로 일치한다. 국적은 외관으로 확인할 수 없다. 눈시울은 붉고 눈물기가 있으며 홍채와 동공은 정상적이다. 왼쪽 화면에는 긴 검은 머리의 여성과 작은 여성 영상이 보이는데, 이 숏에 허용된 현우 외의 인물이다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "통화 화면에 이 숏의 가시 인물로 허용되지 않은 여성을 큰 영상과 작은 영상으로 노출했다."
        ],
        "physics": "머리는 목과 어깨에 정상적으로 연결되고 상의는 몸에 자연스럽게 걸쳐 있다. 하체와 좌석은 구도 밖이므로 지지 상태를 판정할 수 없으며, 공중에 떠 있다고 볼 근거도 없다. 모니터 아래에는 지지대 일부가 보여 무지지 부유물로 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 정면에서 약간 아래에 있는 전경 모니터를 응시한다. 카메라는 그 모니터의 뒷면 또는 상단 외장을 보므로, 표시 면이 현우를 향하는 사용 방향이 성립한다. 왼쪽의 별도 모니터는 현우의 현재 응시 대상이 아니다.",
        "built_space": "뒤쪽에 가로로 놓인 침대 하나와 양 끝 프레임, 큰 창 하나의 여러 창짝, 천장 조명 하나, 왼쪽 수납 가구가 보인다. 흰 벽, 회색 바닥, 창 아래 침대의 관계는 참고 장소와 대체로 이어진다. 모니터는 전경과 왼쪽에 하나씩 총 두 개 보이며, 이 이중 화면 구성은 참고에서 확인되지 않는다. 얼굴뿐 아니라 어깨와 방을 상당히 포함해 요구된 얼굴 클로즈업보다 느슨하다.",
        "entities": "중앙 인물은 참고와 유사한 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 흰 상의를 갖췄다. 눈시울의 붉음과 정상적인 눈 구조가 보인다. 왼쪽 통화 화면에는 현우를 닮은 짧은 검은 머리의 남성이 다시 등장해, 현우만 보여야 하는 숏에 추가 인물 또는 중복 인물상을 만든다. 읽을 수 있는 문자는 확인되지 않는다.",
        "hard_violations": [
         "왼쪽 모니터에 현우를 닮은 남성의 인물상을 추가로 노출해, 현우만 등장해야 한다는 인물 제한을 어겼다."
        ],
        "physics": "현우의 머리와 상체 자세는 자연스럽고, 좌석과 하체는 전경 모니터 및 프레임 밖에 가려져 있다. 왼쪽 모니터는 받침대를 통해 가구 위에 놓여 있다. 전경 모니터의 받침은 구도 밖이므로 지지 여부를 직접 확인할 수 없지만 부유를 시사하는 모습은 없다. 침대와 침구도 정상적인 접촉 관계를 보인다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 붉어진 눈시울은 충실하지만, 시선이 왼쪽 통화 화면에 정확히 닿지 않고 허용되지 않은 여성 영상도 노출된다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "전경 화면을 응시하는 방향과 침실의 공간 연속성이 더 명확하지만, 구도가 다소 넓고 옆 모니터에 현우를 닮은 추가 인물을 보여 완전한 성공은 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 동공은 왼쪽을 향하지만, 시선은 보이는 모니터의 상단 너머로 향하는 듯하다. 통화 상대 영상은 그보다 아래에 있어 화면을 뚫어져라 응시한다는 관계가 명확하지 않다. 모니터의 표시 면은 카메라에 크게 드러나며 현우 쪽으로 충분히 돌아가 있는지도 불명확하다.",
        "built_space": "왼쪽 창 일부, 아래쪽 침대 난간 하나, 왼쪽 모니터 하나, 뒤쪽 벽면 설비 띠와 오른쪽 위 조명 일부가 보인다. 흰 벽과 낮빛은 참고 장소에 부합하지만, 벽면 설비는 참고에서 확인되지 않는다. 얼굴 중심의 좁은 구도이므로 침대 전체나 출입문이 안 보이는 것은 문제가 아니다.",
        "entities": "헝클어진 검은 머리와 흰 상의를 가진 젊은 동아시아계 남성이 보이며 현우의 참고 얼굴과 대체로 일치한다. 국적은 외관으로 확인할 수 없다. 눈시울은 붉고 눈물기가 있으며 홍채와 동공은 정상적이다. 왼쪽 화면에는 긴 검은 머리의 여성과 작은 여성 영상이 보이는데, 이 숏에 허용된 현우 외의 인물이다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "통화 화면에 이 숏의 가시 인물로 허용되지 않은 여성을 큰 영상과 작은 영상으로 노출했다."
        ],
        "physics": "머리는 목과 어깨에 정상적으로 연결되고 상의는 몸에 자연스럽게 걸쳐 있다. 하체와 좌석은 구도 밖이므로 지지 상태를 판정할 수 없으며, 공중에 떠 있다고 볼 근거도 없다. 모니터 아래에는 지지대 일부가 보여 무지지 부유물로 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 정면에서 약간 아래에 있는 전경 모니터를 응시한다. 카메라는 그 모니터의 뒷면 또는 상단 외장을 보므로, 표시 면이 현우를 향하는 사용 방향이 성립한다. 왼쪽의 별도 모니터는 현우의 현재 응시 대상이 아니다.",
        "built_space": "뒤쪽에 가로로 놓인 침대 하나와 양 끝 프레임, 큰 창 하나의 여러 창짝, 천장 조명 하나, 왼쪽 수납 가구가 보인다. 흰 벽, 회색 바닥, 창 아래 침대의 관계는 참고 장소와 대체로 이어진다. 모니터는 전경과 왼쪽에 하나씩 총 두 개 보이며, 이 이중 화면 구성은 참고에서 확인되지 않는다. 얼굴뿐 아니라 어깨와 방을 상당히 포함해 요구된 얼굴 클로즈업보다 느슨하다.",
        "entities": "중앙 인물은 참고와 유사한 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 흰 상의를 갖췄다. 눈시울의 붉음과 정상적인 눈 구조가 보인다. 왼쪽 통화 화면에는 현우를 닮은 짧은 검은 머리의 남성이 다시 등장해, 현우만 보여야 하는 숏에 추가 인물 또는 중복 인물상을 만든다. 읽을 수 있는 문자는 확인되지 않는다.",
        "hard_violations": [
         "왼쪽 모니터에 현우를 닮은 남성의 인물상을 추가로 노출해, 현우만 등장해야 한다는 인물 제한을 어겼다."
        ],
        "physics": "현우의 머리와 상체 자세는 자연스럽고, 좌석과 하체는 전경 모니터 및 프레임 밖에 가려져 있다. 왼쪽 모니터는 받침대를 통해 가구 위에 놓여 있다. 전경 모니터의 받침은 구도 밖이므로 지지 여부를 직접 확인할 수 없지만 부유를 시사하는 모습은 없다. 침대와 침구도 정상적인 접촉 관계를 보인다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 해부학적으로 불가능하게 변형된 붉은 안구 (지침 위반)",
     "[gemini-pro] 레퍼런스의 공간 구조와 모순되는 창문 및 가구 배치",
     "[gpt-high] 왼쪽 모니터에 현우를 닮은 남성의 인물상을 추가로 노출해, 현우만 등장해야 한다는 인물 제한을 어겼다."
    ],
    "B": [
     "[gpt-high] 통화 화면에 이 숏의 가시 인물로 허용되지 않은 여성을 큰 영상과 작은 영상으로 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1500,
   "A": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "레퍼런스의 방 구조를 정확히 반영했으며, 자연스럽게 눈시울이 붉어진 표정과 모니터를 응시하는 시선 연출이 지침에 부합합니다.  ★위반: [gpt-high] 통화 화면에 이 숏의 가시 인물로 허용되지 않은 여성을 큰 영상과 작은 영상으로 노출했다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "방의 창문과 침대 배치가 레퍼런스와 어긋나며, 안구의 흰자위가 비정상적으로 붉게 변형되어 해부학적 지침을 위반했습니다.  ★위반: [gemini-pro] 해부학적으로 불가능하게 변형된 붉은 안구 (지침 위반) / [gemini-pro] 레퍼런스의 공간 구조와 모순되는 창문 및 가구 배치 / [gpt-high] 왼쪽 모니터에 현우를 닮은 남성의 인물상을 추가로 노출해, 현우만 등장해야 한다는 인물 제한을 어겼다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5_sel.png",
    "asset_id": "90a3041b-898b-43a4-87d5-237a14a9a301",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf410-fe97-71b0-95b3-c133c06e6585",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S81sh5"
  }
 },
 "S83sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:01:16.904292+00:00",
  "fingerprint": "dc9f7f55a96c7a412efc4807e79c5b3d522ce3c14d5b1f862039ee44abcb34a9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S83sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S83sh6_sel.png",
  "source_sha256": "bdc7798a8d001f900199cbdca40894de2cd103e646495d36d4064ff3f829f3b5",
  "file": "S83sh6_cine.png",
  "staged_sha256": "2bea3c16d584bde5bb4da841b647daa24deaaff596a09194d5e99997a070b8e8",
  "latency_ms": 10335
 },
 "S83sh7::signage": {
  "fp": "6e2d91105960b25b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S83sh7": {
  "input_fingerprint": "19504fae88b48b19",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 침대 위에 홀로 누운 채 멍하니 천장을 올려다보는 현우의 전신.\n\nLOCATION (lock): On the bed inside the research facility's temporary-care bedroom, beneath fluorescent lighting with daylight outside the sea-facing window. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Bed (Occupied only by 현우 after the call) — The upper surface and footward side are visible from the elevated diagonal position; used as Supports the full-body composition without overwhelming the surrounding space; Room space around the bed (현우 is alone); used as Negative space that increases with the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained room ambience and soft tonal separation, allowing the empty space to feel quiet without inventing a change of time or light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the shelter room, bed, and daytime illumination. Exclude the researcher, who has left, and do not retain an active video-call image after the call ends.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call has ended and the video image has disappeared. 현우: He is alone in his shelter room, lying on the bed and remaining visibly preoccupied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 침대 위에 홀로 누운 채 멍하니 천장을 올려다보는 현우의 전신.\n\nLOCATION (lock): On the bed inside the research facility's temporary-care bedroom, beneath fluorescent lighting with daylight outside the sea-facing window. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Bed (Occupied only by 현우 after the call) — The upper surface and footward side are visible from the elevated diagonal position; used as Supports the full-body composition without overwhelming the surrounding space; Room space around the bed (현우 is alone); used as Negative space that increases with the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained room ambience and soft tonal separation, allowing the empty space to feel quiet without inventing a change of time or light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the shelter room, bed, and daytime illumination. Exclude the researcher, who has left, and do not retain an active video-call image after the call ends.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call has ended and the video image has disappeared. 현우: He is alone in his shelter room, lying on the bed and remaining visibly preoccupied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 침대 위에 홀로 누운 채 멍하니 천장을 올려다보는 현우의 전신.\n\nLOCATION (lock): On the bed inside the research facility's temporary-care bedroom, beneath fluorescent lighting with daylight outside the sea-facing window. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Bed (Occupied only by 현우 after the call) — The upper surface and footward side are visible from the elevated diagonal position; used as Supports the full-body composition without overwhelming the surrounding space; Room space around the bed (현우 is alone); used as Negative space that increases with the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained room ambience and soft tonal separation, allowing the empty space to feel quiet without inventing a change of time or light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the shelter room, bed, and daytime illumination. Exclude the researcher, who has left, and do not retain an active video-call image after the call ends.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call has ended and the video image has disappeared. 현우: He is alone in his shelter room, lying on the bed and remaining visibly preoccupied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 위를 향해 천장에 고정되어 있음.",
    "built_space": "왼쪽 창문, 꺼진 모니터, 중앙 침대 및 우측 협탁이 위치함. 요구된 대각선 하이앵글(elevated diagonal) 구도를 완벽히 따름.",
    "entities": "현우의 앳된 얼굴, 검은 머리 및 흰색 의상이 참조와 일치함. 통화가 끝난 후의 꺼진 모니터가 올바르게 묘사됨.",
    "hard_violations": [],
    "physics": "매트리스 위에 몸의 무게감이 자연스럽게 안착되어 누워 있음."
   },
   {
    "label": "B",
    "direction": "시선은 위를 향해 천장에 고정되어 있음.",
    "built_space": "왼쪽 창문과 중앙 침대가 있으나, 이전 샷의 필수 요소인 모니터가 사라짐. 대각선이 아닌 정면 대칭 구도로 렌더링됨.",
    "entities": "현우의 인물 묘사 및 의상은 일치하나, 모니터 소품이 누락됨.",
    "hard_violations": [],
    "physics": "침대 매트리스에 의해 온몸이 안정적으로 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 대각선 하이앵글 구도를 정확히 구현했으며, 이전 샷의 고정 요소인 모니터를 유지하되 화면을 꺼진 상태로 묘사하여 지시사항을 훌륭히 충족함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "명시된 대각선 구도를 무시하고 정면 앵글로 촬영했으며, 이전 샷에 존재했던 모니터를 완전히 누락시켜 공간의 연속성을 훼손함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 위를 향해 천장에 고정되어 있음.",
        "built_space": "왼쪽 창문, 꺼진 모니터, 중앙 침대 및 우측 협탁이 위치함. 요구된 대각선 하이앵글(elevated diagonal) 구도를 완벽히 따름.",
        "entities": "현우의 앳된 얼굴, 검은 머리 및 흰색 의상이 참조와 일치함. 통화가 끝난 후의 꺼진 모니터가 올바르게 묘사됨.",
        "hard_violations": [],
        "physics": "매트리스 위에 몸의 무게감이 자연스럽게 안착되어 누워 있음."
       },
       {
        "label": "B",
        "direction": "시선은 위를 향해 천장에 고정되어 있음.",
        "built_space": "왼쪽 창문과 중앙 침대가 있으나, 이전 샷의 필수 요소인 모니터가 사라짐. 대각선이 아닌 정면 대칭 구도로 렌더링됨.",
        "entities": "현우의 인물 묘사 및 의상은 일치하나, 모니터 소품이 누락됨.",
        "hard_violations": [],
        "physics": "침대 매트리스에 의해 온몸이 안정적으로 지탱되고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 대각선 하이앵글 구도를 정확히 구현했으며, 이전 샷의 고정 요소인 모니터를 유지하되 화면을 꺼진 상태로 묘사하여 지시사항을 훌륭히 충족함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "명시된 대각선 구도를 무시하고 정면 앵글로 촬영했으며, 이전 샷에 존재했던 모니터를 완전히 누락시켜 공간의 연속성을 훼손함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 위를 향해 천장에 고정되어 있음.",
        "built_space": "왼쪽 창문, 꺼진 모니터, 중앙 침대 및 우측 협탁이 위치함. 요구된 대각선 하이앵글(elevated diagonal) 구도를 완벽히 따름.",
        "entities": "현우의 앳된 얼굴, 검은 머리 및 흰색 의상이 참조와 일치함. 통화가 끝난 후의 꺼진 모니터가 올바르게 묘사됨.",
        "hard_violations": [],
        "physics": "매트리스 위에 몸의 무게감이 자연스럽게 안착되어 누워 있음."
       },
       {
        "label": "B",
        "direction": "시선은 위를 향해 천장에 고정되어 있음.",
        "built_space": "왼쪽 창문과 중앙 침대가 있으나, 이전 샷의 필수 요소인 모니터가 사라짐. 대각선이 아닌 정면 대칭 구도로 렌더링됨.",
        "entities": "현우의 인물 묘사 및 의상은 일치하나, 모니터 소품이 누락됨.",
        "hard_violations": [],
        "physics": "침대 매트리스에 의해 온몸이 안정적으로 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "홀로 누워 천장을 보는 전신과 낮의 병실은 맞지만, 발치 정면에 가까운 구도가 지정된 높은 대각선 시점보다 덜 충실하고 반소매는 인물 참조의 긴소매와 다릅니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "높은 대각선 와이드 구도에서 침대 윗면과 발치 측면, 천장을 응시하는 현우의 전신을 보여주며, 통화가 끝난 모니터와 기존 병실의 연속성도 잘 유지합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 등을 대고 누워 얼굴과 눈을 위로 향하며, 시선의 대상은 화면 밖 천장으로 읽힙니다. 카메라를 직접 응시하거나 다른 물건을 보는 모습은 아닙니다. 무기나 이동 중인 물체는 없습니다.",
        "built_space": "침대 한 개, 머리판과 발치판 각 한 개, 왼쪽의 올라간 난간과 오른쪽의 낮은 측면 레일, 왼쪽 창 하나, 머리 위 벽등 하나와 의료용 벽 패널 하나가 보입니다. 흰 벽과 밝은 바닥, 창밖 낮빛은 이전 장소와 부합합니다. 카메라는 높지만 침대 장축의 발치 정면에 가까워, 요구된 대각선 구도와 발치 측면의 입체적인 노출은 약합니다. 모니터는 보이지 않으며, 이 프레임만으로 제거 여부를 단정할 수 없습니다.",
        "entities": "인물은 현우로 읽히는 젊은 동아시아계 남성 한 명뿐이며, 헝클어진 검은 머리와 흰 환자복을 갖췄습니다. 한국계 미국인이라는 국적은 외관으로 확인할 수 없지만 외형상 충돌은 없습니다. 얼굴은 참조와 대체로 유사하나 인물 참조보다 소매가 짧고 맨발입니다. 연구자나 영상통화 속 인물, 읽을 수 있는 글자는 없습니다. 창밖에는 옅은 바다와 낮빛이 보입니다.",
        "hard_violations": [],
        "physics": "머리와 목은 베개에, 몸통과 골반 및 다리는 매트리스에 지지됩니다. 양손은 몸 옆 침구에 놓였고 발뒤꿈치도 침대에 닿아 있습니다. 침구의 눌림과 누운 자세가 물리적으로 가능하며, 지지 없이 떠 있는 신체나 물건은 없습니다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴은 위쪽을 향하고 눈은 카메라나 모니터가 아닌 천장 쪽에 머뭅니다. 멍하니 누워 생각에 잠긴 행동으로 읽힙니다. 왼쪽 모니터의 화면은 침대 쪽을 향해 있어 사용 방향도 자연스럽습니다.",
        "built_space": "침대 한 개에 머리판과 발치판 각 한 개, 양쪽 난간 두 개가 있고 현우는 그 사이 매트리스 위에 누워 있습니다. 창 하나, 벽등 하나, 긴 의료용 벽 패널 하나가 이전 병실의 배치와 재질을 이어갑니다. 왼쪽에는 모니터 한 대와 받침대, 금속 카트 하나가 있고 오른쪽에는 수납장 하나가 보입니다. 카트와 수납장은 이전의 좁은 화면에서 확인할 수 없었던 부분입니다. 높은 대각선 시점에서 침대 윗면과 발치의 바깥 면, 주변 바닥을 함께 보여줍니다. 꺼진 모니터의 희미한 침구·난간 반사는 화면이 침대를 향한 배치상 가능하며, 명백히 불가능한 반사는 보이지 않습니다.",
        "entities": "헝클어진 검은 머리와 앳된 얼굴의 동아시아계 남성 한 명만 있어 현우의 설정에 부합하고 얼굴도 참조와 가깝습니다. 흰 긴소매 상의와 흰 바지는 인물 참조에 더 충실합니다. 발끝 일부는 발치판에 가려져 신발의 세부는 확정하기 어렵습니다. 침대와 베개, 밝은 벽, 바다를 향한 창, 형광등형 벽등, 통화 영상이 사라진 모니터가 보입니다. 다른 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "머리는 베개에, 등과 골반 및 뻗은 다리는 매트리스에 지지됩니다. 손은 몸 옆 침구 위에 자연스럽게 놓여 있고 발목과 발은 침대 발치 쪽 침구에 받쳐져 있습니다. 모니터는 기둥형 받침대에, 침대와 카트는 바닥의 바퀴에 지지됩니다. 공중에 떠 있거나 불가능하게 결합된 신체·물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "홀로 누워 천장을 보는 전신과 낮의 병실은 맞지만, 발치 정면에 가까운 구도가 지정된 높은 대각선 시점보다 덜 충실하고 반소매는 인물 참조의 긴소매와 다릅니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "높은 대각선 와이드 구도에서 침대 윗면과 발치 측면, 천장을 응시하는 현우의 전신을 보여주며, 통화가 끝난 모니터와 기존 병실의 연속성도 잘 유지합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 등을 대고 누워 얼굴과 눈을 위로 향하며, 시선의 대상은 화면 밖 천장으로 읽힙니다. 카메라를 직접 응시하거나 다른 물건을 보는 모습은 아닙니다. 무기나 이동 중인 물체는 없습니다.",
        "built_space": "침대 한 개, 머리판과 발치판 각 한 개, 왼쪽의 올라간 난간과 오른쪽의 낮은 측면 레일, 왼쪽 창 하나, 머리 위 벽등 하나와 의료용 벽 패널 하나가 보입니다. 흰 벽과 밝은 바닥, 창밖 낮빛은 이전 장소와 부합합니다. 카메라는 높지만 침대 장축의 발치 정면에 가까워, 요구된 대각선 구도와 발치 측면의 입체적인 노출은 약합니다. 모니터는 보이지 않으며, 이 프레임만으로 제거 여부를 단정할 수 없습니다.",
        "entities": "인물은 현우로 읽히는 젊은 동아시아계 남성 한 명뿐이며, 헝클어진 검은 머리와 흰 환자복을 갖췄습니다. 한국계 미국인이라는 국적은 외관으로 확인할 수 없지만 외형상 충돌은 없습니다. 얼굴은 참조와 대체로 유사하나 인물 참조보다 소매가 짧고 맨발입니다. 연구자나 영상통화 속 인물, 읽을 수 있는 글자는 없습니다. 창밖에는 옅은 바다와 낮빛이 보입니다.",
        "hard_violations": [],
        "physics": "머리와 목은 베개에, 몸통과 골반 및 다리는 매트리스에 지지됩니다. 양손은 몸 옆 침구에 놓였고 발뒤꿈치도 침대에 닿아 있습니다. 침구의 눌림과 누운 자세가 물리적으로 가능하며, 지지 없이 떠 있는 신체나 물건은 없습니다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴은 위쪽을 향하고 눈은 카메라나 모니터가 아닌 천장 쪽에 머뭅니다. 멍하니 누워 생각에 잠긴 행동으로 읽힙니다. 왼쪽 모니터의 화면은 침대 쪽을 향해 있어 사용 방향도 자연스럽습니다.",
        "built_space": "침대 한 개에 머리판과 발치판 각 한 개, 양쪽 난간 두 개가 있고 현우는 그 사이 매트리스 위에 누워 있습니다. 창 하나, 벽등 하나, 긴 의료용 벽 패널 하나가 이전 병실의 배치와 재질을 이어갑니다. 왼쪽에는 모니터 한 대와 받침대, 금속 카트 하나가 있고 오른쪽에는 수납장 하나가 보입니다. 카트와 수납장은 이전의 좁은 화면에서 확인할 수 없었던 부분입니다. 높은 대각선 시점에서 침대 윗면과 발치의 바깥 면, 주변 바닥을 함께 보여줍니다. 꺼진 모니터의 희미한 침구·난간 반사는 화면이 침대를 향한 배치상 가능하며, 명백히 불가능한 반사는 보이지 않습니다.",
        "entities": "헝클어진 검은 머리와 앳된 얼굴의 동아시아계 남성 한 명만 있어 현우의 설정에 부합하고 얼굴도 참조와 가깝습니다. 흰 긴소매 상의와 흰 바지는 인물 참조에 더 충실합니다. 발끝 일부는 발치판에 가려져 신발의 세부는 확정하기 어렵습니다. 침대와 베개, 밝은 벽, 바다를 향한 창, 형광등형 벽등, 통화 영상이 사라진 모니터가 보입니다. 다른 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "머리는 베개에, 등과 골반 및 뻗은 다리는 매트리스에 지지됩니다. 손은 몸 옆 침구 위에 자연스럽게 놓여 있고 발목과 발은 침대 발치 쪽 침구에 받쳐져 있습니다. 모니터는 기둥형 받침대에, 침대와 카트는 바닥의 바퀴에 지지됩니다. 공중에 떠 있거나 불가능하게 결합된 신체·물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.349
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.349
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1349
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 대각선 하이앵글 구도를 정확히 구현했으며, 이전 샷의 고정 요소인 모니터를 유지하되 화면을 꺼진 상태로 묘사하여 지시사항을 훌륭히 충족함."
   },
   {
    "label": "B",
    "score": 1349,
    "verdict_ko": "명시된 대각선 구도를 무시하고 정면 앵글로 촬영했으며, 이전 샷에 존재했던 모니터를 완전히 누락시켜 공간의 연속성을 훼손함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh6_sel.png",
    "asset_id": "fa2e6da6-632d-4ffc-9d60-647fd45f5678",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-003a-7893-8559-fe9289ef5cce",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S83sh6"
  }
 },
 "S83sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:02:42.109606+00:00",
  "fingerprint": "cdd56d98bf3cd7ab996c0c86d404db4c6d4a3187b717688cb13418b987c5b9f6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S83sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S83sh7_sel.png",
  "source_sha256": "ba722f6e983d28b048d6733a17d10e839e574f113b9ba1168343a1b28d0efe3b",
  "file": "S83sh7_cine.png",
  "staged_sha256": "b9ce8479d8b67a7c2c8ce0f1ff86f928a65419639ad3033607fac9648571f183",
  "latency_ms": 9614
 },
 "S84sh1::signage": {
  "fp": "f4cb2978eafeee6a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S84sh1": {
  "input_fingerprint": "86a781fedc87ff16",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 실험실의 차가운 침대 위, 눈을 번쩍 뜬 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): On the examination bed inside the research laboratory's glass enclosure, under the blue lighting specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Still supporting 찰리 at the instant of awakening) — Only a narrow portion of its upper surface appears beside his head; used as Grounds the face in its reclining position; Attached wires (Still connected before 찰리 removes them); used as Small lower-edge details that establish his treatment context.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued blue laboratory illumination reveals the worn metal face and newly opened eyes without adding an unsupported eye-light color or flare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is still lying on the stainless-steel bed inside the laboratory's glass enclosure, with multiple wires attached across the body. The eyes have just powered on; the wires have not yet been removed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 실험실의 차가운 침대 위, 눈을 번쩍 뜬 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): On the examination bed inside the research laboratory's glass enclosure, under the blue lighting specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Still supporting 찰리 at the instant of awakening) — Only a narrow portion of its upper surface appears beside his head; used as Grounds the face in its reclining position; Attached wires (Still connected before 찰리 removes them); used as Small lower-edge details that establish his treatment context.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued blue laboratory illumination reveals the worn metal face and newly opened eyes without adding an unsupported eye-light color or flare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is still lying on the stainless-steel bed inside the laboratory's glass enclosure, with multiple wires attached across the body. The eyes have just powered on; the wires have not yet been removed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 실험실의 차가운 침대 위, 눈을 번쩍 뜬 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): On the examination bed inside the research laboratory's glass enclosure, under the blue lighting specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Still supporting 찰리 at the instant of awakening) — Only a narrow portion of its upper surface appears beside his head; used as Grounds the face in its reclining position; Attached wires (Still connected before 찰리 removes them); used as Small lower-edge details that establish his treatment context.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued blue laboratory illumination reveals the worn metal face and newly opened eyes without adding an unsupported eye-light color or flare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is still lying on the stainless-steel bed inside the laboratory's glass enclosure, with multiple wires attached across the body. The eyes have just powered on; the wires have not yet been removed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 고개가 옆으로 뉘어져 있으며 시선은 허공을 향함.",
    "built_space": "연구실 배경과 유리 벽면 앞의 스테인리스 침대 위에 머리가 놓여 있음.",
    "entities": "찰리의 마스크와 얼굴 형태가 레퍼런스와 완전히 다른 기계적인 형태로 발명됨.",
    "hard_violations": [
     "[gemini-pro] 이전 샷에서 고정된 자세(하늘을 향한 머리 방향)를 위반하고 고개를 옆으로 돌림",
     "[gemini-pro] 레퍼런스와 전혀 다른 임의의 얼굴(마스크) 디자인 생성"
    ],
    "physics": "머리와 몸체가 침대 위에 정상적으로 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 고개와 눈동자가 똑바로 위쪽을 향함.",
    "built_space": "스테인리스 침대, 유리 벽, 뒤쪽 연구실 배경과 조작 패널들이 보임.",
    "entities": "찰리의 흰색 마스크형 얼굴과 장갑판이 레퍼런스와 일치하며 인체의 눈이 묘사됨.",
    "hard_violations": [],
    "physics": "침대 위에 등을 대고 안정적으로 누워 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 클로즈업보다 프레임이 다소 넓게 잡혔으나, 캐릭터의 핵심 디자인과 이전 샷의 고정된 누운 자세를 정확히 유지했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레임 지시는 따랐으나, 이전 샷에서 고정된 머리 방향을 무시하고 레퍼런스와 완전히 다른 얼굴을 생성하여 치명적인 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 고개가 옆으로 뉘어져 있으며 시선은 허공을 향함.",
        "built_space": "연구실 배경과 유리 벽면 앞의 스테인리스 침대 위에 머리가 놓여 있음.",
        "entities": "찰리의 마스크와 얼굴 형태가 레퍼런스와 완전히 다른 기계적인 형태로 발명됨.",
        "hard_violations": [
         "이전 샷에서 고정된 자세(하늘을 향한 머리 방향)를 위반하고 고개를 옆으로 돌림",
         "레퍼런스와 전혀 다른 임의의 얼굴(마스크) 디자인 생성"
        ],
        "physics": "머리와 몸체가 침대 위에 정상적으로 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 고개와 눈동자가 똑바로 위쪽을 향함.",
        "built_space": "스테인리스 침대, 유리 벽, 뒤쪽 연구실 배경과 조작 패널들이 보임.",
        "entities": "찰리의 흰색 마스크형 얼굴과 장갑판이 레퍼런스와 일치하며 인체의 눈이 묘사됨.",
        "hard_violations": [],
        "physics": "침대 위에 등을 대고 안정적으로 누워 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 클로즈업보다 프레임이 다소 넓게 잡혔으나, 캐릭터의 핵심 디자인과 이전 샷의 고정된 누운 자세를 정확히 유지했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레임 지시는 따랐으나, 이전 샷에서 고정된 머리 방향을 무시하고 레퍼런스와 완전히 다른 얼굴을 생성하여 치명적인 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 고개가 옆으로 뉘어져 있으며 시선은 허공을 향함.",
        "built_space": "연구실 배경과 유리 벽면 앞의 스테인리스 침대 위에 머리가 놓여 있음.",
        "entities": "찰리의 마스크와 얼굴 형태가 레퍼런스와 완전히 다른 기계적인 형태로 발명됨.",
        "hard_violations": [
         "이전 샷에서 고정된 자세(하늘을 향한 머리 방향)를 위반하고 고개를 옆으로 돌림",
         "레퍼런스와 전혀 다른 임의의 얼굴(마스크) 디자인 생성"
        ],
        "physics": "머리와 몸체가 침대 위에 정상적으로 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 고개와 눈동자가 똑바로 위쪽을 향함.",
        "built_space": "스테인리스 침대, 유리 벽, 뒤쪽 연구실 배경과 조작 패널들이 보임.",
        "entities": "찰리의 흰색 마스크형 얼굴과 장갑판이 레퍼런스와 일치하며 인체의 눈이 묘사됨.",
        "hard_violations": [],
        "physics": "침대 위에 등을 대고 안정적으로 누워 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "푸른 조명과 누운 상태는 맞지만, 얼굴보다 흉부가 크게 잡힌 넓은 구도와 인간의 눈으로 바뀐 외형이 얼굴 클로즈업 및 찰리의 정체성 지시에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "낡은 흰 금속 얼굴과 기계식 눈을 크게 담아 각성 순간에 더 충실하지만, 침대와 실험실 배경의 노출은 요구한 최소 범위보다 넓습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 두 눈은 위쪽 천장 방향을 향하며 렌즈를 직접 응시하지 않습니다. 지정된 응시 대상은 없으므로 방향 자체는 맞지만, 크게 열린 눈은 기계식 광학 장치가 아니라 피부에 둘러싸인 인간의 눈으로 보입니다.",
        "built_space": "검사 침대 한 개의 상판과 테두리가 머리 왼쪽으로 넓게 보이고, 뒤에는 곡면 유리 칸막이와 바닥 배수 홈, 작업대 및 의자가 보입니다. 찰리는 침대에 누워 있으나 머리부터 흉부까지 포함되어 얼굴 클로즈업보다 넓습니다. 중복된 침대나 불가능한 반사는 보이지 않습니다.",
        "entities": "찰리 한 명이 보이며 흰 마스크형 금속 얼굴, 긁힌 샌드 베이지 장갑판, 원형 흉부 부품은 참조와 대체로 맞습니다. 다만 눈 주변의 사람 피부와 자연 안구는 참조의 로봇 얼굴과 다릅니다. 이전 장면처럼 모자와 코트는 없고, 왼쪽 하단에는 여러 전선과 커넥터가 보입니다. 푸른 저조도 조명이며 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "등과 어깨는 금속 침대에 놓여 있고 머리는 목 구조와 뒤쪽 침대 면에 의해 지지되는 자세입니다. 접촉부 일부는 장갑에 가려져 있지만 떠 있다고 볼 근거는 없습니다. 전선은 침대 위에 놓이거나 가장자리로 처지며, 몸을 일으키거나 물건을 들고 있는 동작은 없습니다."
       },
       {
        "label": "B",
        "direction": "비스듬히 누운 얼굴의 두 광학 눈은 천장 및 카메라가 있는 위쪽 공간을 향합니다. 특정 물체를 바라보라는 지시는 없으며, 열린 눈이 선명하게 드러납니다. 별도의 강한 눈빛 발광이나 광선은 보이지 않습니다.",
        "built_space": "검사 침대 한 개가 머리 아래에 있고, 뒤로 곡면 유리 벽과 바닥 배수 홈, 작업대, 모니터 열이 보입니다. 얼굴이 화면 중심을 크게 차지하고 몸통은 오른쪽에서 잘려 A보다 요구한 클로즈업에 가깝습니다. 다만 아래쪽 침대 상판과 뒤쪽 실험실은 '좁은 일부'라는 지시보다 많이 노출됩니다. 명백한 설비 중복이나 불가능한 반사는 없습니다.",
        "entities": "찰리의 마모된 흰 마스크형 얼굴, 선으로 나뉜 입 부위, 기계식 광학 눈, 베이지 장갑과 노출된 목 기구가 보입니다. 인간 피부를 추가하지 않아 참조의 로봇 정체성을 더 잘 유지합니다. 모자와 코트가 없는 상태도 이전 장면과 맞습니다. 여러 전선이 목과 몸통 쪽으로 이어지며 하단에 남아 있습니다. 다른 인물은 명확하게 식별되지 않고, 배경 화면의 글자는 읽을 수 없습니다.",
        "hard_violations": [],
        "physics": "머리 뒤쪽과 측면 하우징이 침대 상판에 닿아 있고, 목 기구는 침대에 누운 몸통으로 이어집니다. 머리를 공중에 들고 있는 자세가 아니며, 전선도 상판에 얹혀 있거나 가장자리 아래로 자연스럽게 늘어집니다. 눈만 열린 채 몸은 계속 지지받는 상태로 읽힙니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "푸른 조명과 누운 상태는 맞지만, 얼굴보다 흉부가 크게 잡힌 넓은 구도와 인간의 눈으로 바뀐 외형이 얼굴 클로즈업 및 찰리의 정체성 지시에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "낡은 흰 금속 얼굴과 기계식 눈을 크게 담아 각성 순간에 더 충실하지만, 침대와 실험실 배경의 노출은 요구한 최소 범위보다 넓습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 두 눈은 위쪽 천장 방향을 향하며 렌즈를 직접 응시하지 않습니다. 지정된 응시 대상은 없으므로 방향 자체는 맞지만, 크게 열린 눈은 기계식 광학 장치가 아니라 피부에 둘러싸인 인간의 눈으로 보입니다.",
        "built_space": "검사 침대 한 개의 상판과 테두리가 머리 왼쪽으로 넓게 보이고, 뒤에는 곡면 유리 칸막이와 바닥 배수 홈, 작업대 및 의자가 보입니다. 찰리는 침대에 누워 있으나 머리부터 흉부까지 포함되어 얼굴 클로즈업보다 넓습니다. 중복된 침대나 불가능한 반사는 보이지 않습니다.",
        "entities": "찰리 한 명이 보이며 흰 마스크형 금속 얼굴, 긁힌 샌드 베이지 장갑판, 원형 흉부 부품은 참조와 대체로 맞습니다. 다만 눈 주변의 사람 피부와 자연 안구는 참조의 로봇 얼굴과 다릅니다. 이전 장면처럼 모자와 코트는 없고, 왼쪽 하단에는 여러 전선과 커넥터가 보입니다. 푸른 저조도 조명이며 읽을 수 있는 문자는 없습니다.",
        "hard_violations": [],
        "physics": "등과 어깨는 금속 침대에 놓여 있고 머리는 목 구조와 뒤쪽 침대 면에 의해 지지되는 자세입니다. 접촉부 일부는 장갑에 가려져 있지만 떠 있다고 볼 근거는 없습니다. 전선은 침대 위에 놓이거나 가장자리로 처지며, 몸을 일으키거나 물건을 들고 있는 동작은 없습니다."
       },
       {
        "label": "A",
        "direction": "비스듬히 누운 얼굴의 두 광학 눈은 천장 및 카메라가 있는 위쪽 공간을 향합니다. 특정 물체를 바라보라는 지시는 없으며, 열린 눈이 선명하게 드러납니다. 별도의 강한 눈빛 발광이나 광선은 보이지 않습니다.",
        "built_space": "검사 침대 한 개가 머리 아래에 있고, 뒤로 곡면 유리 벽과 바닥 배수 홈, 작업대, 모니터 열이 보입니다. 얼굴이 화면 중심을 크게 차지하고 몸통은 오른쪽에서 잘려 A보다 요구한 클로즈업에 가깝습니다. 다만 아래쪽 침대 상판과 뒤쪽 실험실은 '좁은 일부'라는 지시보다 많이 노출됩니다. 명백한 설비 중복이나 불가능한 반사는 없습니다.",
        "entities": "찰리의 마모된 흰 마스크형 얼굴, 선으로 나뉜 입 부위, 기계식 광학 눈, 베이지 장갑과 노출된 목 기구가 보입니다. 인간 피부를 추가하지 않아 참조의 로봇 정체성을 더 잘 유지합니다. 모자와 코트가 없는 상태도 이전 장면과 맞습니다. 여러 전선이 목과 몸통 쪽으로 이어지며 하단에 남아 있습니다. 다른 인물은 명확하게 식별되지 않고, 배경 화면의 글자는 읽을 수 없습니다.",
        "hard_violations": [],
        "physics": "머리 뒤쪽과 측면 하우징이 침대 상판에 닿아 있고, 목 기구는 침대에 누운 몸통으로 이어집니다. 머리를 공중에 들고 있는 자세가 아니며, 전선도 상판에 얹혀 있거나 가장자리 아래로 자연스럽게 늘어집니다. 눈만 열린 채 몸은 계속 지지받는 상태로 읽힙니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.625
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.625
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷에서 고정된 자세(하늘을 향한 머리 방향)를 위반하고 고개를 옆으로 돌림",
     "[gemini-pro] 레퍼런스와 전혀 다른 임의의 얼굴(마스크) 디자인 생성"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1625,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1625,
    "verdict_ko": "지정된 클로즈업보다 프레임이 다소 넓게 잡혔으나, 캐릭터의 핵심 디자인과 이전 샷의 고정된 누운 자세를 정확히 유지했습니다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "클로즈업 프레임 지시는 따랐으나, 이전 샷에서 고정된 머리 방향을 무시하고 레퍼런스와 완전히 다른 얼굴을 생성하여 치명적인 오류를 범했습니다.  ★위반: [gemini-pro] 이전 샷에서 고정된 자세(하늘을 향한 머리 방향)를 위반하고 고개를 옆으로 돌림 / [gemini-pro] 레퍼런스와 전혀 다른 임의의 얼굴(마스크) 디자인 생성"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14_sel.png",
    "asset_id": "c78997e0-2898-4ac1-b041-54b98246752d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-01e5-7c91-b030-c59f852c49b4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh14"
  },
  "locked_char_refs_kept_body_identity": [
   "찰리(C06)"
  ]
 },
 "S84sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:03:39.636046+00:00",
  "fingerprint": "47361c5c11dbc3dfc8d923d70e5fb43a15eeecfcb9aa0422f9d691abf33a4056",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S84sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S84sh1_sel.png",
  "source_sha256": "7232978e05d1930e5124b8055f9c35dcfee1757352f27418b15d33d4f04e2bc4",
  "file": "S84sh1_cine.png",
  "staged_sha256": "c6b767d7d905a7c98f2cd1d07ba01b073bf71a53c93d2f7b3cdcf52a92a24260",
  "latency_ms": 9221
 },
 "S84sh11::signage": {
  "fp": "928caa1df27ed571",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::e4b209e943d61d96": {
  "subjects": [],
  "subject_text": "제주도 연구소 실험실과 유리벽 관찰 구역\n중앙 원형 유리벽 안에 스테인리스 침대와 기계 팔이 설치된 실험실. 바깥에는 컴퓨터 작업대와 관찰 공간이 둘러져 있다.",
  "identity": "canonical",
  "scope_id": "L269",
  "scope_role": "location_interior",
  "scope_sha": "17b47caf6574618e"
 },
 "S84sh11::bgfirst_bg": {
  "input_fingerprint": "f4e1e745bb3245d1",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11__bgfirst_bg.png",
  "asset_id": "d36617d7-e206-4f23-93f4-d6161a0fbc32",
  "input_asset_ids": [
   "b7d06adc-c73d-44f7-8b3e-12527f45e14c",
   "2d7bcd3a-3137-473f-b939-ac3ee04b6230"
  ]
 },
 "S84sh11": {
  "input_fingerprint": "360aeb2daa3ca47c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An old piano stands in the central hall of the modern director's office. Charlie is now mobile beside the piano, with the laboratory wires detached and the glass-enclosure exit left open. 지소영: She remains at the piano in her neat research coat, having paused her playing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An old piano stands in the central hall of the modern director's office. Charlie is now mobile beside the piano, with the laboratory wires detached and the glass-enclosure exit left open. 지소영: She remains at the piano in her neat research coat, having paused her playing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An old piano stands in the central hall of the modern director's office. Charlie is now mobile beside the piano, with the laboratory wires detached and the glass-enclosure exit left open. 지소영: She remains at the piano in her neat research coat, having paused her playing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11__bgfirst_bg.png",
     "asset_id": "d36617d7-e206-4f23-93f4-d6161a0fbc32",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S84sh11.png",
     "asset_id": "b7d06adc-c73d-44f7-8b3e-12527f45e14c",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:703308>",
     "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L269B03.png",
     "asset_id": "2d7bcd3a-3137-473f-b939-ac3ee04b6230",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:703308>",
     "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163969>",
     "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "지소영은 왼쪽 위를 향해 시선을 두고 있으며, 이는 화면 왼쪽 전경에 위치한 찰리의 얼굴(화면 밖)을 향함.",
    "built_space": "원형 유리벽 연구실 내부. 배경에 로봇 팔이 있고, 지소영은 피아노 의자에 앉아 있으며 우측 하단에 건반 일부가 보임.",
    "entities": "지소영은 50대 여성의 외모를 갖췄으나 지정된 연구복 대신 회색 재킷을 입음. 찰리는 모래색 장갑판과 코트 등 특징이 일치함.",
    "hard_violations": [],
    "physics": "지소영은 의자에 안정적으로 앉아 있고, 찰리의 신체 역시 물리적으로 올바르게 지지됨."
   },
   {
    "label": "B",
    "direction": "지소영은 화면 우측 상단의 화면 밖 대상을 바라보고 있음.",
    "built_space": "연구실 내부로 배경에 로봇 팔과 유리벽이 보이며, 전경 전체를 피아노 내부와 건반이 차지함.",
    "entities": "지소영의 얼굴 특징은 일치하나 흰색 재킷을 착용함. 찰리는 프레임에 나타나지 않음.",
    "hard_violations": [
     "[gemini-pro] 지소영이 연주를 멈추고 앉아있다는 설정과 달리, 피아노 건반이 카메라를 향하고 있어 인물이 피아노 반대편에 위치하는 물리적 모순 발생."
    ],
    "physics": "인물이 서 있는 위치가 연주용 좌석이 아니며, 피아노의 기능적 방향(건반)이 사용자(지소영)를 향하지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 어깨를 건 앵글로 지소영의 미소 짓는 얼굴을 포착했으며, 피아노와 배경 묘사가 지시사항에 부합하나 의상이 기준과 다소 다름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "피아노 건반이 카메라를 향해 있어 연주자인 지소영의 위치와 물리적으로 모순되며, 건반이 프레임을 과도하게 차지함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 왼쪽 위를 향해 시선을 두고 있으며, 이는 화면 왼쪽 전경에 위치한 찰리의 얼굴(화면 밖)을 향함.",
        "built_space": "원형 유리벽 연구실 내부. 배경에 로봇 팔이 있고, 지소영은 피아노 의자에 앉아 있으며 우측 하단에 건반 일부가 보임.",
        "entities": "지소영은 50대 여성의 외모를 갖췄으나 지정된 연구복 대신 회색 재킷을 입음. 찰리는 모래색 장갑판과 코트 등 특징이 일치함.",
        "hard_violations": [],
        "physics": "지소영은 의자에 안정적으로 앉아 있고, 찰리의 신체 역시 물리적으로 올바르게 지지됨."
       },
       {
        "label": "B",
        "direction": "지소영은 화면 우측 상단의 화면 밖 대상을 바라보고 있음.",
        "built_space": "연구실 내부로 배경에 로봇 팔과 유리벽이 보이며, 전경 전체를 피아노 내부와 건반이 차지함.",
        "entities": "지소영의 얼굴 특징은 일치하나 흰색 재킷을 착용함. 찰리는 프레임에 나타나지 않음.",
        "hard_violations": [
         "지소영이 연주를 멈추고 앉아있다는 설정과 달리, 피아노 건반이 카메라를 향하고 있어 인물이 피아노 반대편에 위치하는 물리적 모순 발생."
        ],
        "physics": "인물이 서 있는 위치가 연주용 좌석이 아니며, 피아노의 기능적 방향(건반)이 사용자(지소영)를 향하지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 어깨를 건 앵글로 지소영의 미소 짓는 얼굴을 포착했으며, 피아노와 배경 묘사가 지시사항에 부합하나 의상이 기준과 다소 다름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "피아노 건반이 카메라를 향해 있어 연주자인 지소영의 위치와 물리적으로 모순되며, 건반이 프레임을 과도하게 차지함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 왼쪽 위를 향해 시선을 두고 있으며, 이는 화면 왼쪽 전경에 위치한 찰리의 얼굴(화면 밖)을 향함.",
        "built_space": "원형 유리벽 연구실 내부. 배경에 로봇 팔이 있고, 지소영은 피아노 의자에 앉아 있으며 우측 하단에 건반 일부가 보임.",
        "entities": "지소영은 50대 여성의 외모를 갖췄으나 지정된 연구복 대신 회색 재킷을 입음. 찰리는 모래색 장갑판과 코트 등 특징이 일치함.",
        "hard_violations": [],
        "physics": "지소영은 의자에 안정적으로 앉아 있고, 찰리의 신체 역시 물리적으로 올바르게 지지됨."
       },
       {
        "label": "B",
        "direction": "지소영은 화면 우측 상단의 화면 밖 대상을 바라보고 있음.",
        "built_space": "연구실 내부로 배경에 로봇 팔과 유리벽이 보이며, 전경 전체를 피아노 내부와 건반이 차지함.",
        "entities": "지소영의 얼굴 특징은 일치하나 흰색 재킷을 착용함. 찰리는 프레임에 나타나지 않음.",
        "hard_violations": [
         "지소영이 연주를 멈추고 앉아있다는 설정과 달리, 피아노 건반이 카메라를 향하고 있어 인물이 피아노 반대편에 위치하는 물리적 모순 발생."
        ],
        "physics": "인물이 서 있는 위치가 연주용 좌석이 아니며, 피아노의 기능적 방향(건반)이 사용자(지소영)를 향하지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "온화한 미소와 올려다보는 정지 연기, 상대적으로 밀착된 구도가 우세하지만, 얼굴 클로즈업보다 넓고 하단에 건반 대신 피아노 내부가 크게 보인다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "찰리를 올려다보는 관계와 피아노 착석은 명확하지만, 찰리의 몸통과 지소영의 상체를 넓게 담은 구도가 최우선인 얼굴 클로즈업 지시에서 더 멀다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 턱과 눈을 화면 오른쪽 위로 향하며 잔잔하게 웃는다. 시선은 렌즈가 아니라 화면 밖의 높은 상대 위치로 향한다. 찰리가 보이지 않아 실제 눈맞춤 지점은 확인할 수 없지만, 화면 밖 찰리를 올려다보는 배치와 양립한다.",
        "built_space": "전경에 그랜드피아노 한 대의 열린 뚜껑과 현·금속 프레임이 보이고, 지소영은 그 너머에 있다. 뒤에는 곡면 유리 구획 하나, 좌우 로봇 팔 두 개, 중앙 장비 패널과 작업대가 보여 장소 참조의 주요 구조를 따른다. 다만 요청한 하단의 작은 사선 건반 조각은 없고 피아노 내부가 넓게 차지한다. 유리의 흐린 반사에서 명백한 광학적 모순은 보이지 않는다.",
        "entities": "보이는 인물은 지소영 한 명이다. 중년 한국인 여성으로 읽히는 얼굴, 단정한 짧은 검은 머리와 온화한 미소는 참조에 대체로 부합한다. 흰 연구복은 있으나 참조의 높은 목깃과 잠금장치 대신 일반적인 라펠 형태다. 피아노는 실물 그랜드피아노로 읽히며 찰리는 화면 밖이다. 읽을 수 있는 글자나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 상체는 자연스럽게 연결되고 정지 자세도 가능하다. 좌석과 골반은 피아노에 가려져 지지 접점을 확인할 수 없으며, 이 가림만으로 공중 부양이라고 판단할 근거는 없다. 열린 피아노 뚜껑에는 왼쪽의 대각선 지지 구조가 보이고, 현과 금속 프레임은 피아노 몸체 안에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "지소영은 화면 왼쪽 위, 바로 앞에 선 찰리의 화면 밖 머리 위치를 향해 눈과 얼굴을 올린다. 찰리의 장갑 몸통과 팔이 왼쪽 전경에 있어 시선의 상대가 명확하다. 찰리의 얼굴은 잘려 있어 그의 시선 방향은 확인할 수 없다.",
        "built_space": "오른쪽에 오래된 피아노 한 대와 건반이 있고, 지소영 뒤와 아래에 검은 피아노 벤치 하나가 보인다. 그녀는 건반 옆에서 찰리 쪽으로 상체를 돌린 배치다. 배경에는 곡면 유리 구획 하나, 오른쪽 로봇 팔 하나, 뒤쪽 작업대와 회전의자 하나가 보이며 참조의 재료와 공간 구성을 유지한다. 바닥 반사는 조명과 유리벽 위치에 비추어 가능하다. 건반과 공간을 보여주는 범위가 요청한 작은 하단 조각보다 훨씬 넓다.",
        "entities": "지소영의 중년 얼굴과 단정한 검은 머리는 참조에 대체로 부합하며 흰 라펠 연구복은 참조 의상의 목깃 구조와 다르다. 왼쪽 찰리의 육중한 팔, 샌드 베이지 장갑판, 어두운 올리브색 외투는 참조와 맞는다. 잘린 얼굴과 다리는 평가할 수 없다. 피아노와 벤치는 실물로 읽히고, 추가 인물이나 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "지소영의 골반 부근은 하단에 가리지만 바로 뒤의 벤치가 착석을 지지하는 위치에 있으며, 찰리를 향해 상체와 목을 돌리는 자세는 가능하다. 찰리의 팔은 어깨와 몸통에 연결되어 아래로 내려오고 발은 프레임 밖이다. 피아노와 벤치는 바닥에 놓인 가구로 읽히며, 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "온화한 미소와 올려다보는 정지 연기, 상대적으로 밀착된 구도가 우세하지만, 얼굴 클로즈업보다 넓고 하단에 건반 대신 피아노 내부가 크게 보인다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "찰리를 올려다보는 관계와 피아노 착석은 명확하지만, 찰리의 몸통과 지소영의 상체를 넓게 담은 구도가 최우선인 얼굴 클로즈업 지시에서 더 멀다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "지소영은 턱과 눈을 화면 오른쪽 위로 향하며 잔잔하게 웃는다. 시선은 렌즈가 아니라 화면 밖의 높은 상대 위치로 향한다. 찰리가 보이지 않아 실제 눈맞춤 지점은 확인할 수 없지만, 화면 밖 찰리를 올려다보는 배치와 양립한다.",
        "built_space": "전경에 그랜드피아노 한 대의 열린 뚜껑과 현·금속 프레임이 보이고, 지소영은 그 너머에 있다. 뒤에는 곡면 유리 구획 하나, 좌우 로봇 팔 두 개, 중앙 장비 패널과 작업대가 보여 장소 참조의 주요 구조를 따른다. 다만 요청한 하단의 작은 사선 건반 조각은 없고 피아노 내부가 넓게 차지한다. 유리의 흐린 반사에서 명백한 광학적 모순은 보이지 않는다.",
        "entities": "보이는 인물은 지소영 한 명이다. 중년 한국인 여성으로 읽히는 얼굴, 단정한 짧은 검은 머리와 온화한 미소는 참조에 대체로 부합한다. 흰 연구복은 있으나 참조의 높은 목깃과 잠금장치 대신 일반적인 라펠 형태다. 피아노는 실물 그랜드피아노로 읽히며 찰리는 화면 밖이다. 읽을 수 있는 글자나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 상체는 자연스럽게 연결되고 정지 자세도 가능하다. 좌석과 골반은 피아노에 가려져 지지 접점을 확인할 수 없으며, 이 가림만으로 공중 부양이라고 판단할 근거는 없다. 열린 피아노 뚜껑에는 왼쪽의 대각선 지지 구조가 보이고, 현과 금속 프레임은 피아노 몸체 안에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "지소영은 화면 왼쪽 위, 바로 앞에 선 찰리의 화면 밖 머리 위치를 향해 눈과 얼굴을 올린다. 찰리의 장갑 몸통과 팔이 왼쪽 전경에 있어 시선의 상대가 명확하다. 찰리의 얼굴은 잘려 있어 그의 시선 방향은 확인할 수 없다.",
        "built_space": "오른쪽에 오래된 피아노 한 대와 건반이 있고, 지소영 뒤와 아래에 검은 피아노 벤치 하나가 보인다. 그녀는 건반 옆에서 찰리 쪽으로 상체를 돌린 배치다. 배경에는 곡면 유리 구획 하나, 오른쪽 로봇 팔 하나, 뒤쪽 작업대와 회전의자 하나가 보이며 참조의 재료와 공간 구성을 유지한다. 바닥 반사는 조명과 유리벽 위치에 비추어 가능하다. 건반과 공간을 보여주는 범위가 요청한 작은 하단 조각보다 훨씬 넓다.",
        "entities": "지소영의 중년 얼굴과 단정한 검은 머리는 참조에 대체로 부합하며 흰 라펠 연구복은 참조 의상의 목깃 구조와 다르다. 왼쪽 찰리의 육중한 팔, 샌드 베이지 장갑판, 어두운 올리브색 외투는 참조와 맞는다. 잘린 얼굴과 다리는 평가할 수 없다. 피아노와 벤치는 실물로 읽히고, 추가 인물이나 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "지소영의 골반 부근은 하단에 가리지만 바로 뒤의 벤치가 착석을 지지하는 위치에 있으며, 찰리를 향해 상체와 목을 돌리는 자세는 가능하다. 찰리의 팔은 어깨와 몸통에 연결되어 아래로 내려오고 발은 프레임 밖이다. 피아노와 벤치는 바닥에 놓인 가구로 읽히며, 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.833,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.833,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 지소영이 연주를 멈추고 앉아있다는 설정과 달리, 피아노 건반이 카메라를 향하고 있어 인물이 피아노 반대편에 위치하는 물리적 모순 발생."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1833,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1833,
    "verdict_ko": "찰리의 어깨를 건 앵글로 지소영의 미소 짓는 얼굴을 포착했으며, 피아노와 배경 묘사가 지시사항에 부합하나 의상이 기준과 다소 다름."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "피아노 건반이 카메라를 향해 있어 연주자인 지소영의 위치와 물리적으로 모순되며, 건반이 프레임을 과도하게 차지함.  ★위반: [gemini-pro] 지소영이 연주를 멈추고 앉아있다는 설정과 달리, 피아노 건반이 카메라를 향하고 있어 인물이 피아노 반대편에 위치하는 물리적 모순 발생."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L269B03.png",
    "asset_id": "2d7bcd3a-3137-473f-b939-ac3ee04b6230",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:703308>",
    "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-0394-72a2-b316-70f8a69d324d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11__bgfirst_bg.png",
   "bg_asset_id": "d36617d7-e206-4f23-93f4-d6161a0fbc32",
   "bg_record_key": "S84sh11::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "staged_characters_added": [
   "C06"
  ]
 },
 "S84sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:04:46.523830+00:00",
  "fingerprint": "542bc7585aec098f238f76d0974b872468b2cb66040862e8501f04244a319040",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S84sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S84sh11_sel.png",
  "source_sha256": "dcad5fcc993dec50baa1e2902997dc5c0ad44c4d2b515c3753740d1e2391dbf6",
  "file": "S84sh11_cine.png",
  "staged_sha256": "b70a9f6fa89d3dee1b8d7bbb4dbe94e247e2ef5f78184dfd3684b6000a9c26f6",
  "latency_ms": 11919
 },
 "S84sh14::signage": {
  "fp": "9e0bd6fd99ce0dfb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S84sh14": {
  "input_fingerprint": "6864efa4d4e6d522",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영의 몸을 양팔로 빈틈없이 감싸 안은 찰리의 밀착된 상체.\n\nLOCATION (lock): At the piano inside the director's spacious modern suite adjoining the research center, in subdued nighttime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Piano (Old piano, no longer being played) — A partial side and keyboard edge remain behind the joined figures; used as Soft contextual reminder of the music that prompted recognition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and gentle tonal separation preserve the tenderness of the embrace without introducing a new light source or perceptual effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 지소영, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the director's office. Charlie is free of the laboratory wires, with both arms closed in an embrace. 지소영: She remains in her neat research coat near the piano, her eyes gently closed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영의 몸을 양팔로 빈틈없이 감싸 안은 찰리의 밀착된 상체.\n\nLOCATION (lock): At the piano inside the director's spacious modern suite adjoining the research center, in subdued nighttime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Piano (Old piano, no longer being played) — A partial side and keyboard edge remain behind the joined figures; used as Soft contextual reminder of the music that prompted recognition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and gentle tonal separation preserve the tenderness of the embrace without introducing a new light source or perceptual effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 지소영, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the director's office. Charlie is free of the laboratory wires, with both arms closed in an embrace. 지소영: She remains in her neat research coat near the piano, her eyes gently closed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영의 몸을 양팔로 빈틈없이 감싸 안은 찰리의 밀착된 상체.\n\nLOCATION (lock): At the piano inside the director's spacious modern suite adjoining the research center, in subdued nighttime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Piano (Old piano, no longer being played) — A partial side and keyboard edge remain behind the joined figures; used as Soft contextual reminder of the music that prompted recognition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and gentle tonal separation preserve the tenderness of the embrace without introducing a new light source or perceptual effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 지소영, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the director's office. Charlie is free of the laboratory wires, with both arms closed in an embrace. 지소영: She remains in her neat research coat near the piano, her eyes gently closed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 두 팔로 지소영을 틈 없이 감싸 안고 있으며, 지소영은 찰리의 가슴에 기대어 눈을 감고 있음.",
    "built_space": "이전 샷과 연속성이 느껴지는 원형 유리 벽의 연구실이 배경에 보이며, 화면 우측에 피아노가 배치됨.",
    "entities": "찰리는 트렌치코트를 입은 육중한 로봇 외형을 갖춤. 지소영은 지정된 단정한 회색 연구복을 착용함. 피아노 건반 덮개에 'YATKAHA'와 같은 읽을 수 있는 금지된 텍스트가 나타남.",
    "hard_violations": [
     "[gemini-pro] 금지된 텍스트 노출 (피아노에 'YATKAHA'와 같은 브랜드 로고 텍스트가 선명하게 보임)",
     "[gpt-high] 피아노 전면에 읽을 수 있는 영문 상표·로고가 노출되어, 읽을 수 있는 글자와 로고를 전면 금지한 지시를 위반한다."
    ],
    "physics": "두 사람은 바닥에 안정적으로 서서 서로에게 체중을 싣고 자연스럽게 포옹하고 있음."
   },
   {
    "label": "B",
    "direction": "찰리가 왼쪽 팔로 지소영을 깊게 끌어안고, 지소영은 찰리에게 안겨 평온하게 눈을 감고 있음.",
    "built_space": "어두운 조명의 모던한 실내이며 두 사람 바로 뒤쪽에 피아노가 위치해 있음.",
    "entities": "지소영은 회색 연구복 차림으로 기준에 부합함. 찰리는 팔 부분의 녹색 소매만 보이고 코트의 몸통 부분이 누락되어 로봇 골격이 그대로 드러남.",
    "hard_violations": [],
    "physics": "두 사람은 바닥에 지탱하여 서 있으며, 포옹하는 팔과 상체의 밀착이 물리적으로 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시된 구도와 배경의 연속성은 좋으나, 피아노 덮개에 금지된 텍스트 로고가 명확하게 노출된 점이 치명적인 프롬프트 위반입니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 트렌치코트 묘사가 불완전하지만, 치명적인 텍스트 노출 위반 없이 두 사람의 밀착된 포옹과 피아노가 있는 배경을 무난하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 두 팔로 지소영을 틈 없이 감싸 안고 있으며, 지소영은 찰리의 가슴에 기대어 눈을 감고 있음.",
        "built_space": "이전 샷과 연속성이 느껴지는 원형 유리 벽의 연구실이 배경에 보이며, 화면 우측에 피아노가 배치됨.",
        "entities": "찰리는 트렌치코트를 입은 육중한 로봇 외형을 갖춤. 지소영은 지정된 단정한 회색 연구복을 착용함. 피아노 건반 덮개에 'YATKAHA'와 같은 읽을 수 있는 금지된 텍스트가 나타남.",
        "hard_violations": [
         "금지된 텍스트 노출 (피아노에 'YATKAHA'와 같은 브랜드 로고 텍스트가 선명하게 보임)"
        ],
        "physics": "두 사람은 바닥에 안정적으로 서서 서로에게 체중을 싣고 자연스럽게 포옹하고 있음."
       },
       {
        "label": "B",
        "direction": "찰리가 왼쪽 팔로 지소영을 깊게 끌어안고, 지소영은 찰리에게 안겨 평온하게 눈을 감고 있음.",
        "built_space": "어두운 조명의 모던한 실내이며 두 사람 바로 뒤쪽에 피아노가 위치해 있음.",
        "entities": "지소영은 회색 연구복 차림으로 기준에 부합함. 찰리는 팔 부분의 녹색 소매만 보이고 코트의 몸통 부분이 누락되어 로봇 골격이 그대로 드러남.",
        "hard_violations": [],
        "physics": "두 사람은 바닥에 지탱하여 서 있으며, 포옹하는 팔과 상체의 밀착이 물리적으로 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시된 구도와 배경의 연속성은 좋으나, 피아노 덮개에 금지된 텍스트 로고가 명확하게 노출된 점이 치명적인 프롬프트 위반입니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 트렌치코트 묘사가 불완전하지만, 치명적인 텍스트 노출 위반 없이 두 사람의 밀착된 포옹과 피아노가 있는 배경을 무난하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 두 팔로 지소영을 틈 없이 감싸 안고 있으며, 지소영은 찰리의 가슴에 기대어 눈을 감고 있음.",
        "built_space": "이전 샷과 연속성이 느껴지는 원형 유리 벽의 연구실이 배경에 보이며, 화면 우측에 피아노가 배치됨.",
        "entities": "찰리는 트렌치코트를 입은 육중한 로봇 외형을 갖춤. 지소영은 지정된 단정한 회색 연구복을 착용함. 피아노 건반 덮개에 'YATKAHA'와 같은 읽을 수 있는 금지된 텍스트가 나타남.",
        "hard_violations": [
         "금지된 텍스트 노출 (피아노에 'YATKAHA'와 같은 브랜드 로고 텍스트가 선명하게 보임)"
        ],
        "physics": "두 사람은 바닥에 안정적으로 서서 서로에게 체중을 싣고 자연스럽게 포옹하고 있음."
       },
       {
        "label": "B",
        "direction": "찰리가 왼쪽 팔로 지소영을 깊게 끌어안고, 지소영은 찰리에게 안겨 평온하게 눈을 감고 있음.",
        "built_space": "어두운 조명의 모던한 실내이며 두 사람 바로 뒤쪽에 피아노가 위치해 있음.",
        "entities": "지소영은 회색 연구복 차림으로 기준에 부합함. 찰리는 팔 부분의 녹색 소매만 보이고 코트의 몸통 부분이 누락되어 로봇 골격이 그대로 드러남.",
        "hard_violations": [],
        "physics": "두 사람은 바닥에 지탱하여 서 있으며, 포옹하는 팔과 상체의 밀착이 물리적으로 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "양팔로 밀착해 안는 미디엄 숏과 감은 눈을 구현하고 금지된 글자도 없지만, 피아노의 정면 노출이 크고 찰리의 코트와 배경 조명은 참조와 차이가 난다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "밀착 포옹, 연구실 배경과 코트의 연속성은 더 충실하지만, 피아노에 읽을 수 있는 영문 상표가 있어 글자·로고 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개를 지소영의 정수리 쪽으로 숙이고, 지소영은 눈을 감은 채 얼굴을 찰리의 가슴에 붙인다. 찰리의 두 팔은 지소영의 몸을 향해 닫혀 있으며, 한 손은 위쪽 등과 어깨를, 다른 손은 위팔 부근을 감싼다. 피아노를 연주하는 손은 없다.",
        "built_space": "두 인물 뒤에 검은 피아노 한 대가 있으며, 건반이 인물 양옆으로 넓게 드러난다. 요청한 부분적인 측면과 건반 가장자리보다 정면의 비중이 크다. 왼쪽에는 유리 칸막이, 오른쪽에는 수납장과 화병, 뒤쪽에는 따뜻한 조명이 보인다. 참조의 유리와 광택 바닥은 이어지지만 연구실 설비는 보이지 않고, 따뜻한 조명과 가구 배치는 참조에서 확인되지 않는다. 불가능한 인물 반사나 중복 피아노는 보이지 않는다.",
        "entities": "찰리와 지소영 두 인물만 보인다. 지소영은 중년 한국인 여성으로 보이며 정돈된 짧은 검은 머리, 밝은 연구 코트와 감은 눈이 맞는다. 찰리의 육중한 체격, 긴 장갑 팔, 샌드 베이지 장갑판과 밝은 기계식 마스크 얼굴은 맞지만, 캐릭터 참조의 모자가 없고 코트 앞자락 대신 가슴 장갑이 크게 노출된다. 검은 피아노가 있으며 실험실 연결선이나 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 세워져 있고 서로 접촉한다. 찰리의 팔은 어깨와 팔꿈치에서 이어져 지소영을 감싸며 양손도 몸에 닿아 있다. 지소영의 팔은 찰리의 허리를 감싼다. 발과 바닥 접점은 프레임 밖이지만 몸이 들리거나 떠 있는 징후는 없다. 피아노는 하부 구조로 지지되며 공중에 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 지소영의 머리 위로 숙여져 있고, 지소영은 눈을 감고 찰리의 가슴 쪽에 얼굴을 기댄다. 찰리의 두 팔과 손은 지소영의 등과 옆구리 쪽으로 닫혀 밀착 포옹을 만든다. 지소영도 찰리의 몸통을 감싸며, 누구도 건반을 향해 손을 뻗지 않는다.",
        "built_space": "피아노 한 대의 측면과 건반 일부가 오른쪽 배경에 남아 있다. 뒤에는 곡면 유리 칸막이, 원호형 천장 조명, 여러 연구용 모니터와 작업대, 검은 의자 한 개가 보여 이전 숏의 공간과 잘 이어진다. 인물은 피아노 앞에 밀착해 있고 건반이나 가구를 관통하지 않는다. 불가능한 반사나 고정 설비의 명백한 중복은 보이지 않는다.",
        "entities": "찰리와 지소영만 등장한다. 지소영의 중년 얼굴, 짧고 정돈된 검은 머리, 밝은 연구 코트와 감은 눈이 참조 및 지시와 맞는다. 찰리는 베이지 장갑판, 육중한 긴 팔과 올리브색 코트를 유지하지만 참조의 모자는 없다. 얼굴은 아래로 숙여 일부 가려진다. 실험실 연결선은 없고 검은 피아노가 보이지만, 피아노 전면에 판독 가능한 영문 상표가 남아 있다.",
        "hard_violations": [
         "피아노 전면에 읽을 수 있는 영문 상표·로고가 노출되어, 읽을 수 있는 글자와 로고를 전면 금지한 지시를 위반한다."
        ],
        "physics": "찰리의 양팔은 몸통에서 자연스럽게 이어져 지소영을 둘러싸며, 겹친 기계 손들이 지소영의 등과 옆구리에 접촉한다. 지소영의 보이는 손도 찰리의 옆구리에 닿아 있다. 하체와 발은 프레임 밖이므로 지면 접점은 확인되지 않지만, 들어 올리거나 공중에 매달린 자세는 아니다. 피아노 역시 정상적인 고정 가구로 보이며 지지 없이 떠 있는 대상은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "양팔로 밀착해 안는 미디엄 숏과 감은 눈을 구현하고 금지된 글자도 없지만, 피아노의 정면 노출이 크고 찰리의 코트와 배경 조명은 참조와 차이가 난다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "밀착 포옹, 연구실 배경과 코트의 연속성은 더 충실하지만, 피아노에 읽을 수 있는 영문 상표가 있어 글자·로고 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 고개를 지소영의 정수리 쪽으로 숙이고, 지소영은 눈을 감은 채 얼굴을 찰리의 가슴에 붙인다. 찰리의 두 팔은 지소영의 몸을 향해 닫혀 있으며, 한 손은 위쪽 등과 어깨를, 다른 손은 위팔 부근을 감싼다. 피아노를 연주하는 손은 없다.",
        "built_space": "두 인물 뒤에 검은 피아노 한 대가 있으며, 건반이 인물 양옆으로 넓게 드러난다. 요청한 부분적인 측면과 건반 가장자리보다 정면의 비중이 크다. 왼쪽에는 유리 칸막이, 오른쪽에는 수납장과 화병, 뒤쪽에는 따뜻한 조명이 보인다. 참조의 유리와 광택 바닥은 이어지지만 연구실 설비는 보이지 않고, 따뜻한 조명과 가구 배치는 참조에서 확인되지 않는다. 불가능한 인물 반사나 중복 피아노는 보이지 않는다.",
        "entities": "찰리와 지소영 두 인물만 보인다. 지소영은 중년 한국인 여성으로 보이며 정돈된 짧은 검은 머리, 밝은 연구 코트와 감은 눈이 맞는다. 찰리의 육중한 체격, 긴 장갑 팔, 샌드 베이지 장갑판과 밝은 기계식 마스크 얼굴은 맞지만, 캐릭터 참조의 모자가 없고 코트 앞자락 대신 가슴 장갑이 크게 노출된다. 검은 피아노가 있으며 실험실 연결선이나 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 세워져 있고 서로 접촉한다. 찰리의 팔은 어깨와 팔꿈치에서 이어져 지소영을 감싸며 양손도 몸에 닿아 있다. 지소영의 팔은 찰리의 허리를 감싼다. 발과 바닥 접점은 프레임 밖이지만 몸이 들리거나 떠 있는 징후는 없다. 피아노는 하부 구조로 지지되며 공중에 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 지소영의 머리 위로 숙여져 있고, 지소영은 눈을 감고 찰리의 가슴 쪽에 얼굴을 기댄다. 찰리의 두 팔과 손은 지소영의 등과 옆구리 쪽으로 닫혀 밀착 포옹을 만든다. 지소영도 찰리의 몸통을 감싸며, 누구도 건반을 향해 손을 뻗지 않는다.",
        "built_space": "피아노 한 대의 측면과 건반 일부가 오른쪽 배경에 남아 있다. 뒤에는 곡면 유리 칸막이, 원호형 천장 조명, 여러 연구용 모니터와 작업대, 검은 의자 한 개가 보여 이전 숏의 공간과 잘 이어진다. 인물은 피아노 앞에 밀착해 있고 건반이나 가구를 관통하지 않는다. 불가능한 반사나 고정 설비의 명백한 중복은 보이지 않는다.",
        "entities": "찰리와 지소영만 등장한다. 지소영의 중년 얼굴, 짧고 정돈된 검은 머리, 밝은 연구 코트와 감은 눈이 참조 및 지시와 맞는다. 찰리는 베이지 장갑판, 육중한 긴 팔과 올리브색 코트를 유지하지만 참조의 모자는 없다. 얼굴은 아래로 숙여 일부 가려진다. 실험실 연결선은 없고 검은 피아노가 보이지만, 피아노 전면에 판독 가능한 영문 상표가 남아 있다.",
        "hard_violations": [
         "피아노 전면에 읽을 수 있는 영문 상표·로고가 노출되어, 읽을 수 있는 글자와 로고를 전면 금지한 지시를 위반한다."
        ],
        "physics": "찰리의 양팔은 몸통에서 자연스럽게 이어져 지소영을 둘러싸며, 겹친 기계 손들이 지소영의 등과 옆구리에 접촉한다. 지소영의 보이는 손도 찰리의 옆구리에 닿아 있다. 하체와 발은 프레임 밖이므로 지면 접점은 확인되지 않지만, 들어 올리거나 공중에 매달린 자세는 아니다. 피아노 역시 정상적인 고정 가구로 보이며 지지 없이 떠 있는 대상은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.857,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.607,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 금지된 텍스트 노출 (피아노에 'YATKAHA'와 같은 브랜드 로고 텍스트가 선명하게 보임)",
     "[gpt-high] 피아노 전면에 읽을 수 있는 영문 상표·로고가 노출되어, 읽을 수 있는 글자와 로고를 전면 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 607,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 607,
    "verdict_ko": "지시된 구도와 배경의 연속성은 좋으나, 피아노 덮개에 금지된 텍스트 로고가 명확하게 노출된 점이 치명적인 프롬프트 위반입니다.  ★위반: [gemini-pro] 금지된 텍스트 노출 (피아노에 'YATKAHA'와 같은 브랜드 로고 텍스트가 선명하게 보임) / [gpt-high] 피아노 전면에 읽을 수 있는 영문 상표·로고가 노출되어, 읽을 수 있는 글자와 로고를 전면 금지한 지시를 위반한다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "찰리의 트렌치코트 묘사가 불완전하지만, 치명적인 텍스트 노출 위반 없이 두 사람의 밀착된 포옹과 피아노가 있는 배경을 무난하게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 지소영, 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11_sel.png",
    "asset_id": "464f5b11-01db-4cd7-a576-000710c49a03",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:703308>",
    "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-06e5-713b-856e-af59d3ddf554",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S84sh11"
  }
 },
 "S84sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:06:11.505838+00:00",
  "fingerprint": "1f0d6ec49dad470734c531bdd4e82c5ecc70c2fd2a1ae97bed90e1230143e49f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S84sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S84sh14_sel.png",
  "source_sha256": "d0ff3610d593ce703954da681ee9cc1eccd6dfef3eea689e0fd5b66005de6aa2",
  "file": "S84sh14_cine.png",
  "staged_sha256": "8679abce7f0be889abfffdc39708129a01db98c590c98c4dde8dbae604f2cfc4",
  "latency_ms": 10594
 },
 "S85sh11::signage": {
  "fp": "c0fa537ec8cefb43",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S85sh11": {
  "input_fingerprint": "401150e00cda9925",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 회전 중인 한 위치에서 멈춘 가슴 링 틈새에서 짙은 검은 연기가 뿜어져 나오는 순간.\n\nLOCATION (lock): Inside the main research center's glass test enclosure, at the wired test position under laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chest ring (Still rotating and slowing as black smoke emerges from its gaps) — The front and one recessed gap are visible obliquely; used as Primary mechanical detail within the larger torso composition; Attached wires (Connected to 찰리's body); used as Peripheral evidence of the ongoing experiment; Black smoke (Emerging from the ring gap); used as Localized visual evidence of failure, leaving the ring's position legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient illumination with controlled contrast distinguishes the dense black smoke from the surrounding torso without adding an unsupported glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is seated behind the glass wall with multiple wires reattached; black smoke rises from the chest ring as its rotation slows and fails. The experiment display reads “ERROR,” and the previously rising output graphs are falling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 회전 중인 한 위치에서 멈춘 가슴 링 틈새에서 짙은 검은 연기가 뿜어져 나오는 순간.\n\nLOCATION (lock): Inside the main research center's glass test enclosure, at the wired test position under laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chest ring (Still rotating and slowing as black smoke emerges from its gaps) — The front and one recessed gap are visible obliquely; used as Primary mechanical detail within the larger torso composition; Attached wires (Connected to 찰리's body); used as Peripheral evidence of the ongoing experiment; Black smoke (Emerging from the ring gap); used as Localized visual evidence of failure, leaving the ring's position legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient illumination with controlled contrast distinguishes the dense black smoke from the surrounding torso without adding an unsupported glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is seated behind the glass wall with multiple wires reattached; black smoke rises from the chest ring as its rotation slows and fails. The experiment display reads “ERROR,” and the previously rising output graphs are falling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 회전 중인 한 위치에서 멈춘 가슴 링 틈새에서 짙은 검은 연기가 뿜어져 나오는 순간.\n\nLOCATION (lock): Inside the main research center's glass test enclosure, at the wired test position under laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chest ring (Still rotating and slowing as black smoke emerges from its gaps) — The front and one recessed gap are visible obliquely; used as Primary mechanical detail within the larger torso composition; Attached wires (Connected to 찰리's body); used as Peripheral evidence of the ongoing experiment; Black smoke (Emerging from the ring gap); used as Localized visual evidence of failure, leaving the ring's position legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient illumination with controlled contrast distinguishes the dense black smoke from the surrounding torso without adding an unsupported glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is seated behind the glass wall with multiple wires reattached; black smoke rises from the chest ring as its rotation slows and fails. The experiment display reads “ERROR,” and the previously rising output graphs are falling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선은 위를 향하며, 가슴 링 틈새에서 검은 연기가 위로 뿜어져 나옴.",
    "built_space": "이전 샷과 동일하게 유리 벽 내부의 평평한 금속 테이블 위에 피사체가 위치함.",
    "entities": "찰리의 외형, 가슴 링, 연결된 여러 가닥의 전선, 그리고 검은 연기 모두 프롬프트 및 참조 이미지와 일치함.",
    "hard_violations": [],
    "physics": "몸은 평평한 테이블에 의해 완전히 지지되고 있으며, 가슴 링 내부 부품의 회전이 모션 블러로 자연스럽게 표현됨."
   },
   {
    "label": "B",
    "direction": "찰리의 시선은 정면 아래를 향하며, 가슴 링에서 검은 연기가 위로 솟아오름.",
    "built_space": "유리 벽 구조 내부에 위치하나, 이전 샷에서 두껍고 평평했던 금속 테이블이 등받이처럼 구부러져 있어 기존 공간 설정과 충돌함.",
    "entities": "찰리의 외형, 가슴 링, 연결된 전선, 검은 연기가 존재함.",
    "hard_violations": [
     "[gemini-pro] 이전 샷에서 확립된 고정 구조물(평평한 금속 테이블)의 형태를 임의로 구부려 연속성을 심각하게 위반함"
    ],
    "physics": "변형된 테이블에 기대어 앉아 있으며, 연기는 상승하지만 가슴 링의 회전 움직임은 나타나지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 구조와 자세를 정확히 유지하면서 회전하는 가슴 링과 연기를 클로즈업으로 충실히 연출함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷의 평평한 금속 테이블을 억지로 구부려 고정된 공간 구조를 훼손하였으며, 링의 회전감을 보여주지 못함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 위를 향하며, 가슴 링 틈새에서 검은 연기가 위로 뿜어져 나옴.",
        "built_space": "이전 샷과 동일하게 유리 벽 내부의 평평한 금속 테이블 위에 피사체가 위치함.",
        "entities": "찰리의 외형, 가슴 링, 연결된 여러 가닥의 전선, 그리고 검은 연기 모두 프롬프트 및 참조 이미지와 일치함.",
        "hard_violations": [],
        "physics": "몸은 평평한 테이블에 의해 완전히 지지되고 있으며, 가슴 링 내부 부품의 회전이 모션 블러로 자연스럽게 표현됨."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면 아래를 향하며, 가슴 링에서 검은 연기가 위로 솟아오름.",
        "built_space": "유리 벽 구조 내부에 위치하나, 이전 샷에서 두껍고 평평했던 금속 테이블이 등받이처럼 구부러져 있어 기존 공간 설정과 충돌함.",
        "entities": "찰리의 외형, 가슴 링, 연결된 전선, 검은 연기가 존재함.",
        "hard_violations": [
         "이전 샷에서 확립된 고정 구조물(평평한 금속 테이블)의 형태를 임의로 구부려 연속성을 심각하게 위반함"
        ],
        "physics": "변형된 테이블에 기대어 앉아 있으며, 연기는 상승하지만 가슴 링의 회전 움직임은 나타나지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 구조와 자세를 정확히 유지하면서 회전하는 가슴 링과 연기를 클로즈업으로 충실히 연출함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷의 평평한 금속 테이블을 억지로 구부려 고정된 공간 구조를 훼손하였으며, 링의 회전감을 보여주지 못함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 위를 향하며, 가슴 링 틈새에서 검은 연기가 위로 뿜어져 나옴.",
        "built_space": "이전 샷과 동일하게 유리 벽 내부의 평평한 금속 테이블 위에 피사체가 위치함.",
        "entities": "찰리의 외형, 가슴 링, 연결된 여러 가닥의 전선, 그리고 검은 연기 모두 프롬프트 및 참조 이미지와 일치함.",
        "hard_violations": [],
        "physics": "몸은 평평한 테이블에 의해 완전히 지지되고 있으며, 가슴 링 내부 부품의 회전이 모션 블러로 자연스럽게 표현됨."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면 아래를 향하며, 가슴 링에서 검은 연기가 위로 솟아오름.",
        "built_space": "유리 벽 구조 내부에 위치하나, 이전 샷에서 두껍고 평평했던 금속 테이블이 등받이처럼 구부러져 있어 기존 공간 설정과 충돌함.",
        "entities": "찰리의 외형, 가슴 링, 연결된 전선, 검은 연기가 존재함.",
        "hard_violations": [
         "이전 샷에서 확립된 고정 구조물(평평한 금속 테이블)의 형태를 임의로 구부려 연속성을 심각하게 위반함"
        ],
        "physics": "변형된 테이블에 기대어 앉아 있으며, 연기는 상승하지만 가슴 링의 회전 움직임은 나타나지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "가슴보다 얼굴·허벅지·연구실까지 넓게 담고 이전 장면보다 상체를 세웠으며, 연기가 링 틈새가 아닌 중앙 구멍에서 나와 요구된 기계 디테일 클로즈업에 미달한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "기댄 자세와 더 밀착된 가슴 중심 구도가 이전 장면 및 촬영 지시에 더 가깝지만, 중앙 구멍의 연기와 링 아래 별도 회전체 표현은 지정된 링 틈새 고장 묘사와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 눈은 화면 오른쪽 앞쪽을 향하고 있으며 특정 대상은 보이지 않는다. 검은 연기는 가슴 링 중앙의 열린 구멍에서 위로 솟는다. 지정된 둘레의 움푹 들어간 틈새에서 분출되는 모습은 식별되지 않는다.",
        "built_space": "금속 시험 의자 한 개에 찰리 한 명이 앉아 있고, 뒤에는 유리벽과 세로 이음새, 바닥 가장자리 배수 격자가 보인다. 유리 너머에는 검은 모니터 두 개와 사무용 의자 두 개, 작업대 및 하부 수납장이 보인다. 연구실 재질과 분위기는 참조와 유사하지만 촬영 범위가 넓어 배경 비중이 커졌다. 불가능한 반사는 보이지 않는다.",
        "entities": "흰 각진 얼굴 마스크, 정상적인 눈, 긁힌 샌드 베이지 장갑판, 육중한 팔은 이전 장면의 찰리와 부합한다. 중앙 가슴 링 한 개와 왼쪽에서 몸 뒤로 이어지는 여러 전선, 짙은 검은 연기가 보인다. 연결 단자는 몸에 가려져 있다. 다른 인물과 읽을 수 있는 글자는 없다. 링의 회전이나 감속은 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "몸은 금속 시험 의자의 좌석과 등받이에 지지되어 있으며 팔은 몸 옆으로 내려와 있다. 다만 참조보다 상체가 훨씬 세워져 있다. 전선은 의자 가장자리를 따라 중력 방향으로 늘어지고, 링은 흉부 장갑에 고정되어 있다. 연기는 구멍에서 연속적으로 올라오므로 떠 있는 고체나 지지 없는 신체는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 고개를 뒤로 기댄 채 화면 위쪽을 바라보며, 시선의 대상은 프레임 밖이다. 검은 연기는 가슴 링의 중앙 구멍에서 화면 오른쪽 위로 뻗는다. 링 앞면과 비스듬한 깊이는 보이지만 연기의 출발점은 지정된 좁은 틈새보다 중앙 개구부로 읽힌다.",
        "built_space": "찰리 뒤로 금속 시험 의자 한 개의 등받이와 테두리가 보이고, 상단 및 오른쪽 배경에는 유리 enclosure의 하부 경계와 바닥 격자가 이어진다. 오른쪽 바닥에는 케이블이 놓여 있다. 찰리는 참조처럼 등받이에 기댄 위치에 있으며, 배경 설비는 몸통 중심의 가까운 구도 밖으로 대부분 제외되었다. 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 흰 마스크형 얼굴, 정상적인 눈, 마모된 베이지 장갑과 굵은 팔이 보이며 이전 장면의 외형을 잘 따른다. 가슴의 금속 링 한 개, 왼쪽에서 몸 뒤로 이어지는 여러 전선과 검은 연기가 있다. 다만 가슴 링 아래에는 참조의 좁은 검은 통풍구와 다른 큰 회전 기구가 드러나며, 회전 흔적도 주로 여기에 나타난다. 다른 인물이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "등과 몸통은 뒤의 금속 등받이에 받쳐진 기댄 자세이고, 보이는 팔도 어깨와 팔꿈치 관절에 연결되어 있다. 링과 아래 회전 기구는 흉부 내부에 장착된 것으로 보이며 독립적으로 공중에 떠 있지 않다. 전선은 몸 뒤와 의자 옆으로 처지고, 연기는 가슴 개구부에 연결된 연속적인 기둥을 이룬다. 회전 흐림은 있지만 그것만으로 감속까지 확인되지는 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "가슴보다 얼굴·허벅지·연구실까지 넓게 담고 이전 장면보다 상체를 세웠으며, 연기가 링 틈새가 아닌 중앙 구멍에서 나와 요구된 기계 디테일 클로즈업에 미달한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "기댄 자세와 더 밀착된 가슴 중심 구도가 이전 장면 및 촬영 지시에 더 가깝지만, 중앙 구멍의 연기와 링 아래 별도 회전체 표현은 지정된 링 틈새 고장 묘사와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 눈은 화면 오른쪽 앞쪽을 향하고 있으며 특정 대상은 보이지 않는다. 검은 연기는 가슴 링 중앙의 열린 구멍에서 위로 솟는다. 지정된 둘레의 움푹 들어간 틈새에서 분출되는 모습은 식별되지 않는다.",
        "built_space": "금속 시험 의자 한 개에 찰리 한 명이 앉아 있고, 뒤에는 유리벽과 세로 이음새, 바닥 가장자리 배수 격자가 보인다. 유리 너머에는 검은 모니터 두 개와 사무용 의자 두 개, 작업대 및 하부 수납장이 보인다. 연구실 재질과 분위기는 참조와 유사하지만 촬영 범위가 넓어 배경 비중이 커졌다. 불가능한 반사는 보이지 않는다.",
        "entities": "흰 각진 얼굴 마스크, 정상적인 눈, 긁힌 샌드 베이지 장갑판, 육중한 팔은 이전 장면의 찰리와 부합한다. 중앙 가슴 링 한 개와 왼쪽에서 몸 뒤로 이어지는 여러 전선, 짙은 검은 연기가 보인다. 연결 단자는 몸에 가려져 있다. 다른 인물과 읽을 수 있는 글자는 없다. 링의 회전이나 감속은 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "몸은 금속 시험 의자의 좌석과 등받이에 지지되어 있으며 팔은 몸 옆으로 내려와 있다. 다만 참조보다 상체가 훨씬 세워져 있다. 전선은 의자 가장자리를 따라 중력 방향으로 늘어지고, 링은 흉부 장갑에 고정되어 있다. 연기는 구멍에서 연속적으로 올라오므로 떠 있는 고체나 지지 없는 신체는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 고개를 뒤로 기댄 채 화면 위쪽을 바라보며, 시선의 대상은 프레임 밖이다. 검은 연기는 가슴 링의 중앙 구멍에서 화면 오른쪽 위로 뻗는다. 링 앞면과 비스듬한 깊이는 보이지만 연기의 출발점은 지정된 좁은 틈새보다 중앙 개구부로 읽힌다.",
        "built_space": "찰리 뒤로 금속 시험 의자 한 개의 등받이와 테두리가 보이고, 상단 및 오른쪽 배경에는 유리 enclosure의 하부 경계와 바닥 격자가 이어진다. 오른쪽 바닥에는 케이블이 놓여 있다. 찰리는 참조처럼 등받이에 기댄 위치에 있으며, 배경 설비는 몸통 중심의 가까운 구도 밖으로 대부분 제외되었다. 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 흰 마스크형 얼굴, 정상적인 눈, 마모된 베이지 장갑과 굵은 팔이 보이며 이전 장면의 외형을 잘 따른다. 가슴의 금속 링 한 개, 왼쪽에서 몸 뒤로 이어지는 여러 전선과 검은 연기가 있다. 다만 가슴 링 아래에는 참조의 좁은 검은 통풍구와 다른 큰 회전 기구가 드러나며, 회전 흔적도 주로 여기에 나타난다. 다른 인물이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "등과 몸통은 뒤의 금속 등받이에 받쳐진 기댄 자세이고, 보이는 팔도 어깨와 팔꿈치 관절에 연결되어 있다. 링과 아래 회전 기구는 흉부 내부에 장착된 것으로 보이며 독립적으로 공중에 떠 있지 않다. 전선은 몸 뒤와 의자 옆으로 처지고, 연기는 가슴 개구부에 연결된 연속적인 기둥을 이룬다. 회전 흐림은 있지만 그것만으로 감속까지 확인되지는 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.143
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.893
   },
   "violations": {
    "B": [
     "[gemini-pro] 이전 샷에서 확립된 고정 구조물(평평한 금속 테이블)의 형태를 임의로 구부려 연속성을 심각하게 위반함"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 893
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "이전 샷의 구조와 자세를 정확히 유지하면서 회전하는 가슴 링과 연기를 클로즈업으로 충실히 연출함."
   },
   {
    "label": "B",
    "score": 893,
    "verdict_ko": "이전 샷의 평평한 금속 테이블을 억지로 구부려 고정된 공간 구조를 훼손하였으며, 링의 회전감을 보여주지 못함.  ★위반: [gemini-pro] 이전 샷에서 확립된 고정 구조물(평평한 금속 테이블)의 형태를 임의로 구부려 연속성을 심각하게 위반함"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh1_sel.png",
    "asset_id": "5cc0176d-e827-436e-9e8d-0b98f7630b82",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-0890-7752-a0ba-7d83af01b158",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S84sh1"
  }
 },
 "S85sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:07:12.484272+00:00",
  "fingerprint": "70d432c0c1c8bd63f7ae88060c91327cf30cbae12fe6e14eaf870973bb2bc17b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S85sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S85sh11_sel.png",
  "source_sha256": "183b99f30fc03d60a70aa4573a32e26baeb7d53b8c22c0325a635170cef2e87e",
  "file": "S85sh11_cine.png",
  "staged_sha256": "b68fafa65059258b4ed83a9672563fd4840c20d071ec602407ebe7fbbf0e1810",
  "latency_ms": 10421
 },
 "S85sh22::signage": {
  "fp": "44894cb3220c2f6f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::374d1a2346548ec2": {
  "subjects": [],
  "subject_text": "제주도 연구소 메인 연구센터\n발전소처럼 거대한 실내 연구 공간. 각종 컴퓨터와 대형 연구 장치가 배치되고 열대 식물과 꽃이 기계 설비 사이를 채운다.",
  "identity": "canonical",
  "scope_id": "L268",
  "scope_role": "location_interior",
  "scope_sha": "3b89e33f78be0eaf"
 },
 "S85sh22::bgfirst_bg": {
  "input_fingerprint": "ef3385f9c5575580",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22__bgfirst_bg.png",
  "asset_id": "44478bdf-e0e9-48ca-bfd4-2ad6d414a7f9",
  "input_asset_ids": [
   "382f3829-e015-47c7-ab5e-b6be933588ec",
   "a491e625-4c37-4571-b3ba-85a3daf996f9"
  ]
 },
 "S85sh22": {
  "input_fingerprint": "f8f4b1bac3d4aead",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appetizing bread and other food remain on the dining table, still uneaten. 서지민: She is at the dining conversation in her research coat, abruptly checking herself after an unintended disclosure.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appetizing bread and other food remain on the dining table, still uneaten. 서지민: She is at the dining conversation in her research coat, abruptly checking herself after an unintended disclosure.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appetizing bread and other food remain on the dining table, still uneaten. 서지민: She is at the dining conversation in her research coat, abruptly checking herself after an unintended disclosure.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22__bgfirst_bg.png",
     "asset_id": "44478bdf-e0e9-48ca-bfd4-2ad6d414a7f9",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S85sh22.png",
     "asset_id": "382f3829-e015-47c7-ab5e-b6be933588ec",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:930556>",
     "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
     "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:930556>",
     "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "서지민의 시선은 화면 앞쪽에 앉은 남자를 향하고 있습니다.",
    "built_space": "식당 내부의 테이블과 좌석들이 구현되었으나, 제시된 레퍼런스 공간(랩실)과 전혀 일치하지 않습니다.",
    "entities": "서지민, 빵과 샐러드가 담긴 식판이 보입니다. 지문에 없는 앞사람(남자)과 다수의 식사하는 배경 인물들이 존재합니다.",
    "hard_violations": [
     "[gemini-pro] 지문에 없는 인물(전경의 남자 및 배경 인물들) 임의 추가",
     "[gpt-high] 허용되지 않은 대화 상대 남성의 머리와 상체를 전경에 추가했다.",
     "[gpt-high] 서지민만 등장해야 하는 장면에 여섯 명의 추가 배경 인물이 등장한다."
    ],
    "physics": "두 손을 나란히 모아 입을 막고 있으며, 자연스럽게 테이블 앞에 앉아 있습니다."
   },
   {
    "label": "B",
    "direction": "서지민의 시선은 화면 우측 허공을 향하고 있습니다.",
    "built_space": "식당 내부의 창문과 배식구, 테이블이 보이나, 제시된 레퍼런스 사진(유리벽 구조의 랩실)과 공간적 일치도가 없습니다.",
    "entities": "서지민(지정된 디자인과 유사한 연구복 착용), 테이블 위의 빵과 수프가 보입니다. 지문에 언급되지 않은 다수의 배경 인물들이 존재합니다.",
    "hard_violations": [
     "[gemini-pro] 지문에 없는 인물(배경 인물들) 임의 추가",
     "[gpt-high] 서지민만 등장해야 하는 장면에 다수의 추가 배경 인물이 등장한다."
    ],
    "physics": "양손을 교차하여 입을 막고 있으며, 신체는 테이블을 향해 안정적으로 위치해 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업 프레이밍과 지정된 연구복의 구현은 상대적으로 우수하나, 지문에 없는 배경 인물들이 다수 등장하여 주요 규정을 위반했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지문에 없는 앞사람이 등장하여 오버더숄더 샷으로 변질되었으며, 배경에도 다수의 인물이 임의로 추가되어 규정을 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "서지민의 시선은 화면 우측 허공을 향하고 있습니다.",
        "built_space": "식당 내부의 창문과 배식구, 테이블이 보이나, 제시된 레퍼런스 사진(유리벽 구조의 랩실)과 공간적 일치도가 없습니다.",
        "entities": "서지민(지정된 디자인과 유사한 연구복 착용), 테이블 위의 빵과 수프가 보입니다. 지문에 언급되지 않은 다수의 배경 인물들이 존재합니다.",
        "hard_violations": [
         "지문에 없는 인물(배경 인물들) 임의 추가"
        ],
        "physics": "양손을 교차하여 입을 막고 있으며, 신체는 테이블을 향해 안정적으로 위치해 있습니다."
       },
       {
        "label": "A",
        "direction": "서지민의 시선은 화면 앞쪽에 앉은 남자를 향하고 있습니다.",
        "built_space": "식당 내부의 테이블과 좌석들이 구현되었으나, 제시된 레퍼런스 공간(랩실)과 전혀 일치하지 않습니다.",
        "entities": "서지민, 빵과 샐러드가 담긴 식판이 보입니다. 지문에 없는 앞사람(남자)과 다수의 식사하는 배경 인물들이 존재합니다.",
        "hard_violations": [
         "지문에 없는 인물(전경의 남자 및 배경 인물들) 임의 추가"
        ],
        "physics": "두 손을 나란히 모아 입을 막고 있으며, 자연스럽게 테이블 앞에 앉아 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업 프레이밍과 지정된 연구복의 구현은 상대적으로 우수하나, 지문에 없는 배경 인물들이 다수 등장하여 주요 규정을 위반했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지문에 없는 앞사람이 등장하여 오버더숄더 샷으로 변질되었으며, 배경에도 다수의 인물이 임의로 추가되어 규정을 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "서지민의 시선은 화면 우측 허공을 향하고 있습니다.",
        "built_space": "식당 내부의 창문과 배식구, 테이블이 보이나, 제시된 레퍼런스 사진(유리벽 구조의 랩실)과 공간적 일치도가 없습니다.",
        "entities": "서지민(지정된 디자인과 유사한 연구복 착용), 테이블 위의 빵과 수프가 보입니다. 지문에 언급되지 않은 다수의 배경 인물들이 존재합니다.",
        "hard_violations": [
         "지문에 없는 인물(배경 인물들) 임의 추가"
        ],
        "physics": "양손을 교차하여 입을 막고 있으며, 신체는 테이블을 향해 안정적으로 위치해 있습니다."
       },
       {
        "label": "A",
        "direction": "서지민의 시선은 화면 앞쪽에 앉은 남자를 향하고 있습니다.",
        "built_space": "식당 내부의 테이블과 좌석들이 구현되었으나, 제시된 레퍼런스 공간(랩실)과 전혀 일치하지 않습니다.",
        "entities": "서지민, 빵과 샐러드가 담긴 식판이 보입니다. 지문에 없는 앞사람(남자)과 다수의 식사하는 배경 인물들이 존재합니다.",
        "hard_violations": [
         "지문에 없는 인물(전경의 남자 및 배경 인물들) 임의 추가"
        ],
        "physics": "두 손을 나란히 모아 입을 막고 있으며, 자연스럽게 테이블 앞에 앉아 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴 중심의 가까운 구도와 양손으로 입을 틀어막은 당황한 순간은 더 충실하지만, 금지된 배경 인물들과 장소 참조 불일치로 부적합하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "양손의 동작과 검은 단발머리는 부합하지만, 얼굴 클로즈업을 대화 상대의 어깨 너머 상반신 구도로 바꾸었고 추가 인물과 장소 불일치도 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "서지민의 눈은 화면 오른쪽 바깥을 향하며, 그 시선의 상대는 보이지 않는다. 두 손은 서로 포개져 자신의 입을 직접 덮는다. 배경 사람들은 식탁 상대나 배식대 쪽을 향한다. 지정된 시선 표적은 없으므로 주인공의 시선 자체는 위반이 아니다.",
        "built_space": "왼쪽에 연속된 큰 창, 오른쪽에 긴 배식대 하나와 그 위 금속 후드 하나가 있다. 배식대 아래에는 그릇이 놓인 열린 수납칸 하나가 보인다. 전경 식탁 하나, 뒤쪽 식탁과 검은 등받이 의자들이 있으며 주인공은 전경 식탁 바로 뒤에 있다. 낮의 구내식당으로는 읽히지만, 장소 참조의 원형 유리 실험실, 중앙 금속 실험대 하나, 양옆 장비 두 대, 벽면 모니터 배열과는 다른 공간이다.",
        "entities": "주인공 한 명 외에 배경 인물이 약 아홉 명 보인다. 주인공은 검은 머리의 젊은 동아시아계 여성으로 요구된 인물의 대략적인 연령·외형에 부합하지만, 참조의 어깨 길이로 풀어 놓은 머리와 달리 뒤로 묶었다. 흰색 연구복은 보이나 참조의 고글은 없다. 전경에 빵과 수프 그릇이 있으며 먹는 동작은 없다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "서지민만 등장해야 하는 장면에 다수의 추가 배경 인물이 등장한다."
        ],
        "physics": "두 손은 각각 소매 밖 손목과 자연스럽게 연결되어 입과 볼에 닿아 있고, 손가락의 겹침도 가능한 자세다. 빵과 수프 그릇은 식탁 위에 놓여 있다. 주인공의 좌면과 하체는 구도 밖이므로 지지 상태를 직접 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 배경 인물들은 의자에 앉거나 배식대 뒤에 서 있다."
       },
       {
        "label": "B",
        "direction": "서지민은 화면 왼쪽 전경에 있는 남성 쪽을 바라보며, 남성도 서지민을 향한다. 양손은 손끝을 위로 향하게 모아 자신의 입을 덮고 있다. 대화 중 놀란 반응으로는 자연스럽지만, 그 시선 상대를 실제 인물로 추가한 것은 허용되지 않는다.",
        "built_space": "왼쪽에는 연속된 창과 화분, 오른쪽에는 유리 가림막이 있는 배식대 하나, 그 위에는 검은 펜던트등 두 개와 긴 직사각형 천장 구조물이 보인다. 전경 식탁 하나와 여러 후경 식탁, 회색 등받이 의자들이 배치되어 있다. 서지민은 전경 식탁 뒤 의자에 앉아 있다. 식당의 배치는 물리적으로 자연스럽지만 장소 참조의 원형 유리벽과 중앙 실험 장비 공간을 재현하지 않았다.",
        "entities": "서지민 외에 왼쪽 전경 남성 한 명과 배경 인물 여섯 명이 보인다. 주인공은 앳된 동아시아계 여성이고 검은 어깨 길이 머리는 참조에 비교적 가깝다. 흰 연구 가운을 입었지만 참조의 높은 목깃 연구복과는 다르며 고글도 없다. 식탁에는 금속 식판, 빵, 채소류와 물잔이 보이고 음식은 먹지 않은 상태로 놓여 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "허용되지 않은 대화 상대 남성의 머리와 상체를 전경에 추가했다.",
         "서지민만 등장해야 하는 장면에 여섯 명의 추가 배경 인물이 등장한다."
        ],
        "physics": "양팔은 몸통에서 자연스럽게 올라오며 두 손바닥과 손가락이 입 주변에 접촉한다. 의자의 등받이가 몸 뒤에 있고 상체와 식탁의 관계도 앉은 자세로 성립한다. 식판과 물잔은 식탁에 지지되어 있다. 배경 착석자들은 각자의 의자를 사용하며, 떠 있거나 지지 없이 매달린 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴 중심의 가까운 구도와 양손으로 입을 틀어막은 당황한 순간은 더 충실하지만, 금지된 배경 인물들과 장소 참조 불일치로 부적합하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "양손의 동작과 검은 단발머리는 부합하지만, 얼굴 클로즈업을 대화 상대의 어깨 너머 상반신 구도로 바꾸었고 추가 인물과 장소 불일치도 있다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "서지민의 눈은 화면 오른쪽 바깥을 향하며, 그 시선의 상대는 보이지 않는다. 두 손은 서로 포개져 자신의 입을 직접 덮는다. 배경 사람들은 식탁 상대나 배식대 쪽을 향한다. 지정된 시선 표적은 없으므로 주인공의 시선 자체는 위반이 아니다.",
        "built_space": "왼쪽에 연속된 큰 창, 오른쪽에 긴 배식대 하나와 그 위 금속 후드 하나가 있다. 배식대 아래에는 그릇이 놓인 열린 수납칸 하나가 보인다. 전경 식탁 하나, 뒤쪽 식탁과 검은 등받이 의자들이 있으며 주인공은 전경 식탁 바로 뒤에 있다. 낮의 구내식당으로는 읽히지만, 장소 참조의 원형 유리 실험실, 중앙 금속 실험대 하나, 양옆 장비 두 대, 벽면 모니터 배열과는 다른 공간이다.",
        "entities": "주인공 한 명 외에 배경 인물이 약 아홉 명 보인다. 주인공은 검은 머리의 젊은 동아시아계 여성으로 요구된 인물의 대략적인 연령·외형에 부합하지만, 참조의 어깨 길이로 풀어 놓은 머리와 달리 뒤로 묶었다. 흰색 연구복은 보이나 참조의 고글은 없다. 전경에 빵과 수프 그릇이 있으며 먹는 동작은 없다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "서지민만 등장해야 하는 장면에 다수의 추가 배경 인물이 등장한다."
        ],
        "physics": "두 손은 각각 소매 밖 손목과 자연스럽게 연결되어 입과 볼에 닿아 있고, 손가락의 겹침도 가능한 자세다. 빵과 수프 그릇은 식탁 위에 놓여 있다. 주인공의 좌면과 하체는 구도 밖이므로 지지 상태를 직접 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 배경 인물들은 의자에 앉거나 배식대 뒤에 서 있다."
       },
       {
        "label": "A",
        "direction": "서지민은 화면 왼쪽 전경에 있는 남성 쪽을 바라보며, 남성도 서지민을 향한다. 양손은 손끝을 위로 향하게 모아 자신의 입을 덮고 있다. 대화 중 놀란 반응으로는 자연스럽지만, 그 시선 상대를 실제 인물로 추가한 것은 허용되지 않는다.",
        "built_space": "왼쪽에는 연속된 창과 화분, 오른쪽에는 유리 가림막이 있는 배식대 하나, 그 위에는 검은 펜던트등 두 개와 긴 직사각형 천장 구조물이 보인다. 전경 식탁 하나와 여러 후경 식탁, 회색 등받이 의자들이 배치되어 있다. 서지민은 전경 식탁 뒤 의자에 앉아 있다. 식당의 배치는 물리적으로 자연스럽지만 장소 참조의 원형 유리벽과 중앙 실험 장비 공간을 재현하지 않았다.",
        "entities": "서지민 외에 왼쪽 전경 남성 한 명과 배경 인물 여섯 명이 보인다. 주인공은 앳된 동아시아계 여성이고 검은 어깨 길이 머리는 참조에 비교적 가깝다. 흰 연구 가운을 입었지만 참조의 높은 목깃 연구복과는 다르며 고글도 없다. 식탁에는 금속 식판, 빵, 채소류와 물잔이 보이고 음식은 먹지 않은 상태로 놓여 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "허용되지 않은 대화 상대 남성의 머리와 상체를 전경에 추가했다.",
         "서지민만 등장해야 하는 장면에 여섯 명의 추가 배경 인물이 등장한다."
        ],
        "physics": "양팔은 몸통에서 자연스럽게 올라오며 두 손바닥과 손가락이 입 주변에 접촉한다. 의자의 등받이가 몸 뒤에 있고 상체와 식탁의 관계도 앉은 자세로 성립한다. 식판과 물잔은 식탁에 지지되어 있다. 배경 착석자들은 각자의 의자를 사용하며, 떠 있거나 지지 없이 매달린 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.417,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.167,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 지문에 없는 인물(배경 인물들) 임의 추가",
     "[gpt-high] 서지민만 등장해야 하는 장면에 다수의 추가 배경 인물이 등장한다."
    ],
    "A": [
     "[gemini-pro] 지문에 없는 인물(전경의 남자 및 배경 인물들) 임의 추가",
     "[gpt-high] 허용되지 않은 대화 상대 남성의 머리와 상체를 전경에 추가했다.",
     "[gpt-high] 서지민만 등장해야 하는 장면에 여섯 명의 추가 배경 인물이 등장한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 1167
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "클로즈업 프레이밍과 지정된 연구복의 구현은 상대적으로 우수하나, 지문에 없는 배경 인물들이 다수 등장하여 주요 규정을 위반했습니다.  ★위반: [gemini-pro] 지문에 없는 인물(배경 인물들) 임의 추가 / [gpt-high] 서지민만 등장해야 하는 장면에 다수의 추가 배경 인물이 등장한다."
   },
   {
    "label": "A",
    "score": 1167,
    "verdict_ko": "지문에 없는 앞사람이 등장하여 오버더숄더 샷으로 변질되었으며, 배경에도 다수의 인물이 임의로 추가되어 규정을 심각하게 위반했습니다.  ★위반: [gemini-pro] 지문에 없는 인물(전경의 남자 및 배경 인물들) 임의 추가 / [gpt-high] 허용되지 않은 대화 상대 남성의 머리와 상체를 전경에 추가했다. / [gpt-high] 서지민만 등장해야 하는 장면에 여섯 명의 추가 배경 인물이 등장한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
    "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:930556>",
    "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-0a40-7325-8503-9f034d1ec6a0",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22__bgfirst_bg.png",
   "bg_asset_id": "44478bdf-e0e9-48ca-bfd4-2ad6d414a7f9",
   "bg_record_key": "S85sh22::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S85sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:08:23.074033+00:00",
  "fingerprint": "a448c112dd90c4538b2f84452aa7025dcaa927d1712d1e693be08f3d3c8ddcbf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S85sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S85sh22_sel.png",
  "source_sha256": "9c4c263ccb0706d63ac2a1db9ddb3006c1ba07791bc8a811a75f851c43fa90d0",
  "file": "S85sh22_cine.png",
  "staged_sha256": "9196ff1ef648846e6b8612b62cf3f4cd6699cba882a4590bbd4e0499b746711e",
  "latency_ms": 10243
 },
 "S85sh23::signage": {
  "fp": "8760ec37abcff4aa",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S85sh23": {
  "input_fingerprint": "62564c475b62fc25",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두 주먹을 꽉 쥔 채 분노로 일그러진 표정으로 입을 크게 벌리고 있는 현우의 역동적인 상체.\n\nLOCATION (lock): At the same communal cafeteria table inside the research facility, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Bread and food remain in front of 현우) — The near edge runs diagonally across the bottom of the image; used as Maintains the dialogue geography and anchors the fists at a natural scale; Bread and food (Present on the table despite 현우's lack of interest in eating); used as Small peripheral context below the emotional action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the preceding ambient illumination and controlled contrast without a dramatic relighting for the anger.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The bread and other food remain uneaten on the dining table. 현우: He remains at the dining table, now visibly shocked and angry.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두 주먹을 꽉 쥔 채 분노로 일그러진 표정으로 입을 크게 벌리고 있는 현우의 역동적인 상체.\n\nLOCATION (lock): At the same communal cafeteria table inside the research facility, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Bread and food remain in front of 현우) — The near edge runs diagonally across the bottom of the image; used as Maintains the dialogue geography and anchors the fists at a natural scale; Bread and food (Present on the table despite 현우's lack of interest in eating); used as Small peripheral context below the emotional action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the preceding ambient illumination and controlled contrast without a dramatic relighting for the anger.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The bread and other food remain uneaten on the dining table. 현우: He remains at the dining table, now visibly shocked and angry.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두 주먹을 꽉 쥔 채 분노로 일그러진 표정으로 입을 크게 벌리고 있는 현우의 역동적인 상체.\n\nLOCATION (lock): At the same communal cafeteria table inside the research facility, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Bread and food remain in front of 현우) — The near edge runs diagonally across the bottom of the image; used as Maintains the dialogue geography and anchors the fists at a natural scale; Bread and food (Present on the table despite 현우's lack of interest in eating); used as Small peripheral context below the emotional action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the preceding ambient illumination and controlled contrast without a dramatic relighting for the anger.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The bread and other food remain uneaten on the dining table. 현우: He remains at the dining table, now visibly shocked and angry.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 좌측을 향하고 있으며, 크게 벌린 입과 표정으로 허공을 향해 분노를 터뜨리고 있습니다.",
    "built_space": "연구소 식당 내부로, 식탁 모서리가 화면 좌측 하단을 대각선으로 가로지르며 프롬프트의 배치 지시를 충실히 구현했습니다.",
    "entities": "현우(18세 남성, 헝클어진 머리)가 지시된 의상(흰색 옷)을 입고 있으며, 두 주먹을 꽉 쥔 채 탁자 위로 올리고 있습니다. 식탁에는 빵과 수프가 그대로 남아있습니다.",
    "hard_violations": [
     "[gemini-pro] 지시된 인물(현우) 외에는 절대 사람을 추가하지 말라는 지시(never add a person the shot text does not show)를 위반하여 우측 배경에 흐릿한 인물들이 등장함.",
     "[gpt-high] 현우 외 인물이 금지되어 있는데도 오른쪽 먼 배경에 여러 사람과 가장자리의 부분 신체를 추가했다."
    ],
    "physics": "보이지 않는 의자에 앉아 체중을 지탱하고 있으며, 양팔을 들어 주먹을 쥔 상체의 움직임이 자연스럽고 물리적으로 안정적입니다."
   },
   {
    "label": "B",
    "direction": "시선은 화면 우측 허공을 향하고 있으며, 분노하여 소리치는 표정입니다.",
    "built_space": "식당 내부이나, 식탁의 앞쪽 모서리가 대각선이 아닌 수평으로 반듯하게 배치되어 프레이밍 지시를 어겼습니다.",
    "entities": "현우가 지시된 의상을 입고 있으나, 주먹은 한쪽(화면 좌측)만 쥐고 있습니다. 식탁 위에는 빵과 식사가 놓여 있습니다.",
    "hard_violations": [
     "[gemini-pro] 이전 샷(PREVIOUS SHOT STILL)에 등장했던 인물들을 절대 포함하지 말라는 지시를 정면 위반하여 좌측의 여성, 중앙의 흑인 남성 등 배경 인물들을 그대로 복사해 넣음.",
     "[gemini-pro] 명시된 인물 외에 사람을 추가하지 말라는 지시 위반.",
     "[gemini-pro] 두 주먹을 꽉 쥐고(두 주먹) 있어야 하나 한쪽 주먹만 보임.",
     "[gemini-pro] 식탁 모서리가 화면 하단을 대각선으로 가로지르도록 프레이밍하라는 지시 위반(수평으로 배치됨).",
     "[gpt-high] 현우만 허용된 장면에 다수의 배경 식사객과 배식대 주변 인물을 추가했으며, 이전 숏의 인물 배치까지 상당 부분 이어받았다."
    ],
    "physics": "몸을 앞으로 살짝 숙인 채 한쪽 주먹을 쥔 자세이며, 테이블과 의자에 의해 물리적으로 잘 지탱되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 주먹을 쥔 역동적인 상체와 대각선 식탁 모서리 등 프레이밍 지시를 잘 따랐으나, 명시된 인물 외에 배경에 다른 사람들을 추가한 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "이전 샷의 인물들을 절대 가져오지 말라는 지시를 정면으로 위반하였으며, 두 주먹 대신 한쪽 주먹만 보이고 식탁 모서리의 대각선 프레이밍 지시도 놓쳤습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 좌측을 향하고 있으며, 크게 벌린 입과 표정으로 허공을 향해 분노를 터뜨리고 있습니다.",
        "built_space": "연구소 식당 내부로, 식탁 모서리가 화면 좌측 하단을 대각선으로 가로지르며 프롬프트의 배치 지시를 충실히 구현했습니다.",
        "entities": "현우(18세 남성, 헝클어진 머리)가 지시된 의상(흰색 옷)을 입고 있으며, 두 주먹을 꽉 쥔 채 탁자 위로 올리고 있습니다. 식탁에는 빵과 수프가 그대로 남아있습니다.",
        "hard_violations": [
         "지시된 인물(현우) 외에는 절대 사람을 추가하지 말라는 지시(never add a person the shot text does not show)를 위반하여 우측 배경에 흐릿한 인물들이 등장함."
        ],
        "physics": "보이지 않는 의자에 앉아 체중을 지탱하고 있으며, 양팔을 들어 주먹을 쥔 상체의 움직임이 자연스럽고 물리적으로 안정적입니다."
       },
       {
        "label": "B",
        "direction": "시선은 화면 우측 허공을 향하고 있으며, 분노하여 소리치는 표정입니다.",
        "built_space": "식당 내부이나, 식탁의 앞쪽 모서리가 대각선이 아닌 수평으로 반듯하게 배치되어 프레이밍 지시를 어겼습니다.",
        "entities": "현우가 지시된 의상을 입고 있으나, 주먹은 한쪽(화면 좌측)만 쥐고 있습니다. 식탁 위에는 빵과 식사가 놓여 있습니다.",
        "hard_violations": [
         "이전 샷(PREVIOUS SHOT STILL)에 등장했던 인물들을 절대 포함하지 말라는 지시를 정면 위반하여 좌측의 여성, 중앙의 흑인 남성 등 배경 인물들을 그대로 복사해 넣음.",
         "명시된 인물 외에 사람을 추가하지 말라는 지시 위반.",
         "두 주먹을 꽉 쥐고(두 주먹) 있어야 하나 한쪽 주먹만 보임.",
         "식탁 모서리가 화면 하단을 대각선으로 가로지르도록 프레이밍하라는 지시 위반(수평으로 배치됨)."
        ],
        "physics": "몸을 앞으로 살짝 숙인 채 한쪽 주먹을 쥔 자세이며, 테이블과 의자에 의해 물리적으로 잘 지탱되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 주먹을 쥔 역동적인 상체와 대각선 식탁 모서리 등 프레이밍 지시를 잘 따랐으나, 명시된 인물 외에 배경에 다른 사람들을 추가한 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "이전 샷의 인물들을 절대 가져오지 말라는 지시를 정면으로 위반하였으며, 두 주먹 대신 한쪽 주먹만 보이고 식탁 모서리의 대각선 프레이밍 지시도 놓쳤습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 좌측을 향하고 있으며, 크게 벌린 입과 표정으로 허공을 향해 분노를 터뜨리고 있습니다.",
        "built_space": "연구소 식당 내부로, 식탁 모서리가 화면 좌측 하단을 대각선으로 가로지르며 프롬프트의 배치 지시를 충실히 구현했습니다.",
        "entities": "현우(18세 남성, 헝클어진 머리)가 지시된 의상(흰색 옷)을 입고 있으며, 두 주먹을 꽉 쥔 채 탁자 위로 올리고 있습니다. 식탁에는 빵과 수프가 그대로 남아있습니다.",
        "hard_violations": [
         "지시된 인물(현우) 외에는 절대 사람을 추가하지 말라는 지시(never add a person the shot text does not show)를 위반하여 우측 배경에 흐릿한 인물들이 등장함."
        ],
        "physics": "보이지 않는 의자에 앉아 체중을 지탱하고 있으며, 양팔을 들어 주먹을 쥔 상체의 움직임이 자연스럽고 물리적으로 안정적입니다."
       },
       {
        "label": "B",
        "direction": "시선은 화면 우측 허공을 향하고 있으며, 분노하여 소리치는 표정입니다.",
        "built_space": "식당 내부이나, 식탁의 앞쪽 모서리가 대각선이 아닌 수평으로 반듯하게 배치되어 프레이밍 지시를 어겼습니다.",
        "entities": "현우가 지시된 의상을 입고 있으나, 주먹은 한쪽(화면 좌측)만 쥐고 있습니다. 식탁 위에는 빵과 식사가 놓여 있습니다.",
        "hard_violations": [
         "이전 샷(PREVIOUS SHOT STILL)에 등장했던 인물들을 절대 포함하지 말라는 지시를 정면 위반하여 좌측의 여성, 중앙의 흑인 남성 등 배경 인물들을 그대로 복사해 넣음.",
         "명시된 인물 외에 사람을 추가하지 말라는 지시 위반.",
         "두 주먹을 꽉 쥐고(두 주먹) 있어야 하나 한쪽 주먹만 보임.",
         "식탁 모서리가 화면 하단을 대각선으로 가로지르도록 프레이밍하라는 지시 위반(수평으로 배치됨)."
        ],
        "physics": "몸을 앞으로 살짝 숙인 채 한쪽 주먹을 쥔 자세이며, 테이블과 의자에 의해 물리적으로 잘 지탱되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "분노한 상체와 식당의 연속성은 구현했지만, 금지된 배경 인물을 다수 추가했고 두 주먹 중 하나만 보여 핵심 행동 전달도 부족하다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "두 주먹을 꽉 쥐고 입을 크게 벌린 역동적인 상체를 미디엄 숏으로 더 충실히 구현했으나, 금지된 배경 인물 때문에 최종 사용에는 부적합하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸을 앞으로 기울이고 화면 오른쪽의 프레임 밖을 바라보며 소리친다. 시선의 구체적인 대상은 보이지 않으며 지문도 대상을 지정하지 않는다. 화면 왼쪽 주먹은 식탁을 향해 내려와 있고 다른 손은 보이지 않는다.",
        "built_space": "왼쪽의 연속된 큰 창과 흰 기둥, 오른쪽의 긴 금속 후드 한 줄과 배식대 한 줄, 배식대 아래 열린 수납부가 이전 장소와 잘 대응한다. 전경에는 식탁 하나와 쟁반 하나가 있고, 검은 테두리 의자들이 현우 주변과 뒤쪽 식탁에 배치되어 있다. 현우는 전경 식탁 바로 뒤에 있다. 식탁의 눈에 띄는 경계가 거의 수평이어서 요구한 하단 대각선 구도는 약하다.",
        "entities": "주인공은 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리, 마른 체격, 흰색 단추 상의가 현우 참조와 대체로 일치한다. 국적은 외양만으로 확인할 수 없다. 입을 크게 벌리고 얼굴을 찡그렸으며 한쪽 주먹은 명확히 쥐고 있다. 빵과 잘린 빵이 담긴 접시 하나, 음식이 남은 그릇 하나가 식탁 위에 있다. 그러나 왼쪽 식사객들과 오른쪽 배식대 주변 인물 등 여러 추가 인물이 보인다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "현우만 허용된 장면에 다수의 배경 식사객과 배식대 주변 인물을 추가했으며, 이전 숏의 인물 배치까지 상당 부분 이어받았다."
        ],
        "physics": "화면 왼쪽 주먹의 아래쪽이 식탁에 닿아 있고 팔은 어깨와 자연스럽게 연결된다. 몸통을 앞으로 숙인 자세는 가능한 동작이며 하체의 지지는 식탁에 가려 확인되지 않는다. 접시와 음식 그릇은 쟁반에, 쟁반은 식탁에 놓여 있다. 명백히 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽의 프레임 밖을 바라보며 입을 크게 벌린다. 특정 상대는 보이지 않지만 지정된 시선 대상도 없다. 두 주먹은 몸통 앞 양쪽에서 위로 올라와 있으며 특정 인물을 때리거나 물체를 겨누는 자세가 아니라 분노로 움켜쥔 동작으로 읽힌다.",
        "built_space": "왼쪽의 큰 창들과 흰 기둥, 뒤로 이어지는 식탁 및 검은 테두리 의자가 이전 식당의 재료와 공간감을 유지한다. 오른쪽 배식대는 선택된 카메라 각도 밖에 있어 확인할 수 없다. 전경에는 현우 앞 식탁 하나와 왼쪽 인접 식탁 일부가 있고, 하단 식탁 경계가 비스듬히 지나간다. 현우의 상체와 두 주먹이 중심을 차지하고 음식은 하단 주변부에 놓여 요구 구도에 더 가깝다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 검은 헝클어진 머리, 얼굴 윤곽, 마른 체격과 흰색 단추 상의가 참조에 대체로 부합한다. 두 주먹 모두 꽉 쥐었고 찡그린 얼굴과 크게 벌린 입이 명확하다. 접시 하나에는 빵과 다른 음식이 남아 있고, 옆의 그릇 하나에도 음식이 남아 있다. 오른쪽 먼 배경에는 서 있는 인물, 앉은 인물과 가장자리의 잘린 인물 등 추가 인물이 보인다. 벽의 게시물은 흐려서 글자를 읽을 수 없다.",
        "hard_violations": [
         "현우 외 인물이 금지되어 있는데도 오른쪽 먼 배경에 여러 사람과 가장자리의 부분 신체를 추가했다."
        ],
        "physics": "들어 올린 두 주먹은 각각 손목과 굽힌 팔을 통해 몸통에 연결되어 있으며 팔 근육으로 유지 가능한 자세다. 상체의 긴장과 전방 기울임도 자연스럽다. 하체와 좌석 접촉은 식탁에 가려져 있으므로 지지 불능으로 볼 근거가 없다. 빵과 음식은 접시 및 그릇에 담겨 식탁 위에 놓여 있고 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "분노한 상체와 식당의 연속성은 구현했지만, 금지된 배경 인물을 다수 추가했고 두 주먹 중 하나만 보여 핵심 행동 전달도 부족하다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "두 주먹을 꽉 쥐고 입을 크게 벌린 역동적인 상체를 미디엄 숏으로 더 충실히 구현했으나, 금지된 배경 인물 때문에 최종 사용에는 부적합하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸을 앞으로 기울이고 화면 오른쪽의 프레임 밖을 바라보며 소리친다. 시선의 구체적인 대상은 보이지 않으며 지문도 대상을 지정하지 않는다. 화면 왼쪽 주먹은 식탁을 향해 내려와 있고 다른 손은 보이지 않는다.",
        "built_space": "왼쪽의 연속된 큰 창과 흰 기둥, 오른쪽의 긴 금속 후드 한 줄과 배식대 한 줄, 배식대 아래 열린 수납부가 이전 장소와 잘 대응한다. 전경에는 식탁 하나와 쟁반 하나가 있고, 검은 테두리 의자들이 현우 주변과 뒤쪽 식탁에 배치되어 있다. 현우는 전경 식탁 바로 뒤에 있다. 식탁의 눈에 띄는 경계가 거의 수평이어서 요구한 하단 대각선 구도는 약하다.",
        "entities": "주인공은 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리, 마른 체격, 흰색 단추 상의가 현우 참조와 대체로 일치한다. 국적은 외양만으로 확인할 수 없다. 입을 크게 벌리고 얼굴을 찡그렸으며 한쪽 주먹은 명확히 쥐고 있다. 빵과 잘린 빵이 담긴 접시 하나, 음식이 남은 그릇 하나가 식탁 위에 있다. 그러나 왼쪽 식사객들과 오른쪽 배식대 주변 인물 등 여러 추가 인물이 보인다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "현우만 허용된 장면에 다수의 배경 식사객과 배식대 주변 인물을 추가했으며, 이전 숏의 인물 배치까지 상당 부분 이어받았다."
        ],
        "physics": "화면 왼쪽 주먹의 아래쪽이 식탁에 닿아 있고 팔은 어깨와 자연스럽게 연결된다. 몸통을 앞으로 숙인 자세는 가능한 동작이며 하체의 지지는 식탁에 가려 확인되지 않는다. 접시와 음식 그릇은 쟁반에, 쟁반은 식탁에 놓여 있다. 명백히 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽의 프레임 밖을 바라보며 입을 크게 벌린다. 특정 상대는 보이지 않지만 지정된 시선 대상도 없다. 두 주먹은 몸통 앞 양쪽에서 위로 올라와 있으며 특정 인물을 때리거나 물체를 겨누는 자세가 아니라 분노로 움켜쥔 동작으로 읽힌다.",
        "built_space": "왼쪽의 큰 창들과 흰 기둥, 뒤로 이어지는 식탁 및 검은 테두리 의자가 이전 식당의 재료와 공간감을 유지한다. 오른쪽 배식대는 선택된 카메라 각도 밖에 있어 확인할 수 없다. 전경에는 현우 앞 식탁 하나와 왼쪽 인접 식탁 일부가 있고, 하단 식탁 경계가 비스듬히 지나간다. 현우의 상체와 두 주먹이 중심을 차지하고 음식은 하단 주변부에 놓여 요구 구도에 더 가깝다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 검은 헝클어진 머리, 얼굴 윤곽, 마른 체격과 흰색 단추 상의가 참조에 대체로 부합한다. 두 주먹 모두 꽉 쥐었고 찡그린 얼굴과 크게 벌린 입이 명확하다. 접시 하나에는 빵과 다른 음식이 남아 있고, 옆의 그릇 하나에도 음식이 남아 있다. 오른쪽 먼 배경에는 서 있는 인물, 앉은 인물과 가장자리의 잘린 인물 등 추가 인물이 보인다. 벽의 게시물은 흐려서 글자를 읽을 수 없다.",
        "hard_violations": [
         "현우 외 인물이 금지되어 있는데도 오른쪽 먼 배경에 여러 사람과 가장자리의 부분 신체를 추가했다."
        ],
        "physics": "들어 올린 두 주먹은 각각 손목과 굽힌 팔을 통해 몸통에 연결되어 있으며 팔 근육으로 유지 가능한 자세다. 상체의 긴장과 전방 기울임도 자연스럽다. 하체와 좌석 접촉은 식탁에 가려져 있으므로 지지 불능으로 볼 근거가 없다. 빵과 음식은 접시 및 그릇에 담겨 식탁 위에 놓여 있고 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.833
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.583
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시된 인물(현우) 외에는 절대 사람을 추가하지 말라는 지시(never add a person the shot text does not show)를 위반하여 우측 배경에 흐릿한 인물들이 등장함.",
     "[gpt-high] 현우 외 인물이 금지되어 있는데도 오른쪽 먼 배경에 여러 사람과 가장자리의 부분 신체를 추가했다."
    ],
    "B": [
     "[gemini-pro] 이전 샷(PREVIOUS SHOT STILL)에 등장했던 인물들을 절대 포함하지 말라는 지시를 정면 위반하여 좌측의 여성, 중앙의 흑인 남성 등 배경 인물들을 그대로 복사해 넣음.",
     "[gemini-pro] 명시된 인물 외에 사람을 추가하지 말라는 지시 위반.",
     "[gemini-pro] 두 주먹을 꽉 쥐고(두 주먹) 있어야 하나 한쪽 주먹만 보임.",
     "[gemini-pro] 식탁 모서리가 화면 하단을 대각선으로 가로지르도록 프레이밍하라는 지시 위반(수평으로 배치됨).",
     "[gpt-high] 현우만 허용된 장면에 다수의 배경 식사객과 배식대 주변 인물을 추가했으며, 이전 숏의 인물 배치까지 상당 부분 이어받았다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 583
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "두 주먹을 쥔 역동적인 상체와 대각선 식탁 모서리 등 프레이밍 지시를 잘 따랐으나, 명시된 인물 외에 배경에 다른 사람들을 추가한 점이 감점 요인입니다.  ★위반: [gemini-pro] 지시된 인물(현우) 외에는 절대 사람을 추가하지 말라는 지시(never add a person the shot text does not show)를 위반하여 우측 배경에 흐릿한 인물들이 등장함. / [gpt-high] 현우 외 인물이 금지되어 있는데도 오른쪽 먼 배경에 여러 사람과 가장자리의 부분 신체를 추가했다."
   },
   {
    "label": "B",
    "score": 583,
    "verdict_ko": "이전 샷의 인물들을 절대 가져오지 말라는 지시를 정면으로 위반하였으며, 두 주먹 대신 한쪽 주먹만 보이고 식탁 모서리의 대각선 프레이밍 지시도 놓쳤습니다.  ★위반: [gemini-pro] 이전 샷(PREVIOUS SHOT STILL)에 등장했던 인물들을 절대 포함하지 말라는 지시를 정면 위반하여 좌측의 여성, 중앙의 흑인 남성 등 배경 인물들을 그대로 복사해 넣음. / [gemini-pro] 명시된 인물 외에 사람을 추가하지 말라는 지시 위반. / [gemini-pro] 두 주먹을 꽉 쥐고(두 주먹) 있어야 하나 한쪽 주먹만 보임. / [gemini-pro] 식탁 모서리가 화면 하단을 대각선으로 가로지르도록 프레이밍하라는 지시 위반(수평으로 배치됨). / [gpt-high] 현우만 허용된 장면에 다수의 배경 식사객과 배식대 주변 인물을 추가했으며, 이전 숏의 인물 배치까지 상당 부분 이어받았다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22_sel.png",
    "asset_id": "82a44334-e8a8-434b-994f-4761905fe13b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-0d8b-7189-8bc5-cabb442455d1",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S85sh22"
  }
 },
 "S85sh23::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:09:16.785945+00:00",
  "fingerprint": "7a0836f1f803ae8a4a8952eae11ab0b8e615c55c4b01da0688b83926c745c5f7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S85sh23_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S85sh23_sel.png",
  "source_sha256": "20189314fcd4be5b8c434148d063cb166d178b872c75c2d5f7041ba0ba31a0bd",
  "file": "S85sh23_cine.png",
  "staged_sha256": "f54169aee8cf40041f665b08d6058f1744f27c0991db0a400d3c9f5ff9d18c3a",
  "latency_ms": 11170
 },
 "S86sh2::signage": {
  "fp": "6611a6f01424ad0c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S86sh2": {
  "input_fingerprint": "b8ba19bade77099f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영을 향해 핏발 선 눈으로 삿대질한 채 고함치는 현우의 분노한 상체.\n\nLOCATION (lock): In the open conversation area inside the research director's office, under nighttime office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office doorway (Open following 현우's entrance) — Seen obliquely behind him on the established entrance side; used as Peripheral spatial anchor connecting his entrance to the confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate ambient illumination with controlled facial contrast, preserving the bloodshot eyes without inventing a colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 현우: He has entered the director's office and remains standing in an agitated confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영을 향해 핏발 선 눈으로 삿대질한 채 고함치는 현우의 분노한 상체.\n\nLOCATION (lock): In the open conversation area inside the research director's office, under nighttime office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office doorway (Open following 현우's entrance) — Seen obliquely behind him on the established entrance side; used as Peripheral spatial anchor connecting his entrance to the confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate ambient illumination with controlled facial contrast, preserving the bloodshot eyes without inventing a colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 현우: He has entered the director's office and remains standing in an agitated confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영을 향해 핏발 선 눈으로 삿대질한 채 고함치는 현우의 분노한 상체.\n\nLOCATION (lock): In the open conversation area inside the research director's office, under nighttime office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office doorway (Open following 현우's entrance) — Seen obliquely behind him on the established entrance side; used as Peripheral spatial anchor connecting his entrance to the confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate ambient illumination with controlled facial contrast, preserving the bloodshot eyes without inventing a colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 현우: He has entered the director's office and remains standing in an agitated confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 삿대질하는 왼손이 전경 우측에 위치한 여성(지소영)을 향해 정확히 겨냥됨.",
    "built_space": "소장실 내부. 좌측에 열린 나무 문이 있고, 배경에 피아노와 장식장이 보이며 곡선형 천장 조명 등 기존 레퍼런스의 구조와 일치함.",
    "entities": "현우는 캐릭터 레퍼런스의 얼굴, 헤어스타일, 흰색 유니폼과 정확히 일치하며 핏발 선 눈이 잘 표현됨. 전경에 지소영의 뒷모습이 보임.",
    "hard_violations": [
     "[gpt-high] 현우 외의 인물을 허용하지 않은 프레임에 여성의 머리와 상체를 추가했으며, 이전 숏의 여성 외모와 의상을 다시 가져온 것으로 보인다."
    ],
    "physics": "자연스럽게 서서 체중을 지탱하고 있으며, 손의 방향과 구조가 삿대질하는 동작에 맞게 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "현우의 시선과 삿대질하는 손이 전경 우측에 위치한 여성을 향하고 있음.",
    "built_space": "사무실 내부. 우측에 열린 문과 배경에 피아노가 있으나, 천장에 레퍼런스에 없는 밝은 사각형 패널 조명들이 추가됨.",
    "entities": "현우의 얼굴과 핏발 선 눈은 표현되었으나, 지정된 흰색 유니폼 대신 어두운 재킷과 티셔츠를 입고 있어 복장 설정과 어긋남.",
    "hard_violations": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (왼팔 끝에 엄지손가락 방향이 반대인 오른손이 달려 있음)",
     "[gemini-pro] 공간 레퍼런스 위반 (원래 공간에 없는 밝은 사각형 천장 조명을 임의로 추가함)",
     "[gpt-high] 현우만 등장하도록 제한된 프레임에 별도의 여성 머리와 상체를 추가했으며, 이전 숏 인물을 가져오지 말라는 조건에도 어긋난다."
    ],
    "physics": "서 있는 자세 자체는 유지되나, 왼팔에서 이어지는 손의 형태가 오른손 구조를 하고 있어 해부학적 물리법칙을 위반함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "핏발 선 눈과 분노한 표정을 훌륭하게 연출했으며, 지정된 흰색 유니폼 복장과 레퍼런스의 공간(곡선형 조명, 피아노 배치 등)을 완벽하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지정된 복장을 무시하고 어두운 재킷을 입혔으며 레퍼런스에 없는 사각형 천장 조명을 창작했고, 결정적으로 왼팔에 오른손이 달려있는 치명적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 삿대질하는 왼손이 전경 우측에 위치한 여성(지소영)을 향해 정확히 겨냥됨.",
        "built_space": "소장실 내부. 좌측에 열린 나무 문이 있고, 배경에 피아노와 장식장이 보이며 곡선형 천장 조명 등 기존 레퍼런스의 구조와 일치함.",
        "entities": "현우는 캐릭터 레퍼런스의 얼굴, 헤어스타일, 흰색 유니폼과 정확히 일치하며 핏발 선 눈이 잘 표현됨. 전경에 지소영의 뒷모습이 보임.",
        "hard_violations": [],
        "physics": "자연스럽게 서서 체중을 지탱하고 있으며, 손의 방향과 구조가 삿대질하는 동작에 맞게 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "현우의 시선과 삿대질하는 손이 전경 우측에 위치한 여성을 향하고 있음.",
        "built_space": "사무실 내부. 우측에 열린 문과 배경에 피아노가 있으나, 천장에 레퍼런스에 없는 밝은 사각형 패널 조명들이 추가됨.",
        "entities": "현우의 얼굴과 핏발 선 눈은 표현되었으나, 지정된 흰색 유니폼 대신 어두운 재킷과 티셔츠를 입고 있어 복장 설정과 어긋남.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (왼팔 끝에 엄지손가락 방향이 반대인 오른손이 달려 있음)",
         "공간 레퍼런스 위반 (원래 공간에 없는 밝은 사각형 천장 조명을 임의로 추가함)"
        ],
        "physics": "서 있는 자세 자체는 유지되나, 왼팔에서 이어지는 손의 형태가 오른손 구조를 하고 있어 해부학적 물리법칙을 위반함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "핏발 선 눈과 분노한 표정을 훌륭하게 연출했으며, 지정된 흰색 유니폼 복장과 레퍼런스의 공간(곡선형 조명, 피아노 배치 등)을 완벽하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지정된 복장을 무시하고 어두운 재킷을 입혔으며 레퍼런스에 없는 사각형 천장 조명을 창작했고, 결정적으로 왼팔에 오른손이 달려있는 치명적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 삿대질하는 왼손이 전경 우측에 위치한 여성(지소영)을 향해 정확히 겨냥됨.",
        "built_space": "소장실 내부. 좌측에 열린 나무 문이 있고, 배경에 피아노와 장식장이 보이며 곡선형 천장 조명 등 기존 레퍼런스의 구조와 일치함.",
        "entities": "현우는 캐릭터 레퍼런스의 얼굴, 헤어스타일, 흰색 유니폼과 정확히 일치하며 핏발 선 눈이 잘 표현됨. 전경에 지소영의 뒷모습이 보임.",
        "hard_violations": [],
        "physics": "자연스럽게 서서 체중을 지탱하고 있으며, 손의 방향과 구조가 삿대질하는 동작에 맞게 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "현우의 시선과 삿대질하는 손이 전경 우측에 위치한 여성을 향하고 있음.",
        "built_space": "사무실 내부. 우측에 열린 문과 배경에 피아노가 있으나, 천장에 레퍼런스에 없는 밝은 사각형 패널 조명들이 추가됨.",
        "entities": "현우의 얼굴과 핏발 선 눈은 표현되었으나, 지정된 흰색 유니폼 대신 어두운 재킷과 티셔츠를 입고 있어 복장 설정과 어긋남.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (왼팔 끝에 엄지손가락 방향이 반대인 오른손이 달려 있음)",
         "공간 레퍼런스 위반 (원래 공간에 없는 밝은 사각형 천장 조명을 임의로 추가함)"
        ],
        "physics": "서 있는 자세 자체는 유지되나, 왼팔에서 이어지는 손의 형태가 오른손 구조를 하고 있어 해부학적 물리법칙을 위반함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "분노한 상체와 삿대질은 구현했지만, 금지된 여성 인물을 전경에 추가했고 현우의 흰색 의상을 짙은 재킷으로 바꿨다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 흰옷·얼굴·충혈된 눈과 고함치는 동작은 더 충실하지만, 금지된 여성 인물 추가와 전경으로 옮겨진 출입문 때문에 요구한 단독 상체 숏에는 실패했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈과 뻗은 검지가 모두 화면 오른쪽 전경 여성의 얼굴을 향한다. 삿대질과 시선의 목표는 일치하지만, 지소영을 화면 밖 대상으로 처리하지 않고 실제 여성 인물을 넣었다. 여성은 현우 쪽으로 얼굴을 돌리고 있다.",
        "built_space": "현우를 허리 부근에서 자른 미디엄 숏이다. 오른쪽 뒤에 비스듬히 열린 목재 문 하나와 손잡이 하나가 보인다. 왼쪽에는 유리 칸막이, 뒤에는 검은 피아노 한 대와 의자 하나, 낮은 수납장과 조명이 보인다. 출입문의 주변부 배치는 요구에 가깝지만, 현우가 유리 경계와 문 바로 앞에 있어 사무실 안 개방형 대화 공간에 들어선 상태는 덜 명확하다. 천장의 밝은 사각 조명들은 참조보다 두드러진다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 검은 헝클어진 머리와 앳된 동아시아계 남성 외모로 참조의 기본 인상에 가깝다. 한국계 미국인이라는 국적·배경 자체는 화면으로 확인할 수 없다. 입을 벌려 고함치고 눈 주변이 붉지만, 참조에 없는 뺨의 상처 같은 자국이 보인다. 참조의 흰 단추 상의 대신 짙은 회녹색 재킷과 안옷을 입었다. 검은 업라이트 피아노는 유지되어 있다. 오른쪽에는 이전 숏의 여성과 유사한 짧은 검은 머리와 흰옷의 인물이 추가되었다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "현우만 등장하도록 제한된 프레임에 별도의 여성 머리와 상체를 추가했으며, 이전 숏 인물을 가져오지 말라는 조건에도 어긋난다."
        ],
        "physics": "현우의 상체는 서 있는 자세로 이어지고, 삿대질하는 팔은 어깨·팔꿈치·손목으로 자연스럽게 연결된다. 발은 프레임 밖이므로 접지점은 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 문은 문틀에 연결되어 있고 피아노와 의자는 바닥 위에 놓여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 전경 여성의 얼굴을 바라보며 검지를 같은 방향으로 뻗고 있다. 고함치는 입과 몸의 전진 방향도 그 여성을 향한다. 목표 방향은 일관되지만, 화면 밖이어야 할 상대를 여성의 머리와 어깨로 직접 보여준다.",
        "built_space": "현우의 허리 부근까지 담은 미디엄 숏이며, 왼쪽 전경에 열린 목재 문 하나와 경첩·손잡이가 보인다. 뒤에는 유리 벽, 검은 업라이트 피아노 한 대, 왼쪽 낮은 수납장 하나와 작은 조명 하나가 있다. 오른쪽 아래에는 낮은 목재 탁자 일부가 보인다. 밤의 유리 공간과 회색 벽, 따뜻한 간접조명은 참조와 대체로 이어진다. 그러나 출입문이 현우 뒤의 주변부 표지가 아니라 카메라 바로 옆 전경에 있어 요구한 공간 관계와 다르다. 중복된 피아노나 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 참조와 유사한 앳된 동아시아계 남성 얼굴, 검은 머리, 흰색 둥근 목 단추 상의를 갖췄다. 머리는 참조보다 조금 정돈되어 있지만 A보다 의상과 인물 인상이 충실하다. 눈은 정상적인 홍채와 동공을 유지하면서 붉게 충혈되어 있고, 찌푸린 미간과 크게 열린 입이 분노와 고함을 보여준다. 국적은 외형만으로 확인할 수 없다. 뒤의 오래된 검은 피아노는 유지되었다. 오른쪽에는 허용되지 않은 짧은 검은 머리의 흰옷 여성 인물이 보인다. 판독 가능한 글자는 없다.",
        "hard_violations": [
         "현우 외의 인물을 허용하지 않은 프레임에 여성의 머리와 상체를 추가했으며, 이전 숏의 여성 외모와 의상을 다시 가져온 것으로 보인다."
        ],
        "physics": "현우는 허리를 약간 앞으로 기울인 채 팔꿈치를 굽혀 삿대질한다. 팔과 손의 연결 및 상체 기울기는 서서 격앙되게 말하는 동작으로 가능하다. 발은 잘려 있어 직접적인 접지는 확인되지 않지만 부유를 나타내는 모습은 없다. 열린 문은 보이는 경첩으로 지지되고, 수납장과 피아노는 바닥에 놓여 있다. 탁자 위 작은 직사각형 물체도 상판에 얹혀 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "분노한 상체와 삿대질은 구현했지만, 금지된 여성 인물을 전경에 추가했고 현우의 흰색 의상을 짙은 재킷으로 바꿨다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우의 흰옷·얼굴·충혈된 눈과 고함치는 동작은 더 충실하지만, 금지된 여성 인물 추가와 전경으로 옮겨진 출입문 때문에 요구한 단독 상체 숏에는 실패했다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈과 뻗은 검지가 모두 화면 오른쪽 전경 여성의 얼굴을 향한다. 삿대질과 시선의 목표는 일치하지만, 지소영을 화면 밖 대상으로 처리하지 않고 실제 여성 인물을 넣었다. 여성은 현우 쪽으로 얼굴을 돌리고 있다.",
        "built_space": "현우를 허리 부근에서 자른 미디엄 숏이다. 오른쪽 뒤에 비스듬히 열린 목재 문 하나와 손잡이 하나가 보인다. 왼쪽에는 유리 칸막이, 뒤에는 검은 피아노 한 대와 의자 하나, 낮은 수납장과 조명이 보인다. 출입문의 주변부 배치는 요구에 가깝지만, 현우가 유리 경계와 문 바로 앞에 있어 사무실 안 개방형 대화 공간에 들어선 상태는 덜 명확하다. 천장의 밝은 사각 조명들은 참조보다 두드러진다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 검은 헝클어진 머리와 앳된 동아시아계 남성 외모로 참조의 기본 인상에 가깝다. 한국계 미국인이라는 국적·배경 자체는 화면으로 확인할 수 없다. 입을 벌려 고함치고 눈 주변이 붉지만, 참조에 없는 뺨의 상처 같은 자국이 보인다. 참조의 흰 단추 상의 대신 짙은 회녹색 재킷과 안옷을 입었다. 검은 업라이트 피아노는 유지되어 있다. 오른쪽에는 이전 숏의 여성과 유사한 짧은 검은 머리와 흰옷의 인물이 추가되었다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "현우만 등장하도록 제한된 프레임에 별도의 여성 머리와 상체를 추가했으며, 이전 숏 인물을 가져오지 말라는 조건에도 어긋난다."
        ],
        "physics": "현우의 상체는 서 있는 자세로 이어지고, 삿대질하는 팔은 어깨·팔꿈치·손목으로 자연스럽게 연결된다. 발은 프레임 밖이므로 접지점은 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 문은 문틀에 연결되어 있고 피아노와 의자는 바닥 위에 놓여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 전경 여성의 얼굴을 바라보며 검지를 같은 방향으로 뻗고 있다. 고함치는 입과 몸의 전진 방향도 그 여성을 향한다. 목표 방향은 일관되지만, 화면 밖이어야 할 상대를 여성의 머리와 어깨로 직접 보여준다.",
        "built_space": "현우의 허리 부근까지 담은 미디엄 숏이며, 왼쪽 전경에 열린 목재 문 하나와 경첩·손잡이가 보인다. 뒤에는 유리 벽, 검은 업라이트 피아노 한 대, 왼쪽 낮은 수납장 하나와 작은 조명 하나가 있다. 오른쪽 아래에는 낮은 목재 탁자 일부가 보인다. 밤의 유리 공간과 회색 벽, 따뜻한 간접조명은 참조와 대체로 이어진다. 그러나 출입문이 현우 뒤의 주변부 표지가 아니라 카메라 바로 옆 전경에 있어 요구한 공간 관계와 다르다. 중복된 피아노나 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 참조와 유사한 앳된 동아시아계 남성 얼굴, 검은 머리, 흰색 둥근 목 단추 상의를 갖췄다. 머리는 참조보다 조금 정돈되어 있지만 A보다 의상과 인물 인상이 충실하다. 눈은 정상적인 홍채와 동공을 유지하면서 붉게 충혈되어 있고, 찌푸린 미간과 크게 열린 입이 분노와 고함을 보여준다. 국적은 외형만으로 확인할 수 없다. 뒤의 오래된 검은 피아노는 유지되었다. 오른쪽에는 허용되지 않은 짧은 검은 머리의 흰옷 여성 인물이 보인다. 판독 가능한 글자는 없다.",
        "hard_violations": [
         "현우 외의 인물을 허용하지 않은 프레임에 여성의 머리와 상체를 추가했으며, 이전 숏의 여성 외모와 의상을 다시 가져온 것으로 보인다."
        ],
        "physics": "현우는 허리를 약간 앞으로 기울인 채 팔꿈치를 굽혀 삿대질한다. 팔과 손의 연결 및 상체 기울기는 서서 격앙되게 말하는 동작으로 가능하다. 발은 잘려 있어 직접적인 접지는 확인되지 않지만 부유를 나타내는 모습은 없다. 열린 문은 보이는 경첩으로 지지되고, 수납장과 피아노는 바닥에 놓여 있다. 탁자 위 작은 직사각형 물체도 상판에 얹혀 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (왼팔 끝에 엄지손가락 방향이 반대인 오른손이 달려 있음)",
     "[gemini-pro] 공간 레퍼런스 위반 (원래 공간에 없는 밝은 사각형 천장 조명을 임의로 추가함)",
     "[gpt-high] 현우만 등장하도록 제한된 프레임에 별도의 여성 머리와 상체를 추가했으며, 이전 숏 인물을 가져오지 말라는 조건에도 어긋난다."
    ],
    "A": [
     "[gpt-high] 현우 외의 인물을 허용하지 않은 프레임에 여성의 머리와 상체를 추가했으며, 이전 숏의 여성 외모와 의상을 다시 가져온 것으로 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "핏발 선 눈과 분노한 표정을 훌륭하게 연출했으며, 지정된 흰색 유니폼 복장과 레퍼런스의 공간(곡선형 조명, 피아노 배치 등)을 완벽하게 재현했습니다.  ★위반: [gpt-high] 현우 외의 인물을 허용하지 않은 프레임에 여성의 머리와 상체를 추가했으며, 이전 숏의 여성 외모와 의상을 다시 가져온 것으로 보인다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "지정된 복장을 무시하고 어두운 재킷을 입혔으며 레퍼런스에 없는 사각형 천장 조명을 창작했고, 결정적으로 왼팔에 오른손이 달려있는 치명적 오류가 있습니다.  ★위반: [gemini-pro] 해부학적으로 불가능한 신체 구조 (왼팔 끝에 엄지손가락 방향이 반대인 오른손이 달려 있음) / [gemini-pro] 공간 레퍼런스 위반 (원래 공간에 없는 밝은 사각형 천장 조명을 임의로 추가함) / [gpt-high] 현우만 등장하도록 제한된 프레임에 별도의 여성 머리와 상체를 추가했으며, 이전 숏 인물을 가져오지 말라는 조건에도 어긋난다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh14_sel.png",
    "asset_id": "affa89dc-6c1f-42dc-ae01-e1736910fb19",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-0f2e-7392-aed2-e284c05e2c94",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S84sh14"
  }
 },
 "S86sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:10:21.386430+00:00",
  "fingerprint": "f0bd7368856904a66aa263a646b9c142ad58dcd571b8c5ea4a64d7ff8128d5e3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S86sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S86sh2_sel.png",
  "source_sha256": "c9ee0cc0bbdcccd8d7b53b040eaaa01da8b75c4d8a76f762ef322e7f553014f0",
  "file": "S86sh2_cine.png",
  "staged_sha256": "c095438e162a64f17cb97adb4ed931e73859df028531e3de589dcec59e8611e9",
  "latency_ms": 9977
 },
 "S86sh5::signage": {
  "fp": "740ce7685831109a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S86sh5": {
  "input_fingerprint": "d6c4d0da6043aa42",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양손을 펴 보인 채 애써 차분하게 달래는 기색의 지소영 굳은 상체.\n\nLOCATION (lock): Inside the director's office adjoining the research center, facing the visitor under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained ambient illumination with sufficient tonal separation to read her open hands and the tension in her face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 지소영: She remains in her neat research coat, trying to maintain composure during the confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양손을 펴 보인 채 애써 차분하게 달래는 기색의 지소영 굳은 상체.\n\nLOCATION (lock): Inside the director's office adjoining the research center, facing the visitor under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained ambient illumination with sufficient tonal separation to read her open hands and the tension in her face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 지소영: She remains in her neat research coat, trying to maintain composure during the confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양손을 펴 보인 채 애써 차분하게 달래는 기색의 지소영 굳은 상체.\n\nLOCATION (lock): Inside the director's office adjoining the research center, facing the visitor under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained ambient illumination with sufficient tonal separation to read her open hands and the tension in her face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 지소영: She remains in her neat research coat, trying to maintain composure during the confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라 앞쪽(보이지 않는 현우)을 향해 시선과 양손이 향함.",
    "built_space": "지정된 감독실 내부. 이전 샷의 문틀, 유리벽, 피아노, 책상 등이 올바른 위치에 있음.",
    "entities": "지소영 단독 등장. 레퍼런스의 얼굴, 단정한 흑발, 연구복 코트를 정확히 반영함.",
    "hard_violations": [],
    "physics": "바닥에 안정적으로 서 있으며, 들어올린 양팔이 손을 지지함."
   },
   {
    "label": "B",
    "direction": "화면 우측에 서 있는 남성을 향해 시선과 손이 향함.",
    "built_space": "감독실 내부. 수납장과 피아노 등 배경 요소가 적절히 배치됨.",
    "entities": "지소영이 지정된 연구복이 아닌 일반 재킷을 입음. 프롬프트에 없는 남성이 우측에 있음.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에서 엄격히 금지한 추가 인물(이전 샷의 남성) 포함",
     "[gpt-high] 이번 화면에는 지소영만 등장할 수 있는데, 오른쪽 전경에 남성의 머리와 상체를 추가하여 이전 숏 인물을 가져오지 말라는 지시를 위반했다."
    ],
    "physics": "바닥에 서서 양팔을 들어 자연스럽게 손을 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 단일 인물 제약과 의상 레퍼런스를 완벽히 준수하며, 요구된 미디엄 샷과 양손을 편 자세를 정확하게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "등장인물 제한 지침을 어기고 화면에 남성을 추가했으며, 지정된 연구복 의상도 일치하지 않아 치명적 위반임."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라 앞쪽(보이지 않는 현우)을 향해 시선과 양손이 향함.",
        "built_space": "지정된 감독실 내부. 이전 샷의 문틀, 유리벽, 피아노, 책상 등이 올바른 위치에 있음.",
        "entities": "지소영 단독 등장. 레퍼런스의 얼굴, 단정한 흑발, 연구복 코트를 정확히 반영함.",
        "hard_violations": [],
        "physics": "바닥에 안정적으로 서 있으며, 들어올린 양팔이 손을 지지함."
       },
       {
        "label": "B",
        "direction": "화면 우측에 서 있는 남성을 향해 시선과 손이 향함.",
        "built_space": "감독실 내부. 수납장과 피아노 등 배경 요소가 적절히 배치됨.",
        "entities": "지소영이 지정된 연구복이 아닌 일반 재킷을 입음. 프롬프트에 없는 남성이 우측에 있음.",
        "hard_violations": [
         "프롬프트에서 엄격히 금지한 추가 인물(이전 샷의 남성) 포함"
        ],
        "physics": "바닥에 서서 양팔을 들어 자연스럽게 손을 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 단일 인물 제약과 의상 레퍼런스를 완벽히 준수하며, 요구된 미디엄 샷과 양손을 편 자세를 정확하게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "등장인물 제한 지침을 어기고 화면에 남성을 추가했으며, 지정된 연구복 의상도 일치하지 않아 치명적 위반임."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라 앞쪽(보이지 않는 현우)을 향해 시선과 양손이 향함.",
        "built_space": "지정된 감독실 내부. 이전 샷의 문틀, 유리벽, 피아노, 책상 등이 올바른 위치에 있음.",
        "entities": "지소영 단독 등장. 레퍼런스의 얼굴, 단정한 흑발, 연구복 코트를 정확히 반영함.",
        "hard_violations": [],
        "physics": "바닥에 안정적으로 서 있으며, 들어올린 양팔이 손을 지지함."
       },
       {
        "label": "B",
        "direction": "화면 우측에 서 있는 남성을 향해 시선과 손이 향함.",
        "built_space": "감독실 내부. 수납장과 피아노 등 배경 요소가 적절히 배치됨.",
        "entities": "지소영이 지정된 연구복이 아닌 일반 재킷을 입음. 프롬프트에 없는 남성이 우측에 있음.",
        "hard_violations": [
         "프롬프트에서 엄격히 금지한 추가 인물(이전 샷의 남성) 포함"
        ],
        "physics": "바닥에 서서 양팔을 들어 자연스럽게 손을 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우를 향한 달래는 동작은 보이지만, 등장 금지된 남성을 오른쪽 전경에 추가했고 연구복도 인물 참조와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지소영만 담은 미디엄 숏에서 현우 쪽으로 펼친 양손과 굳은 상체, 억제된 긴장 표정을 구현하며 참조의 연구복과 야간 원장실을 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영의 눈과 얼굴은 화면 오른쪽 남성을 향하고, 두 손도 그에게 뻗어 있다. 남성은 지소영을 마주 본다. 달래는 대상은 명확하지만 양 손바닥은 정면으로 보여 주기보다 서로 안쪽을 향해 다소 비스듬하다.",
        "built_space": "왼쪽에 목재 수납장 하나와 유리 칸막이, 뒤쪽에 검은 피아노 한 대와 벤치 하나, 오른쪽 아래에 목재 책상 하나가 보인다. 곡선 천장 조명과 벽의 간접조명은 참조 장소와 유사하다. 지소영은 책상 앞 빈 공간에 서 있고 남성은 오른쪽 전경을 차지한다. 유리에 보이는 조명 반사는 불가능해 보이지 않는다.",
        "entities": "검은 단발의 중년 한국인 여성으로 묘사된 지소영은 참조 얼굴과 대체로 유사하지만, 높은 깃과 비대칭 여밈의 연구복 대신 일반적인 라펠형 흰 가운을 입었다. 오른쪽에는 이전 숏의 현우처럼 보이는 검은 머리와 흰옷의 남성이 추가되어 있다. 오래된 검은 피아노는 유지되며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이번 화면에는 지소영만 등장할 수 있는데, 오른쪽 전경에 남성의 머리와 상체를 추가하여 이전 숏 인물을 가져오지 말라는 지시를 위반했다."
        ],
        "physics": "지소영의 손은 손목과 팔에 정상적으로 연결되어 있고 굽힌 팔꿈치로 양손을 내미는 동작이 가능하다. 두 사람의 발은 프레임 밖이지만 상체가 공중에 뜬 징후는 없다. 책상과 수납장은 다리로 바닥에 지지되며 소품은 상판 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "지소영은 렌즈보다 약간 왼쪽의 화면 밖 상대를 바라본다. 양손을 가슴 높이로 들어 손바닥을 카메라 쪽 방문객에게 펼쳐 보인다. 현우 자체는 보이지 않지만 시선과 제스처는 같은 전방 상대를 달래는 관계로 읽힌다.",
        "built_space": "왼쪽 전경에 열린 목재 문 하나와 경첩·손잡이가 있고, 그 뒤로 목재 수납장 하나와 유리 칸막이가 이어진다. 뒤 오른쪽에는 검은 피아노 한 대와 벤치 하나, 아래 오른쪽에는 목재 책상 하나와 부분적으로 보이는 의자 하나가 있다. 지소영은 문 안쪽에서 책상 앞에 서 있다. 참조의 곡선 천장, 유리의 조명 반사, 회색 바닥과 따뜻한 간접조명이 이어지며 중복된 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 사람은 지소영 한 명뿐이다. 중년 한국인 여성의 얼굴, 단정한 검은 단발, 체형이 인물 참조와 잘 맞는다. 높은 깃, 비대칭 금속 여밈과 회색 배색선이 있는 밝은 연구복도 참조를 따른다. 뒤의 검은 피아노가 유지되고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손은 각각 손목과 굽힌 팔에 연결되어 자연스럽게 유지된다. 펼친 손가락과 긴장된 어깨는 움직임을 억제하며 상대를 진정시키는 자세로 가능하다. 하체는 미디엄 숏 밖이어서 발 접점은 확인할 수 없지만 부유를 시사하는 자세는 아니다. 가구와 상판 위 물건에도 지지 관계의 이상은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우를 향한 달래는 동작은 보이지만, 등장 금지된 남성을 오른쪽 전경에 추가했고 연구복도 인물 참조와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지소영만 담은 미디엄 숏에서 현우 쪽으로 펼친 양손과 굳은 상체, 억제된 긴장 표정을 구현하며 참조의 연구복과 야간 원장실을 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "지소영의 눈과 얼굴은 화면 오른쪽 남성을 향하고, 두 손도 그에게 뻗어 있다. 남성은 지소영을 마주 본다. 달래는 대상은 명확하지만 양 손바닥은 정면으로 보여 주기보다 서로 안쪽을 향해 다소 비스듬하다.",
        "built_space": "왼쪽에 목재 수납장 하나와 유리 칸막이, 뒤쪽에 검은 피아노 한 대와 벤치 하나, 오른쪽 아래에 목재 책상 하나가 보인다. 곡선 천장 조명과 벽의 간접조명은 참조 장소와 유사하다. 지소영은 책상 앞 빈 공간에 서 있고 남성은 오른쪽 전경을 차지한다. 유리에 보이는 조명 반사는 불가능해 보이지 않는다.",
        "entities": "검은 단발의 중년 한국인 여성으로 묘사된 지소영은 참조 얼굴과 대체로 유사하지만, 높은 깃과 비대칭 여밈의 연구복 대신 일반적인 라펠형 흰 가운을 입었다. 오른쪽에는 이전 숏의 현우처럼 보이는 검은 머리와 흰옷의 남성이 추가되어 있다. 오래된 검은 피아노는 유지되며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이번 화면에는 지소영만 등장할 수 있는데, 오른쪽 전경에 남성의 머리와 상체를 추가하여 이전 숏 인물을 가져오지 말라는 지시를 위반했다."
        ],
        "physics": "지소영의 손은 손목과 팔에 정상적으로 연결되어 있고 굽힌 팔꿈치로 양손을 내미는 동작이 가능하다. 두 사람의 발은 프레임 밖이지만 상체가 공중에 뜬 징후는 없다. 책상과 수납장은 다리로 바닥에 지지되며 소품은 상판 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "지소영은 렌즈보다 약간 왼쪽의 화면 밖 상대를 바라본다. 양손을 가슴 높이로 들어 손바닥을 카메라 쪽 방문객에게 펼쳐 보인다. 현우 자체는 보이지 않지만 시선과 제스처는 같은 전방 상대를 달래는 관계로 읽힌다.",
        "built_space": "왼쪽 전경에 열린 목재 문 하나와 경첩·손잡이가 있고, 그 뒤로 목재 수납장 하나와 유리 칸막이가 이어진다. 뒤 오른쪽에는 검은 피아노 한 대와 벤치 하나, 아래 오른쪽에는 목재 책상 하나와 부분적으로 보이는 의자 하나가 있다. 지소영은 문 안쪽에서 책상 앞에 서 있다. 참조의 곡선 천장, 유리의 조명 반사, 회색 바닥과 따뜻한 간접조명이 이어지며 중복된 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 사람은 지소영 한 명뿐이다. 중년 한국인 여성의 얼굴, 단정한 검은 단발, 체형이 인물 참조와 잘 맞는다. 높은 깃, 비대칭 금속 여밈과 회색 배색선이 있는 밝은 연구복도 참조를 따른다. 뒤의 검은 피아노가 유지되고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손은 각각 손목과 굽힌 팔에 연결되어 자연스럽게 유지된다. 펼친 손가락과 긴장된 어깨는 움직임을 억제하며 상대를 진정시키는 자세로 가능하다. 하체는 미디엄 숏 밖이어서 발 접점은 확인할 수 없지만 부유를 시사하는 자세는 아니다. 가구와 상판 위 물건에도 지지 관계의 이상은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.762
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.512
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에서 엄격히 금지한 추가 인물(이전 샷의 남성) 포함",
     "[gpt-high] 이번 화면에는 지소영만 등장할 수 있는데, 오른쪽 전경에 남성의 머리와 상체를 추가하여 이전 숏 인물을 가져오지 말라는 지시를 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 512
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 단일 인물 제약과 의상 레퍼런스를 완벽히 준수하며, 요구된 미디엄 샷과 양손을 편 자세를 정확하게 구현함."
   },
   {
    "label": "B",
    "score": 512,
    "verdict_ko": "등장인물 제한 지침을 어기고 화면에 남성을 추가했으며, 지정된 연구복 의상도 일치하지 않아 치명적 위반임.  ★위반: [gemini-pro] 프롬프트에서 엄격히 금지한 추가 인물(이전 샷의 남성) 포함 / [gpt-high] 이번 화면에는 지소영만 등장할 수 있는데, 오른쪽 전경에 남성의 머리와 상체를 추가하여 이전 숏 인물을 가져오지 말라는 지시를 위반했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S86sh2_sel.png",
    "asset_id": "72181ee6-5ad7-419a-8fa9-0b5ede022b3d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:703308>",
    "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-10d6-74a5-aab7-14fc652ff39a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S86sh2"
  }
 },
 "S86sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:11:13.085570+00:00",
  "fingerprint": "0c03663a2f972d995b7ef19d4b11226b492641bee29176208266c01d3dc5e442",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S86sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S86sh5_sel.png",
  "source_sha256": "bad2d8172e6909765d1a44b266b2bb1e62edab0608aaad55e5dde3e7b30326f8",
  "file": "S86sh5_cine.png",
  "staged_sha256": "33cd92cae716c2c572cb3938a9223d2ba13522fd769d188c5105460d647398a2",
  "latency_ms": 10593
 },
 "S87sh4::signage": {
  "fp": "4579be01bc0f5441",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::acc0168d60745657": {
  "subjects": [],
  "subject_text": "특임대 이강준의 사무실\n문으로 출입하는 군 부대 사무실. 업무용 책상과 문건을 놓을 수 있는 실무 공간으로 구성된다.",
  "identity": "canonical",
  "scope_id": "L273",
  "scope_role": "location_interior",
  "scope_sha": "d3aa87b72484aa83"
 },
 "groupbg::special_unit_office": {
  "input_fingerprint": "4b5b811772f286dc",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "special_unit_office",
    "tags": [
     "S87sh4"
    ]
   },
   "context_sig": "0a0ddcecc943e9d8"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 절도 있는 보폭으로 걸어가는 군홧발. 틸업하면 특임대장 사무실\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 절도 있는 보폭으로 걸어가는 군홧발. 틸업하면 특임대장 사무실\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_special_unit_office_2f9882.png",
  "asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2",
  "input_asset_ids": [
   "50017786-0851-4a08-8e8f-72323e26d3ef"
  ],
  "origin_tag": "S87sh4",
  "place_text": "Inside the special-operations commander's office, in the briefing area under nighttime office lighting.",
  "origin_inputs": {
   "place_text": "Inside the special-operations commander's office, in the briefing area under nighttime office lighting.",
   "time_of_day_en": "night",
   "conti_asset_id": "50017786-0851-4a08-8e8f-72323e26d3ef"
  }
 },
 "S87sh4::bgfirst_bg": {
  "input_fingerprint": "07f6cefe9afb97ea",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh4__bgfirst_bg.png",
  "asset_id": "4e972bfb-a762-4f2f-9abf-8faa5c7623ed",
  "input_asset_ids": [
   "50017786-0851-4a08-8e8f-72323e26d3ef",
   "4964e681-31dd-4c79-aad2-e3da5f4b52d2"
  ]
 },
 "S87sh4": {
  "input_fingerprint": "ce921913108e96db",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 이강준 (한국인, 성인 남성, 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 이강준 (한국인, 성인 남성, 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 이강준 (한국인, 성인 남성, 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh4__bgfirst_bg.png",
     "asset_id": "4e972bfb-a762-4f2f-9abf-8faa5c7623ed",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S87sh4.png",
     "asset_id": "50017786-0851-4a08-8e8f-72323e26d3ef",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 이강준: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1327185>",
     "asset_id": "6dbc9554-4775-4399-b597-16c2aa160d07",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_special_unit_office_2f9882.png",
     "asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 이강준: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1327185>",
     "asset_id": "6dbc9554-4775-4399-b597-16c2aa160d07",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "주 인물은 화면 오른쪽 전경에 있는 남자를 향해 시선을 두고 소리치고 있음.",
    "built_space": "야간 지휘관실. 데스크, 대형 스크린, 회의 테이블, 깃발 등 레퍼런스의 구조와 일치함.",
    "entities": "이강준(성인 남성, 짧은 검은 머리)이 묘사된 행동을 취하고 있으나, 프롬프트에 없는 남자가 전경에 나타남.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)",
     "[gpt-high] 화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
    ],
    "physics": "인물은 바닥에 발을 딛고 자연스럽게 서서 체중을 지탱함."
   },
   {
    "label": "B",
    "direction": "인물은 화면 왼쪽 밖 허공을 향해 시선을 두고 주먹을 들어 올림.",
    "built_space": "야간 지휘관실. 스크린, 깃발, 회의 테이블과 의자 배치가 레퍼런스와 일치함.",
    "entities": "이강준(성인 남성, 짧은 검은 머리)이 레퍼런스의 외형과 일치하며, 지시된 표정과 주먹 쥔 상체를 보여줌.",
    "hard_violations": [],
    "physics": "인물은 안정적인 자세로 서서 행동을 취하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 언급되지 않은 인물이 전경에 추가되어 결정적인 위반 사항이 발생했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미디엄 샷 구도 내에서 허공을 향한 주먹, 핏대 선 목, 벌린 입 등 명시된 인물의 행동과 표정을 정확하게 재현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "주 인물은 화면 오른쪽 전경에 있는 남자를 향해 시선을 두고 소리치고 있음.",
        "built_space": "야간 지휘관실. 데스크, 대형 스크린, 회의 테이블, 깃발 등 레퍼런스의 구조와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 묘사된 행동을 취하고 있으나, 프롬프트에 없는 남자가 전경에 나타남.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)"
        ],
        "physics": "인물은 바닥에 발을 딛고 자연스럽게 서서 체중을 지탱함."
       },
       {
        "label": "B",
        "direction": "인물은 화면 왼쪽 밖 허공을 향해 시선을 두고 주먹을 들어 올림.",
        "built_space": "야간 지휘관실. 스크린, 깃발, 회의 테이블과 의자 배치가 레퍼런스와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 레퍼런스의 외형과 일치하며, 지시된 표정과 주먹 쥔 상체를 보여줌.",
        "hard_violations": [],
        "physics": "인물은 안정적인 자세로 서서 행동을 취하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 언급되지 않은 인물이 전경에 추가되어 결정적인 위반 사항이 발생했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미디엄 샷 구도 내에서 허공을 향한 주먹, 핏대 선 목, 벌린 입 등 명시된 인물의 행동과 표정을 정확하게 재현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "주 인물은 화면 오른쪽 전경에 있는 남자를 향해 시선을 두고 소리치고 있음.",
        "built_space": "야간 지휘관실. 데스크, 대형 스크린, 회의 테이블, 깃발 등 레퍼런스의 구조와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 묘사된 행동을 취하고 있으나, 프롬프트에 없는 남자가 전경에 나타남.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)"
        ],
        "physics": "인물은 바닥에 발을 딛고 자연스럽게 서서 체중을 지탱함."
       },
       {
        "label": "B",
        "direction": "인물은 화면 왼쪽 밖 허공을 향해 시선을 두고 주먹을 들어 올림.",
        "built_space": "야간 지휘관실. 스크린, 깃발, 회의 테이블과 의자 배치가 레퍼런스와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 레퍼런스의 외형과 일치하며, 지시된 표정과 주먹 쥔 상체를 보여줌.",
        "hard_violations": [],
        "physics": "인물은 안정적인 자세로 서서 행동을 취하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "상체 중심의 미디엄 숏에서 허공에 쥔 주먹, 크게 벌린 입과 목의 핏대를 정확히 살렸으며, 기준 인물의 정장과 야간 사무실도 유지했다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "허공의 주먹과 외치는 표정은 맞지만, 금지된 두 번째 인물을 추가했고 허벅지까지 넓힌 구도와 검은 티셔츠도 지정된 상체 숏 및 기준 복장에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "이강준은 화면 오른쪽 바깥을 바라보며 입을 크게 벌리고 있다. 화면 왼쪽의 주먹은 머리보다 높게 들어 빈 공간을 향하며, 사람이나 물체를 가격하는 방향이 아니다. 출동을 지시하는 듯한 시선과 허공에 쥔 주먹이라는 지시가 맞는다.",
        "built_space": "왼쪽 집무 책상 1개, 뒤쪽 벽면의 다중 화면 설비 1세트, 깃발 1개, 지도 부조 1개, 낮은 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽에는 회의 테이블 1개와 명확히 구분되는 의자 4개, 야경이 보이는 창이 있다. 인물은 책상 앞과 회의 구역 사이에 서 있으며 가구와 겹쳐 관통하지 않는다. 회색 벽체와 따뜻한 간접조명, 창의 위치가 장소 기준과 잘 맞고 불가능한 반사는 보이지 않는다.",
        "entities": "성인 남성 1명만 등장하며, 한국인으로 설정된 기준 인물의 짧은 검은 머리, 얼굴 특징과 체격에 대체로 부합한다. 짙은 정장, 흰 셔츠와 어두운 넥타이도 유지된다. 크게 열린 입과 도드라진 목의 힘줄이 보이고 눈은 정상적인 사람의 눈이다. 셔츠 깃과 넥타이는 기준 사진보다 느슨하지만 동작 중 복장으로 자연스럽다. 화면이나 사물에서 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "들어 올린 주먹은 손목과 굽힌 팔, 어깨에 자연스럽게 연결되어 근육으로 지지된다. 앞으로 기운 상체는 화면 아래로 이어지는 몸통에 연결되며 떠 있는 모습이 아니다. 발은 상체 중심 구도 밖이므로 접지는 확인할 수 없지만, 지지 없는 공중 자세를 시사하지 않는다. 책상과 의자는 바닥에 놓이고 화분은 수납장 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "중앙 남성은 오른쪽 전경에 있는 다른 남성을 바라보며 외치고, 전경 남성도 그를 향해 몸과 얼굴을 돌리고 있다. 올라간 주먹 자체는 머리 위의 빈 공간을 향해 있어 허공에 쥔 주먹이라는 동작은 맞지만, 시선의 수신자로 허용되지 않은 인물을 화면에 추가했다.",
        "built_space": "왼쪽 집무 책상 1개와 높은 등받이 의자 1개, 깃발 1개, 벽면 화면 설비 1세트, 지도 부조 1개, 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽 회의 구역에는 전경과 후경으로 나뉘어 보이는 탁자 면과 최소 4개의 의자 등받이가 있다. 회색 벽체, 간접조명과 창밖 야경은 기준 장소를 따른다. 주인공은 책상 앞에 서 있지만 오른쪽 전경을 추가 인물의 어깨와 머리가 크게 가려, 지정된 단독 상체 장면이 대화 상대의 어깨 너머 구도로 바뀌었다.",
        "entities": "외치는 성인 남성은 짧은 검은 머리와 얼굴에서 이강준 기준을 대체로 따르지만, 정장·흰 셔츠·넥타이 대신 몸에 붙는 검은 긴팔 티셔츠를 입었다. 오른쪽에는 검은 머리에 정장을 입은 별도의 성인 남성이 부분적으로 등장한다. 허용 인물은 이강준 한 명뿐이므로 이 추가 인물은 명백한 위반이다. 주인공의 열린 입과 목의 핏대는 잘 보이며, 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
        ],
        "physics": "주인공의 올라간 주먹은 굽힌 팔과 어깨로 지지되고, 내려간 주먹도 팔 끝에 자연스럽게 연결된다. 상체는 벌린 다리 쪽으로 이어져 외치는 서 있는 자세로 가능하다. 발은 화면 밖이지만 몸이 공중에 떠 있는 징후는 없다. 전경 남성 역시 몸통이 화면 아래로 이어지고, 보이는 가구와 소품도 바닥이나 가구 면에 지지되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "상체 중심의 미디엄 숏에서 허공에 쥔 주먹, 크게 벌린 입과 목의 핏대를 정확히 살렸으며, 기준 인물의 정장과 야간 사무실도 유지했다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "허공의 주먹과 외치는 표정은 맞지만, 금지된 두 번째 인물을 추가했고 허벅지까지 넓힌 구도와 검은 티셔츠도 지정된 상체 숏 및 기준 복장에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "이강준은 화면 오른쪽 바깥을 바라보며 입을 크게 벌리고 있다. 화면 왼쪽의 주먹은 머리보다 높게 들어 빈 공간을 향하며, 사람이나 물체를 가격하는 방향이 아니다. 출동을 지시하는 듯한 시선과 허공에 쥔 주먹이라는 지시가 맞는다.",
        "built_space": "왼쪽 집무 책상 1개, 뒤쪽 벽면의 다중 화면 설비 1세트, 깃발 1개, 지도 부조 1개, 낮은 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽에는 회의 테이블 1개와 명확히 구분되는 의자 4개, 야경이 보이는 창이 있다. 인물은 책상 앞과 회의 구역 사이에 서 있으며 가구와 겹쳐 관통하지 않는다. 회색 벽체와 따뜻한 간접조명, 창의 위치가 장소 기준과 잘 맞고 불가능한 반사는 보이지 않는다.",
        "entities": "성인 남성 1명만 등장하며, 한국인으로 설정된 기준 인물의 짧은 검은 머리, 얼굴 특징과 체격에 대체로 부합한다. 짙은 정장, 흰 셔츠와 어두운 넥타이도 유지된다. 크게 열린 입과 도드라진 목의 힘줄이 보이고 눈은 정상적인 사람의 눈이다. 셔츠 깃과 넥타이는 기준 사진보다 느슨하지만 동작 중 복장으로 자연스럽다. 화면이나 사물에서 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "들어 올린 주먹은 손목과 굽힌 팔, 어깨에 자연스럽게 연결되어 근육으로 지지된다. 앞으로 기운 상체는 화면 아래로 이어지는 몸통에 연결되며 떠 있는 모습이 아니다. 발은 상체 중심 구도 밖이므로 접지는 확인할 수 없지만, 지지 없는 공중 자세를 시사하지 않는다. 책상과 의자는 바닥에 놓이고 화분은 수납장 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "중앙 남성은 오른쪽 전경에 있는 다른 남성을 바라보며 외치고, 전경 남성도 그를 향해 몸과 얼굴을 돌리고 있다. 올라간 주먹 자체는 머리 위의 빈 공간을 향해 있어 허공에 쥔 주먹이라는 동작은 맞지만, 시선의 수신자로 허용되지 않은 인물을 화면에 추가했다.",
        "built_space": "왼쪽 집무 책상 1개와 높은 등받이 의자 1개, 깃발 1개, 벽면 화면 설비 1세트, 지도 부조 1개, 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽 회의 구역에는 전경과 후경으로 나뉘어 보이는 탁자 면과 최소 4개의 의자 등받이가 있다. 회색 벽체, 간접조명과 창밖 야경은 기준 장소를 따른다. 주인공은 책상 앞에 서 있지만 오른쪽 전경을 추가 인물의 어깨와 머리가 크게 가려, 지정된 단독 상체 장면이 대화 상대의 어깨 너머 구도로 바뀌었다.",
        "entities": "외치는 성인 남성은 짧은 검은 머리와 얼굴에서 이강준 기준을 대체로 따르지만, 정장·흰 셔츠·넥타이 대신 몸에 붙는 검은 긴팔 티셔츠를 입었다. 오른쪽에는 검은 머리에 정장을 입은 별도의 성인 남성이 부분적으로 등장한다. 허용 인물은 이강준 한 명뿐이므로 이 추가 인물은 명백한 위반이다. 주인공의 열린 입과 목의 핏대는 잘 보이며, 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
        ],
        "physics": "주인공의 올라간 주먹은 굽힌 팔과 어깨로 지지되고, 내려간 주먹도 팔 끝에 자연스럽게 연결된다. 상체는 벌린 다리 쪽으로 이어져 외치는 서 있는 자세로 가능하다. 발은 화면 밖이지만 몸이 공중에 떠 있는 징후는 없다. 전경 남성 역시 몸통이 화면 아래로 이어지고, 보이는 가구와 소품도 바닥이나 가구 면에 지지되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.651,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.401,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)",
     "[gpt-high] 화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 401,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 401,
    "verdict_ko": "프롬프트에 언급되지 않은 인물이 전경에 추가되어 결정적인 위반 사항이 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽) / [gpt-high] 화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "미디엄 샷 구도 내에서 허공을 향한 주먹, 핏대 선 목, 벌린 입 등 명시된 인물의 행동과 표정을 정확하게 재현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_special_unit_office_2f9882.png",
    "asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 이강준: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1327185>",
    "asset_id": "6dbc9554-4775-4399-b597-16c2aa160d07",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-127d-7ef8-b7cd-7afed4aa4113",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh4__bgfirst_bg.png",
   "bg_asset_id": "4e972bfb-a762-4f2f-9abf-8faa5c7623ed",
   "bg_record_key": "S87sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "special_unit_office",
   "groupbg_asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S87sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:33:38.921272+00:00",
  "fingerprint": "6af1100358607380c68f50dab733b9cd28acde35d816c6110a5bd59093453ee3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S87sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S87sh4_sel.png",
  "source_sha256": "2b4ad413de014687fa5dd1fdffcd9b712b6877b8fdc6b6c041cb0c101273f908",
  "file": "S87sh4_cine.png",
  "staged_sha256": "8f357df37bfb071a99c61ad02da5674b5a3b62f8b857eed14a93a51653f8efec",
  "latency_ms": 9562
 },
 "S87sh5::signage": {
  "fp": "9ee41674796092ed",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::corporate_chair_office": {
  "input_fingerprint": "748a8191ab997abe",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "corporate_chair_office",
    "tags": [
     "S87sh5"
    ]
   },
   "context_sig": "98a913aacd2b6fe2"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /유빅사 윤성찬 회장실 –N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /유빅사 윤성찬 회장실 –N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_corporate_chair_office_dd85e5.png",
  "asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308",
  "input_asset_ids": [
   "f85ce6f3-3160-4f21-847d-73d55d234f47"
  ],
  "origin_tag": "S87sh5",
  "place_text": "Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.",
  "origin_inputs": {
   "place_text": "Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.",
   "time_of_day_en": "night",
   "conti_asset_id": "f85ce6f3-3160-4f21-847d-73d55d234f47"
  }
 },
 "S87sh5::bgfirst_bg": {
  "input_fingerprint": "3a302fccdeace35b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh5__bgfirst_bg.png",
  "asset_id": "ea1e3ce7-0a30-49e0-a731-7053a0115a46",
  "input_asset_ids": [
   "f85ce6f3-3160-4f21-847d-73d55d234f47",
   "0bb05930-fe04-4dc3-b470-8ae77659d308"
  ]
 },
 "S87sh5": {
  "input_fingerprint": "891bd1738d677ba3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 윤성찬 right now, so 윤성찬's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 윤성찬: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 윤성찬 right now, so 윤성찬's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 윤성찬: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 윤성찬 right now, so 윤성찬's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 윤성찬: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh5__bgfirst_bg.png",
     "asset_id": "ea1e3ce7-0a30-49e0-a731-7053a0115a46",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S87sh5.png",
     "asset_id": "f85ce6f3-3160-4f21-847d-73d55d234f47",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:886743>",
     "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_corporate_chair_office_dd85e5.png",
     "asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:886743>",
     "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 우측 밖을 향하며, 오른손은 오른쪽 귀에 꽂힌 이어폰을 향해 있음.",
    "built_space": "야경이 보이는 창문, 데스크, 모니터 등 지정된 사무실 배경이 보이나 인물이 배경 공간과 분리되어 합성된 것처럼 배치됨.",
    "entities": "윤성찬(얼굴과 안경 일치 / 정장과 넥타이 누락). 귀에 꽂힌 이어폰과 손이 존재함. 표정은 비열하지 않고 온화함.",
    "hard_violations": [
     "[gpt-high] 참조에 있는 회장석과 감시석을 그대로 남긴 채 인물용 업무 의자를 전경에 하나 더 추가했다.",
     "[gpt-high] 인물을 지정된 통신 감시석이 아니라 별도로 만든 전경 좌석에 앉혔다."
    ],
    "physics": "손가락이 귀와 이어폰에 닿아 있으나, 인물과 배경 사이의 조명과 깊이감이 물리적으로 어색함."
   },
   {
    "label": "B",
    "direction": "시선은 화면 우측 하단을 향하고, 오른손은 오른쪽 귀의 이어폰에 닿아 있음.",
    "built_space": "좌측의 창문과 우측의 모니터 데스크 등 지정된 사무실 공간이 아웃포커싱되어 자연스럽게 배치됨.",
    "entities": "윤성찬(얼굴, 주름, 의상 일치 / 안경 누락). 이어폰에 손을 댄 모습과 비열하게 입꼬리를 올린 표정 연기가 지시문과 정확히 일치함.",
    "hard_violations": [
     "[gpt-high] 통신 감시 위치에서 벌어져야 하는 장면인데, 감시 장비를 뒤에 둔 전면 회장석에 인물을 배치했다."
    ],
    "physics": "오른손이 이어폰을 쥔 채 귀에 자연스럽게 밀착되어 있으며, 인물에 드리워진 조명이 공간과 잘 어우러짐."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 안경이 누락되었으나, 지시된 의상(스트라이프 정장)을 준수하고 '비열한 미소'와 제한된 조명 분위기를 매우 사실적으로 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 정장 자켓과 넥타이가 누락되었으며, 지시된 비열한 표정 대신 온화한 미소를 띠고 있어 프롬프트의 핵심 무드와 어긋남."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 밖을 향하며, 오른손은 오른쪽 귀에 꽂힌 이어폰을 향해 있음.",
        "built_space": "야경이 보이는 창문, 데스크, 모니터 등 지정된 사무실 배경이 보이나 인물이 배경 공간과 분리되어 합성된 것처럼 배치됨.",
        "entities": "윤성찬(얼굴과 안경 일치 / 정장과 넥타이 누락). 귀에 꽂힌 이어폰과 손이 존재함. 표정은 비열하지 않고 온화함.",
        "hard_violations": [],
        "physics": "손가락이 귀와 이어폰에 닿아 있으나, 인물과 배경 사이의 조명과 깊이감이 물리적으로 어색함."
       },
       {
        "label": "B",
        "direction": "시선은 화면 우측 하단을 향하고, 오른손은 오른쪽 귀의 이어폰에 닿아 있음.",
        "built_space": "좌측의 창문과 우측의 모니터 데스크 등 지정된 사무실 공간이 아웃포커싱되어 자연스럽게 배치됨.",
        "entities": "윤성찬(얼굴, 주름, 의상 일치 / 안경 누락). 이어폰에 손을 댄 모습과 비열하게 입꼬리를 올린 표정 연기가 지시문과 정확히 일치함.",
        "hard_violations": [],
        "physics": "오른손이 이어폰을 쥔 채 귀에 자연스럽게 밀착되어 있으며, 인물에 드리워진 조명이 공간과 잘 어우러짐."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 안경이 누락되었으나, 지시된 의상(스트라이프 정장)을 준수하고 '비열한 미소'와 제한된 조명 분위기를 매우 사실적으로 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 정장 자켓과 넥타이가 누락되었으며, 지시된 비열한 표정 대신 온화한 미소를 띠고 있어 프롬프트의 핵심 무드와 어긋남."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 밖을 향하며, 오른손은 오른쪽 귀에 꽂힌 이어폰을 향해 있음.",
        "built_space": "야경이 보이는 창문, 데스크, 모니터 등 지정된 사무실 배경이 보이나 인물이 배경 공간과 분리되어 합성된 것처럼 배치됨.",
        "entities": "윤성찬(얼굴과 안경 일치 / 정장과 넥타이 누락). 귀에 꽂힌 이어폰과 손이 존재함. 표정은 비열하지 않고 온화함.",
        "hard_violations": [],
        "physics": "손가락이 귀와 이어폰에 닿아 있으나, 인물과 배경 사이의 조명과 깊이감이 물리적으로 어색함."
       },
       {
        "label": "B",
        "direction": "시선은 화면 우측 하단을 향하고, 오른손은 오른쪽 귀의 이어폰에 닿아 있음.",
        "built_space": "좌측의 창문과 우측의 모니터 데스크 등 지정된 사무실 공간이 아웃포커싱되어 자연스럽게 배치됨.",
        "entities": "윤성찬(얼굴, 주름, 의상 일치 / 안경 누락). 이어폰에 손을 댄 모습과 비열하게 입꼬리를 올린 표정 연기가 지시문과 정확히 일치함.",
        "hard_violations": [],
        "physics": "오른손이 이어폰을 쥔 채 귀에 자연스럽게 밀착되어 있으며, 인물에 드리워진 조명이 공간과 잘 어우러짐."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "얼굴 중심 클로즈업과 이어폰 접촉, 치켜올린 입꼬리는 더 충실하지만, 통신 감시석이 아닌 전면 회장석에 배치되어 있으며 참조의 안경도 빠졌다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "안경과 고령의 외모는 맞지만, 별도 의자를 추가한 전경 배치가 지정 위치를 벗어나고 구도도 상반신과 사무실을 과하게 넓혀 보여준다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "눈은 카메라 가까운 화면 오른쪽의 화면 밖을 향하며, 특정 감시 화면을 보고 있지는 않다. 시선 대상은 지문에 고정되어 있지 않다. 들어 올린 검지는 가까운 귀의 이어폰 바깥 부분에 닿아 있어 동작의 대상은 정확하다.",
        "built_space": "왼쪽에 야경이 보이는 통유리와 창틀, 인물 바로 뒤에 큰 검은 가죽 의자 하나, 왼쪽 가장자리에 책상 조명 하나가 보인다. 오른쪽 뒤에는 감시 화면 일부와 탁상등 하나, 검은 깃발 일부가 보인다. 재료와 야간 조명은 장소 참조에 부합하지만, 인물은 감시 장비 앞이 아니라 그보다 앞쪽 회장석에 놓여 있다. 창의 선형 조명 반사에는 뚜렷한 광학적 모순이 없다.",
        "entities": "고령의 한국인 남성으로 보이는 인물 한 명이며, 회색으로 빗어 넘긴 머리와 깊은 이마·눈가 주름은 윤성찬의 설정에 맞는다. 참조와 달리 안경이 없고 얼굴 인상에도 차이가 있다. 흰 셔츠, 어두운 넥타이, 줄무늬 정장 상의가 보이며 참조의 짙은 외투 차림과는 다르다. 귀에 검정·은색 통신 이어폰 하나가 있고 외부 부분이 식별된다. 입꼬리를 올리고 눈을 좁힌 웃음은 요구한 비열한 표정에 비교적 가깝다. 읽을 수 있는 문자는 없다.",
        "hard_violations": [
         "통신 감시 위치에서 벌어져야 하는 장면인데, 감시 장비를 뒤에 둔 전면 회장석에 인물을 배치했다."
        ],
        "physics": "이어폰은 귀에 삽입되어 지지되고 손가락이 바깥 부분에 접촉한다. 주름진 손은 인물의 나이와 어울리며 손목과 소매가 프레임 아래로 자연스럽게 이어진다. 몸통은 화면 아래로 이어지고 뒤에 의자 등받이가 있어 부유한 자세로 보이지 않는다. 손이나 기기의 지지 불가능, 명백한 해부학적 오류는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "얼굴과 두 눈은 화면 오른쪽의 화면 밖을 향한다. 배경 감시 모니터를 보는 방향은 아니지만 지문은 특정 시선 대상을 요구하지 않는다. 검지는 가까운 귀에 꽂힌 이어폰의 노출 부분을 직접 누르고 있어 사용 방향이 맞는다.",
        "built_space": "통유리 야경, 천장 간접조명, 오른쪽 대리석 벽과 깃발, 책상 조명 두 개가 참조의 공간을 재현한다. 뒤쪽 감시대에는 모니터 두 면이 드러나고 나머지 영역은 인물에 가려져 전체 수량을 단정할 수 없다. 의자는 왼쪽 회장석 하나, 뒤쪽 감시석 하나, 인물이 앉은 전경 의자 하나로 세 개가 보인다. 참조의 두 업무용 의자 외에 전경 의자를 추가했고 인물은 지정된 감시석과 떨어져 있다. 창에 비친 선형 조명은 가능한 반사로 보인다.",
        "entities": "고령의 한국인 남성 한 명으로, 회색 머리와 안경, 깊은 주름은 참조 윤성찬의 특징에 부합한다. 다만 회색 셔츠와 줄무늬 조끼를 입고 넥타이와 외투가 없어 참조의 복장과 다르다. 귀에는 작은 유선형 이어폰 하나가 있으며 손가락 접촉도 보인다. 웃음은 뚜렷하지만 비열하게 입꼬리를 한껏 올린 표정보다는 온화한 미소에 가깝다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "참조에 있는 회장석과 감시석을 그대로 남긴 채 인물용 업무 의자를 전경에 하나 더 추가했다.",
         "인물을 지정된 통신 감시석이 아니라 별도로 만든 전경 좌석에 앉혔다."
        ],
        "physics": "등과 몸통 뒤에 가죽 의자 등받이가 있으며 앉은 상체의 지지는 자연스럽다. 이어폰은 귀에 걸리고 삽입된 부분으로 지지되며, 검지가 외부 부분에 닿는다. 손목과 전완은 셔츠 소매로 이어지고 손의 나이도 얼굴과 어울린다. 이어폰 선은 아래로 늘어져 중력 방향에 맞으며, 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 중심 클로즈업과 이어폰 접촉, 치켜올린 입꼬리는 더 충실하지만, 통신 감시석이 아닌 전면 회장석에 배치되어 있으며 참조의 안경도 빠졌다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "안경과 고령의 외모는 맞지만, 별도 의자를 추가한 전경 배치가 지정 위치를 벗어나고 구도도 상반신과 사무실을 과하게 넓혀 보여준다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "눈은 카메라 가까운 화면 오른쪽의 화면 밖을 향하며, 특정 감시 화면을 보고 있지는 않다. 시선 대상은 지문에 고정되어 있지 않다. 들어 올린 검지는 가까운 귀의 이어폰 바깥 부분에 닿아 있어 동작의 대상은 정확하다.",
        "built_space": "왼쪽에 야경이 보이는 통유리와 창틀, 인물 바로 뒤에 큰 검은 가죽 의자 하나, 왼쪽 가장자리에 책상 조명 하나가 보인다. 오른쪽 뒤에는 감시 화면 일부와 탁상등 하나, 검은 깃발 일부가 보인다. 재료와 야간 조명은 장소 참조에 부합하지만, 인물은 감시 장비 앞이 아니라 그보다 앞쪽 회장석에 놓여 있다. 창의 선형 조명 반사에는 뚜렷한 광학적 모순이 없다.",
        "entities": "고령의 한국인 남성으로 보이는 인물 한 명이며, 회색으로 빗어 넘긴 머리와 깊은 이마·눈가 주름은 윤성찬의 설정에 맞는다. 참조와 달리 안경이 없고 얼굴 인상에도 차이가 있다. 흰 셔츠, 어두운 넥타이, 줄무늬 정장 상의가 보이며 참조의 짙은 외투 차림과는 다르다. 귀에 검정·은색 통신 이어폰 하나가 있고 외부 부분이 식별된다. 입꼬리를 올리고 눈을 좁힌 웃음은 요구한 비열한 표정에 비교적 가깝다. 읽을 수 있는 문자는 없다.",
        "hard_violations": [
         "통신 감시 위치에서 벌어져야 하는 장면인데, 감시 장비를 뒤에 둔 전면 회장석에 인물을 배치했다."
        ],
        "physics": "이어폰은 귀에 삽입되어 지지되고 손가락이 바깥 부분에 접촉한다. 주름진 손은 인물의 나이와 어울리며 손목과 소매가 프레임 아래로 자연스럽게 이어진다. 몸통은 화면 아래로 이어지고 뒤에 의자 등받이가 있어 부유한 자세로 보이지 않는다. 손이나 기기의 지지 불가능, 명백한 해부학적 오류는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "얼굴과 두 눈은 화면 오른쪽의 화면 밖을 향한다. 배경 감시 모니터를 보는 방향은 아니지만 지문은 특정 시선 대상을 요구하지 않는다. 검지는 가까운 귀에 꽂힌 이어폰의 노출 부분을 직접 누르고 있어 사용 방향이 맞는다.",
        "built_space": "통유리 야경, 천장 간접조명, 오른쪽 대리석 벽과 깃발, 책상 조명 두 개가 참조의 공간을 재현한다. 뒤쪽 감시대에는 모니터 두 면이 드러나고 나머지 영역은 인물에 가려져 전체 수량을 단정할 수 없다. 의자는 왼쪽 회장석 하나, 뒤쪽 감시석 하나, 인물이 앉은 전경 의자 하나로 세 개가 보인다. 참조의 두 업무용 의자 외에 전경 의자를 추가했고 인물은 지정된 감시석과 떨어져 있다. 창에 비친 선형 조명은 가능한 반사로 보인다.",
        "entities": "고령의 한국인 남성 한 명으로, 회색 머리와 안경, 깊은 주름은 참조 윤성찬의 특징에 부합한다. 다만 회색 셔츠와 줄무늬 조끼를 입고 넥타이와 외투가 없어 참조의 복장과 다르다. 귀에는 작은 유선형 이어폰 하나가 있으며 손가락 접촉도 보인다. 웃음은 뚜렷하지만 비열하게 입꼬리를 한껏 올린 표정보다는 온화한 미소에 가깝다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "참조에 있는 회장석과 감시석을 그대로 남긴 채 인물용 업무 의자를 전경에 하나 더 추가했다.",
         "인물을 지정된 통신 감시석이 아니라 별도로 만든 전경 좌석에 앉혔다."
        ],
        "physics": "등과 몸통 뒤에 가죽 의자 등받이가 있으며 앉은 상체의 지지는 자연스럽다. 이어폰은 귀에 걸리고 삽입된 부분으로 지지되며, 검지가 외부 부분에 닿는다. 손목과 전완은 셔츠 소매로 이어지고 손의 나이도 얼굴과 어울린다. 이어폰 선은 아래로 늘어져 중력 방향에 맞으며, 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.071,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.821,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gpt-high] 통신 감시 위치에서 벌어져야 하는 장면인데, 감시 장비를 뒤에 둔 전면 회장석에 인물을 배치했다."
    ],
    "A": [
     "[gpt-high] 참조에 있는 회장석과 감시석을 그대로 남긴 채 인물용 업무 의자를 전경에 하나 더 추가했다.",
     "[gpt-high] 인물을 지정된 통신 감시석이 아니라 별도로 만든 전경 좌석에 앉혔다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 821
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "레퍼런스의 안경이 누락되었으나, 지시된 의상(스트라이프 정장)을 준수하고 '비열한 미소'와 제한된 조명 분위기를 매우 사실적으로 구현함.  ★위반: [gpt-high] 통신 감시 위치에서 벌어져야 하는 장면인데, 감시 장비를 뒤에 둔 전면 회장석에 인물을 배치했다."
   },
   {
    "label": "A",
    "score": 821,
    "verdict_ko": "지정된 정장 자켓과 넥타이가 누락되었으며, 지시된 비열한 표정 대신 온화한 미소를 띠고 있어 프롬프트의 핵심 무드와 어긋남.  ★위반: [gpt-high] 참조에 있는 회장석과 감시석을 그대로 남긴 채 인물용 업무 의자를 전경에 하나 더 추가했다. / [gpt-high] 인물을 지정된 통신 감시석이 아니라 별도로 만든 전경 좌석에 앉혔다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_corporate_chair_office_dd85e5.png",
    "asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-174c-7921-a4e7-29ef6a842e56",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh5__bgfirst_bg.png",
   "bg_asset_id": "ea1e3ce7-0a30-49e0-a731-7053a0115a46",
   "bg_record_key": "S87sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "corporate_chair_office",
   "groupbg_asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S87sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:12:19.502666+00:00",
  "fingerprint": "1e1b3a2ab08ef5897c14d7d116fded4bb00b79413431eb947349c5f71de88dcd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S87sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S87sh5_sel.png",
  "source_sha256": "2f5b6e36d5bd49257b8a6e0b97273d52310ed01fa4d7a87dbf59f5e80287f923",
  "file": "S87sh5_cine.png",
  "staged_sha256": "47b498d1e814835c7013b419c6fcc1ec07fb8b5dcb4f9ea626be5295ae1e9f93",
  "latency_ms": 12243
 },
 "S88sh7::signage": {
  "fp": "060df420176958a3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S88sh7": {
  "input_fingerprint": "332ccbe47d001f85",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 켜진 실험실 안, 차가운 스테인리스 침대 위에 굵은 끈으로 결박당한 채 누워있는 찰리의 낡은 전신.\n\nLOCATION (lock): On the stainless-steel examination bed inside the research laboratory's glass enclosure, under blue laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Stainless-steel bed, headward end toward upper right in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (Supporting 찰리's restrained body) — The upper surface, near long edge, and footward end are visible from above; used as Diagonal spatial anchor that exposes the full-body restraint arrangement; Thick restraints (Secured around 찰리, holding him to the bed); used as Readable evidence of confinement across the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The specified blue laboratory illumination gives the restrained body a subdued, cold appearance while preserving readable detail in the restraints.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the stainless-steel bed with his body restrained against it, including his hands before they are released. The source does not specify the restraint points, his head's direction, or the exact arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies restrained on the stainless-steel laboratory bed; the chest-ring power system remains impaired after the failed test. The laboratory's glass enclosure and monitoring equipment remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 켜진 실험실 안, 차가운 스테인리스 침대 위에 굵은 끈으로 결박당한 채 누워있는 찰리의 낡은 전신.\n\nLOCATION (lock): On the stainless-steel examination bed inside the research laboratory's glass enclosure, under blue laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Stainless-steel bed, headward end toward upper right in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (Supporting 찰리's restrained body) — The upper surface, near long edge, and footward end are visible from above; used as Diagonal spatial anchor that exposes the full-body restraint arrangement; Thick restraints (Secured around 찰리, holding him to the bed); used as Readable evidence of confinement across the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The specified blue laboratory illumination gives the restrained body a subdued, cold appearance while preserving readable detail in the restraints.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the stainless-steel bed with his body restrained against it, including his hands before they are released. The source does not specify the restraint points, his head's direction, or the exact arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies restrained on the stainless-steel laboratory bed; the chest-ring power system remains impaired after the failed test. The laboratory's glass enclosure and monitoring equipment remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 켜진 실험실 안, 차가운 스테인리스 침대 위에 굵은 끈으로 결박당한 채 누워있는 찰리의 낡은 전신.\n\nLOCATION (lock): On the stainless-steel examination bed inside the research laboratory's glass enclosure, under blue laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Stainless-steel bed, headward end toward upper right in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (Supporting 찰리's restrained body) — The upper surface, near long edge, and footward end are visible from above; used as Diagonal spatial anchor that exposes the full-body restraint arrangement; Thick restraints (Secured around 찰리, holding him to the bed); used as Readable evidence of confinement across the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The specified blue laboratory illumination gives the restrained body a subdued, cold appearance while preserving readable detail in the restraints.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the stainless-steel bed with his body restrained against it, including his hands before they are released. The source does not specify the restraint points, his head's direction, or the exact arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies restrained on the stainless-steel laboratory bed; the chest-ring power system remains impaired after the failed test. The laboratory's glass enclosure and monitoring equipment remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 머리는 프레임 우측 상단을 향하고 있음.",
    "built_space": "유리벽 안 실험실 중앙에 스테인리스 침대가 배치됨. 배경의 연구원석은 지시대로 모두 비어 있음.",
    "entities": "찰리(베이지색 장갑판 로봇 몸체), 굵은 결박 끈. 찰리 외의 다른 인물은 없음.",
    "hard_violations": [],
    "physics": "찰리는 침대 위에 중력에 맞게 누워있고 굵은 끈에 의해 물리적으로 자연스럽게 고정되어 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 머리는 프레임 우측 상단을 향하고 있음.",
    "built_space": "유리벽 안 실험실 중앙에 스테인리스 침대가 배치됨. 배경에 여러 연구원들이 배치되어 있음.",
    "entities": "찰리, 결박 끈, 그리고 프롬프트가 금지한 다수의 배경 연구원들.",
    "hard_violations": [
     "[gemini-pro] 프롬프트가 명시적으로 금지한 배경 인물들(연구원)이 등장함.",
     "[gpt-high] 찰리 외에는 아무도 등장하지 말라는 명시적 지시를 어기고 이전 장면의 연구원 7명을 배경에 그대로 남겼다."
    ],
    "physics": "찰리는 침대 위에 누워있고 끈으로 고정됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "배경 인물 제거 지시를 완벽히 준수하고 결박된 찰리의 모습을 프레임과 조명 설정에 맞게 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷의 배경 인물을 제거하라는 절대적인 제약 조건을 위반하고 다수의 인물을 그대로 렌더링하여 감점됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 프레임 우측 상단을 향하고 있음.",
        "built_space": "유리벽 안 실험실 중앙에 스테인리스 침대가 배치됨. 배경의 연구원석은 지시대로 모두 비어 있음.",
        "entities": "찰리(베이지색 장갑판 로봇 몸체), 굵은 결박 끈. 찰리 외의 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "찰리는 침대 위에 중력에 맞게 누워있고 굵은 끈에 의해 물리적으로 자연스럽게 고정되어 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 프레임 우측 상단을 향하고 있음.",
        "built_space": "유리벽 안 실험실 중앙에 스테인리스 침대가 배치됨. 배경에 여러 연구원들이 배치되어 있음.",
        "entities": "찰리, 결박 끈, 그리고 프롬프트가 금지한 다수의 배경 연구원들.",
        "hard_violations": [
         "프롬프트가 명시적으로 금지한 배경 인물들(연구원)이 등장함."
        ],
        "physics": "찰리는 침대 위에 누워있고 끈으로 고정됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "배경 인물 제거 지시를 완벽히 준수하고 결박된 찰리의 모습을 프레임과 조명 설정에 맞게 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷의 배경 인물을 제거하라는 절대적인 제약 조건을 위반하고 다수의 인물을 그대로 렌더링하여 감점됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 프레임 우측 상단을 향하고 있음.",
        "built_space": "유리벽 안 실험실 중앙에 스테인리스 침대가 배치됨. 배경의 연구원석은 지시대로 모두 비어 있음.",
        "entities": "찰리(베이지색 장갑판 로봇 몸체), 굵은 결박 끈. 찰리 외의 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "찰리는 침대 위에 중력에 맞게 누워있고 굵은 끈에 의해 물리적으로 자연스럽게 고정되어 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 프레임 우측 상단을 향하고 있음.",
        "built_space": "유리벽 안 실험실 중앙에 스테인리스 침대가 배치됨. 배경에 여러 연구원들이 배치되어 있음.",
        "entities": "찰리, 결박 끈, 그리고 프롬프트가 금지한 다수의 배경 연구원들.",
        "hard_violations": [
         "프롬프트가 명시적으로 금지한 배경 인물들(연구원)이 등장함."
        ],
        "physics": "찰리는 침대 위에 누워있고 끈으로 고정됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "전신 결박과 푸른 조명은 구현했지만, 등장하면 안 되는 연구원 7명을 그대로 남긴 것이 결정적인 위반이며 침대도 요구된 중경보다 전경을 크게 차지한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "찰리만 등장하고 머리가 우상단을 향하는 전신 결박 구도를 충실히 구현했으나, 침대의 전경 비중과 오른쪽 검사 장치의 위치는 장소·배치 지시와 다소 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 화면 우상단, 발은 좌하단을 향하며 얼굴은 위쪽으로 놓여 있다. 특정 인물을 응시하는 모습은 아니다. 뒤편 연구원들은 각자의 작업대와 화면을 향한다. 왼쪽 검사 장치의 끝은 침대 쪽을 향하며 무기나 이동 중인 신체는 없다.",
        "built_space": "곡면 유리 격리실 하나, 중앙 금속 침대 하나, 둘레 배수 홈, 주변 작업대와 벽면 모니터들이 보인다. 침대 왼쪽에는 받침대가 있는 검사 장치 하나가 보이고, 참고 사진의 오른쪽 장치는 커진 침대와 몸체에 가려졌는지 확인할 수 없다. 주변 작업대에 연구원 7명이 남아 있다. 침대 상면·가까운 긴 변·발치 끝은 보이지만 침대가 중경보다 전경을 크게 점유한다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 한 개체이며 흰 마스크형 얼굴, 각진 낡은 장갑판, 육중한 긴 팔과 짧은 다리를 갖춘 참고의 기계형 외형이다. 장갑판은 푸른 조명 때문에 샌드 베이지보다 청회색으로 보인다. 이전 장면처럼 모자와 외투는 없고 가슴 고리는 밝게 작동하지 않는다. 몸통·손목 부근·하퇴에 굵은 끈이 보인다. 스테인리스 침대, 유리 격리실, 감시 장비는 있으나 흰 가운 연구원 7명은 이번 장면에 허용되지 않은 인물이다.",
        "hard_violations": [
         "찰리 외에는 아무도 등장하지 말라는 명시적 지시를 어기고 이전 장면의 연구원 7명을 배경에 그대로 남겼다."
        ],
        "physics": "찰리의 몸통과 다리는 침대 상판에 놓이고, 손은 침대 가장자리와 팔의 연결부에 의해 지지된다. 발은 뒤꿈치 쪽이 상판에 닿아 있다. 결박 끈은 몸과 다리를 감싸 침대 가장자리 쪽으로 이어진다. 침대는 금속 하부 받침으로 바닥에 지지되며 검사 장치는 기둥형 받침 위에 있다. 지지 없이 떠 있는 몸이나 물체는 확인되지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 우상단, 발은 좌하단을 향해 침대의 대각선 축과 일치한다. 얼굴은 천장 쪽으로 향하고 별도의 응시 대상은 없다. 양쪽 검사 장치의 끝은 격리실 안쪽과 침대 방향을 향한다. 움직이는 인물이나 조준 중인 무기는 없다.",
        "built_space": "곡면 유리 격리실 하나와 스테인리스 침대 하나가 있으며, 둘레 배수 홈과 외곽 작업대·모니터·빈 의자가 유지된다. 왼쪽 검사 장치는 받침대까지 보이고 오른쪽 장치는 화면 가장자리에 끝부분만 보인다. 오른쪽 장치는 참고보다 외곽으로 옮겨진 듯하다. 침대 상면, 가까운 긴 변, 발치 끝과 하부 받침이 모두 보인다. 전신을 담은 넓은 구도이지만 침대는 요구된 중경 배치보다 전경 쪽 비중이 크다. 추가 인물이나 불가능한 인물 반사는 없다.",
        "entities": "찰리만 한 개체 등장한다. 흰 마스크형 얼굴과 검은 점·선, 샌드 베이지 장갑판, 큰 어깨와 긴 팔, 짧은 다리, 표면의 마모가 참고 외형과 부합한다. 이전 장면에 없던 모자나 외투를 추가하지 않았다. 가슴의 원형 장치는 어둡고 몸통·허벅지·하퇴·손목 부근의 두꺼운 결박 끈이 보인다. 금속 침대, 유리 격리실, 주변 감시 장비와 바닥 케이블이 유지된다. 판독 가능한 문구는 뚜렷하게 보이지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반, 다리는 침대에 실려 있고 손은 상판 가장자리 쪽에 내려놓여 있다. 발은 뒤꿈치로 상판에 지지되며 힘을 주어 들어 올린 자세로 보이지 않는다. 굵은 끈은 신체를 눌러 감싸고 침대 가장자리로 이어진다. 침대의 하부 받침은 바닥에 닿고 케이블은 장치에서 내려와 바닥에 늘어진다. 오른쪽 장치의 축은 화면 밖으로 이어져 전체 받침은 확인할 수 없지만, 허공에 독립적으로 떠 있는 물체로 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "전신 결박과 푸른 조명은 구현했지만, 등장하면 안 되는 연구원 7명을 그대로 남긴 것이 결정적인 위반이며 침대도 요구된 중경보다 전경을 크게 차지한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찰리만 등장하고 머리가 우상단을 향하는 전신 결박 구도를 충실히 구현했으나, 침대의 전경 비중과 오른쪽 검사 장치의 위치는 장소·배치 지시와 다소 어긋난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리는 화면 우상단, 발은 좌하단을 향하며 얼굴은 위쪽으로 놓여 있다. 특정 인물을 응시하는 모습은 아니다. 뒤편 연구원들은 각자의 작업대와 화면을 향한다. 왼쪽 검사 장치의 끝은 침대 쪽을 향하며 무기나 이동 중인 신체는 없다.",
        "built_space": "곡면 유리 격리실 하나, 중앙 금속 침대 하나, 둘레 배수 홈, 주변 작업대와 벽면 모니터들이 보인다. 침대 왼쪽에는 받침대가 있는 검사 장치 하나가 보이고, 참고 사진의 오른쪽 장치는 커진 침대와 몸체에 가려졌는지 확인할 수 없다. 주변 작업대에 연구원 7명이 남아 있다. 침대 상면·가까운 긴 변·발치 끝은 보이지만 침대가 중경보다 전경을 크게 점유한다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 한 개체이며 흰 마스크형 얼굴, 각진 낡은 장갑판, 육중한 긴 팔과 짧은 다리를 갖춘 참고의 기계형 외형이다. 장갑판은 푸른 조명 때문에 샌드 베이지보다 청회색으로 보인다. 이전 장면처럼 모자와 외투는 없고 가슴 고리는 밝게 작동하지 않는다. 몸통·손목 부근·하퇴에 굵은 끈이 보인다. 스테인리스 침대, 유리 격리실, 감시 장비는 있으나 흰 가운 연구원 7명은 이번 장면에 허용되지 않은 인물이다.",
        "hard_violations": [
         "찰리 외에는 아무도 등장하지 말라는 명시적 지시를 어기고 이전 장면의 연구원 7명을 배경에 그대로 남겼다."
        ],
        "physics": "찰리의 몸통과 다리는 침대 상판에 놓이고, 손은 침대 가장자리와 팔의 연결부에 의해 지지된다. 발은 뒤꿈치 쪽이 상판에 닿아 있다. 결박 끈은 몸과 다리를 감싸 침대 가장자리 쪽으로 이어진다. 침대는 금속 하부 받침으로 바닥에 지지되며 검사 장치는 기둥형 받침 위에 있다. 지지 없이 떠 있는 몸이나 물체는 확인되지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리의 머리는 우상단, 발은 좌하단을 향해 침대의 대각선 축과 일치한다. 얼굴은 천장 쪽으로 향하고 별도의 응시 대상은 없다. 양쪽 검사 장치의 끝은 격리실 안쪽과 침대 방향을 향한다. 움직이는 인물이나 조준 중인 무기는 없다.",
        "built_space": "곡면 유리 격리실 하나와 스테인리스 침대 하나가 있으며, 둘레 배수 홈과 외곽 작업대·모니터·빈 의자가 유지된다. 왼쪽 검사 장치는 받침대까지 보이고 오른쪽 장치는 화면 가장자리에 끝부분만 보인다. 오른쪽 장치는 참고보다 외곽으로 옮겨진 듯하다. 침대 상면, 가까운 긴 변, 발치 끝과 하부 받침이 모두 보인다. 전신을 담은 넓은 구도이지만 침대는 요구된 중경 배치보다 전경 쪽 비중이 크다. 추가 인물이나 불가능한 인물 반사는 없다.",
        "entities": "찰리만 한 개체 등장한다. 흰 마스크형 얼굴과 검은 점·선, 샌드 베이지 장갑판, 큰 어깨와 긴 팔, 짧은 다리, 표면의 마모가 참고 외형과 부합한다. 이전 장면에 없던 모자나 외투를 추가하지 않았다. 가슴의 원형 장치는 어둡고 몸통·허벅지·하퇴·손목 부근의 두꺼운 결박 끈이 보인다. 금속 침대, 유리 격리실, 주변 감시 장비와 바닥 케이블이 유지된다. 판독 가능한 문구는 뚜렷하게 보이지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반, 다리는 침대에 실려 있고 손은 상판 가장자리 쪽에 내려놓여 있다. 발은 뒤꿈치로 상판에 지지되며 힘을 주어 들어 올린 자세로 보이지 않는다. 굵은 끈은 신체를 눌러 감싸고 침대 가장자리로 이어진다. 침대의 하부 받침은 바닥에 닿고 케이블은 장치에서 내려와 바닥에 늘어진다. 오른쪽 장치의 축은 화면 밖으로 이어져 전체 받침은 확인할 수 없지만, 허공에 독립적으로 떠 있는 물체로 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트가 명시적으로 금지한 배경 인물들(연구원)이 등장함.",
     "[gpt-high] 찰리 외에는 아무도 등장하지 말라는 명시적 지시를 어기고 이전 장면의 연구원 7명을 배경에 그대로 남겼다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "배경 인물 제거 지시를 완벽히 준수하고 결박된 찰리의 모습을 프레임과 조명 설정에 맞게 정확히 구현함."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "이전 샷의 배경 인물을 제거하라는 절대적인 제약 조건을 위반하고 다수의 인물을 그대로 렌더링하여 감점됨.  ★위반: [gemini-pro] 프롬프트가 명시적으로 금지한 배경 인물들(연구원)이 등장함. / [gpt-high] 찰리 외에는 아무도 등장하지 말라는 명시적 지시를 어기고 이전 장면의 연구원 7명을 배경에 그대로 남겼다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14_sel.png",
    "asset_id": "c78997e0-2898-4ac1-b041-54b98246752d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-1c03-7c7d-8f30-d2f062e39b3c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh14"
  }
 },
 "S88sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:14:46.009892+00:00",
  "fingerprint": "c1209631d4c224358be2e59d030fad1d187034714947141e50e21bd9eac01756",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S88sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S88sh7_sel.png",
  "source_sha256": "7d080a0d4f6e8fe08484f88103c0d363f9d7478cbcf9a2fd6a2e0a5c236a9a73",
  "file": "S88sh7_cine.png",
  "staged_sha256": "aff566e2d8800a79dcb7e95e696aa05001b17f0ea30c95b7d6b0e90f0e842f5d",
  "latency_ms": 10160
 },
 "S88sh24::signage": {
  "fp": "077f014043012f7e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S88sh24": {
  "input_fingerprint": "d4680cca150e00e0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 슬픈 표정으로 눈물을 흘리는 현우의 뺨에 커다란 금속 손가락을 가만히 얹은 찰리의 다정한 상체.\n\nLOCATION (lock): At the examination bedside inside the glass-walled research laboratory, under its established cool lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (찰리 remains lying on it during the bedside exchange) — A narrow portion of the near long edge and upper surface is visible beneath his shoulder; used as Shared spatial anchor beneath the two figures without obstructing the touch.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued laboratory illumination and gentle tonal separation so the tears and metal finger remain readable without adding a new source or signaling the later alarm.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, laboratory equipment, and cool base lighting. Exclude the earlier fully secured restraint state; the hand restraints have been released.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie remains recumbent on the stainless-steel bed after his hands are released, extending one freed hand to rest a metal finger against Hyunwoo's cheek. The source does not specify his other hand's position, his legs' arrangement, or which remaining restraints are still fastened.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains lying on the stainless-steel bed, but the hand restraints have been released; no release of the remaining restraints has yet been established. The chest-ring system is still impaired following the failed experiment. 현우: He is beside the laboratory bed, crying heavily; the pole stand taken during the confrontation remains in his possession.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 슬픈 표정으로 눈물을 흘리는 현우의 뺨에 커다란 금속 손가락을 가만히 얹은 찰리의 다정한 상체.\n\nLOCATION (lock): At the examination bedside inside the glass-walled research laboratory, under its established cool lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (찰리 remains lying on it during the bedside exchange) — A narrow portion of the near long edge and upper surface is visible beneath his shoulder; used as Shared spatial anchor beneath the two figures without obstructing the touch.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued laboratory illumination and gentle tonal separation so the tears and metal finger remain readable without adding a new source or signaling the later alarm.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, laboratory equipment, and cool base lighting. Exclude the earlier fully secured restraint state; the hand restraints have been released.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie remains recumbent on the stainless-steel bed after his hands are released, extending one freed hand to rest a metal finger against Hyunwoo's cheek. The source does not specify his other hand's position, his legs' arrangement, or which remaining restraints are still fastened.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains lying on the stainless-steel bed, but the hand restraints have been released; no release of the remaining restraints has yet been established. The chest-ring system is still impaired following the failed experiment. 현우: He is beside the laboratory bed, crying heavily; the pole stand taken during the confrontation remains in his possession.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 슬픈 표정으로 눈물을 흘리는 현우의 뺨에 커다란 금속 손가락을 가만히 얹은 찰리의 다정한 상체.\n\nLOCATION (lock): At the examination bedside inside the glass-walled research laboratory, under its established cool lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (찰리 remains lying on it during the bedside exchange) — A narrow portion of the near long edge and upper surface is visible beneath his shoulder; used as Shared spatial anchor beneath the two figures without obstructing the touch.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued laboratory illumination and gentle tonal separation so the tears and metal finger remain readable without adding a new source or signaling the later alarm.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, laboratory equipment, and cool base lighting. Exclude the earlier fully secured restraint state; the hand restraints have been released.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie remains recumbent on the stainless-steel bed after his hands are released, extending one freed hand to rest a metal finger against Hyunwoo's cheek. The source does not specify his other hand's position, his legs' arrangement, or which remaining restraints are still fastened.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains lying on the stainless-steel bed, but the hand restraints have been released; no release of the remaining restraints has yet been established. The chest-ring system is still impaired following the failed experiment. 현우: He is beside the laboratory bed, crying heavily; the pole stand taken during the confrontation remains in his possession.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 금속 손가락이 현우의 뺨에 정확히 닿아 있으며, 현우는 약간 아래를 응시함.",
    "built_space": "스테인리스 침대와 유리벽이 있는 실험실 내부. 인물들의 위치와 카메라 구도가 공간에 맞게 배치됨.",
    "entities": "찰리의 외형, 현우의 눈물과 얼굴 특징은 잘 반영되었으나, 현우가 레퍼런스의 흰옷 대신 검은 옷을 입고 있음.",
    "hard_violations": [],
    "physics": "찰리는 침대에 누워 팔을 뻗은 자세가 자연스러우며, 현우는 한 손으로 폴대를 쥐고 침대에 기대어 체중을 지탱함."
   },
   {
    "label": "B",
    "direction": "찰리의 손가락이 실제 현우의 뺨이 아닌 오른쪽에 떠 있는 현우의 두 번째 얼굴을 향하고 있음.",
    "built_space": "실험실 구조물과 침대가 있으나, 카메라 위치상 물리적으로 불가능한 각도로 현우의 얼굴이 유리에 겹쳐짐.",
    "entities": "현우가 입은 흰색 의상은 맞으나, 지시되지 않은 배경 인물(연구원 2명)이 등장함.",
    "hard_violations": [
     "[gemini-pro] invented people (배경의 엑스트라 2명)",
     "[gemini-pro] duplicated or extra bodies (현우의 얼굴 중복)",
     "[gemini-pro] physically impossible staging (불가능한 반사 및 신체 분리)",
     "[gpt-high] 현우가 중앙의 정상 크기 상반신과 오른쪽의 크게 합성된 얼굴로 중복되어 나타납니다.",
     "[gpt-high] 연구실 배경에 허용되지 않은 인물 두 명이 추가되었습니다."
    ],
    "physics": "오른쪽에 위치한 현우의 두 번째 얼굴이 신체 없이 허공(또는 유리창)에 부자연스럽게 떠 있어 지지대가 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 구도와 슬픈 감정선, 로봇의 포즈를 잘 구현했으나 현우의 의상 색상이 캐릭터 레퍼런스와 일치하지 않습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 인물들이 배경에 나타났으며, 현우의 얼굴이 중복되고 기형적으로 겹쳐져 치명적인 구조적 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 금속 손가락이 현우의 뺨에 정확히 닿아 있으며, 현우는 약간 아래를 응시함.",
        "built_space": "스테인리스 침대와 유리벽이 있는 실험실 내부. 인물들의 위치와 카메라 구도가 공간에 맞게 배치됨.",
        "entities": "찰리의 외형, 현우의 눈물과 얼굴 특징은 잘 반영되었으나, 현우가 레퍼런스의 흰옷 대신 검은 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "찰리는 침대에 누워 팔을 뻗은 자세가 자연스러우며, 현우는 한 손으로 폴대를 쥐고 침대에 기대어 체중을 지탱함."
       },
       {
        "label": "B",
        "direction": "찰리의 손가락이 실제 현우의 뺨이 아닌 오른쪽에 떠 있는 현우의 두 번째 얼굴을 향하고 있음.",
        "built_space": "실험실 구조물과 침대가 있으나, 카메라 위치상 물리적으로 불가능한 각도로 현우의 얼굴이 유리에 겹쳐짐.",
        "entities": "현우가 입은 흰색 의상은 맞으나, 지시되지 않은 배경 인물(연구원 2명)이 등장함.",
        "hard_violations": [
         "invented people (배경의 엑스트라 2명)",
         "duplicated or extra bodies (현우의 얼굴 중복)",
         "physically impossible staging (불가능한 반사 및 신체 분리)"
        ],
        "physics": "오른쪽에 위치한 현우의 두 번째 얼굴이 신체 없이 허공(또는 유리창)에 부자연스럽게 떠 있어 지지대가 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 구도와 슬픈 감정선, 로봇의 포즈를 잘 구현했으나 현우의 의상 색상이 캐릭터 레퍼런스와 일치하지 않습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 인물들이 배경에 나타났으며, 현우의 얼굴이 중복되고 기형적으로 겹쳐져 치명적인 구조적 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 금속 손가락이 현우의 뺨에 정확히 닿아 있으며, 현우는 약간 아래를 응시함.",
        "built_space": "스테인리스 침대와 유리벽이 있는 실험실 내부. 인물들의 위치와 카메라 구도가 공간에 맞게 배치됨.",
        "entities": "찰리의 외형, 현우의 눈물과 얼굴 특징은 잘 반영되었으나, 현우가 레퍼런스의 흰옷 대신 검은 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "찰리는 침대에 누워 팔을 뻗은 자세가 자연스러우며, 현우는 한 손으로 폴대를 쥐고 침대에 기대어 체중을 지탱함."
       },
       {
        "label": "B",
        "direction": "찰리의 손가락이 실제 현우의 뺨이 아닌 오른쪽에 떠 있는 현우의 두 번째 얼굴을 향하고 있음.",
        "built_space": "실험실 구조물과 침대가 있으나, 카메라 위치상 물리적으로 불가능한 각도로 현우의 얼굴이 유리에 겹쳐짐.",
        "entities": "현우가 입은 흰색 의상은 맞으나, 지시되지 않은 배경 인물(연구원 2명)이 등장함.",
        "hard_violations": [
         "invented people (배경의 엑스트라 2명)",
         "duplicated or extra bodies (현우의 얼굴 중복)",
         "physically impossible staging (불가능한 반사 및 신체 분리)"
        ],
        "physics": "오른쪽에 위치한 현우의 두 번째 얼굴이 신체 없이 허공(또는 유리창)에 부자연스럽게 떠 있어 지지대가 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "현우의 중복 인물과 배경의 추가 인물로 실격이며, 찰리의 다리와 침대를 넓게 보여 주어 요청한 상체 중심 미디엄 숏도 벗어납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "침대에 누운 찰리가 우는 현우의 뺨에 손가락을 대는 미디엄 숏과 지지 관계를 충실히 구현했으나, 현우의 검은 의상은 참조의 흰 의상과 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 오른쪽 위 현우 쪽을 향하고, 뻗은 금속 검지는 오른쪽의 크게 표현된 현우 얼굴에서 코 옆 뺨에 닿습니다. 그 얼굴의 눈은 아래쪽 찰리 방향으로 내려가 있습니다. 중앙에 별도로 나타난 현우도 아래를 보며, 배경의 흰옷 인물 둘은 서로 또는 실험 장비 쪽을 향합니다.",
        "built_space": "전경에 스테인리스 침대 한 대가 있으며 상판과 긴 측면이 넓게 노출됩니다. 뒤에는 유리 칸막이, 작업대와 여러 장비가 있고 중앙 상단에는 고리가 달린 스탠드 하나, 오른쪽에는 손으로 잡은 별도 금속 봉이 보입니다. 차가운 연구실 재질은 이어지지만, 찰리의 어깨 아래 좁은 침대 부분만 공간 기준으로 쓰라는 구도보다 훨씬 넓습니다. 배경 인물 둘은 이 숏에 허용되지 않습니다.",
        "entities": "찰리는 흰 마스크형 얼굴, 베이지 장갑판, 육중한 금속 팔을 갖췄고 이전 숏처럼 모자와 코트가 없습니다. 흉부 고리는 밝게 작동하지 않으며 하체에는 구속 띠가 보입니다. 검은 머리의 젊은 동아시아계 남성이 흰옷을 입고 있으나 현우가 중앙의 상반신과 오른쪽의 거대한 얼굴로 중복됩니다. 오른쪽 얼굴에는 눈물이 있고 금속 봉을 잡은 손도 보입니다. 배경에는 별도의 흰옷 인물 두 명이 추가되었습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "현우가 중앙의 정상 크기 상반신과 오른쪽의 크게 합성된 얼굴로 중복되어 나타납니다.",
         "연구실 배경에 허용되지 않은 인물 두 명이 추가되었습니다."
        ],
        "physics": "찰리의 어깨와 몸통, 하체는 침대 상판에 받쳐져 있고, 뻗은 손은 팔꿈치와 손목 관절로 이어집니다. 손을 드는 동작 자체는 해제된 손으로 뺨을 만지라는 지정 동작에 맞습니다. 오른쪽 금속 봉에는 손의 파지가 보입니다. 다만 크게 삽입된 현우 얼굴은 중앙의 별도 몸과 연속된 단일 인체로 읽히지 않아 정상적인 두 인물의 접촉 장면이 성립하지 않습니다."
       },
       {
        "label": "B",
        "direction": "찰리는 오른쪽 현우 쪽으로 얼굴을 돌리고, 들어 올린 금속 검지 끝을 현우의 눈물 흐르는 뺨에 직접 댑니다. 현우의 시선은 왼쪽 아래 찰리의 얼굴 방향을 향합니다. 시선과 접촉 대상이 모두 지정된 상대에게 맞습니다.",
        "built_space": "스테인리스 침대 한 대의 상판과 가까운 가장자리가 찰리의 어깨 아래 왼쪽 하단에 좁게 보입니다. 현우는 침대 옆에서 상체를 가까이 기울이고 있습니다. 배경에는 유리 칸막이와 흐릿한 작업대, 장비 및 모니터들이 있으며, 오른쪽에는 현우가 잡은 금속 봉 하나가 있습니다. 상체 중심 미디엄 숏 안에서 침대가 두 인물의 공통 공간 기준으로 기능하고, 차가운 연구실 조명도 유지됩니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리와 눈물 자국, 슬픈 표정이 보이며 참조 인물의 외형에 가깝습니다. 다만 상의가 참조의 흰색이 아니라 검은색입니다. 찰리는 흰 마스크형 얼굴, 각진 샌드 베이지 장갑판과 긴 금속 팔을 유지하며 이전 숏에 없던 모자나 코트를 추가하지 않았습니다. 손목에 채워진 구속구는 보이지 않고 몸통 아래쪽에는 띠가 남아 있습니다. 흉부 고리는 발광하지 않습니다. 현우의 손에는 금속 봉이 들려 있으며 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "찰리의 어깨와 등은 침대에 기대어 지지되고, 들어 올린 팔은 어깨·팔꿈치·손목으로 자연스럽게 연결됩니다. 이는 명시된 뺨 접촉 동작이며 허공에 떠 있는 신체가 아닙니다. 금속 검지 끝은 현우의 뺨 표면에 닿고, 현우의 손은 수직 금속 봉을 실제로 감싸 잡습니다. 현우의 하체와 봉의 밑부분은 프레임 밖이므로 보이지 않는 지지 상태를 결함으로 판단할 근거는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "현우의 중복 인물과 배경의 추가 인물로 실격이며, 찰리의 다리와 침대를 넓게 보여 주어 요청한 상체 중심 미디엄 숏도 벗어납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "침대에 누운 찰리가 우는 현우의 뺨에 손가락을 대는 미디엄 숏과 지지 관계를 충실히 구현했으나, 현우의 검은 의상은 참조의 흰 의상과 다릅니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 오른쪽 위 현우 쪽을 향하고, 뻗은 금속 검지는 오른쪽의 크게 표현된 현우 얼굴에서 코 옆 뺨에 닿습니다. 그 얼굴의 눈은 아래쪽 찰리 방향으로 내려가 있습니다. 중앙에 별도로 나타난 현우도 아래를 보며, 배경의 흰옷 인물 둘은 서로 또는 실험 장비 쪽을 향합니다.",
        "built_space": "전경에 스테인리스 침대 한 대가 있으며 상판과 긴 측면이 넓게 노출됩니다. 뒤에는 유리 칸막이, 작업대와 여러 장비가 있고 중앙 상단에는 고리가 달린 스탠드 하나, 오른쪽에는 손으로 잡은 별도 금속 봉이 보입니다. 차가운 연구실 재질은 이어지지만, 찰리의 어깨 아래 좁은 침대 부분만 공간 기준으로 쓰라는 구도보다 훨씬 넓습니다. 배경 인물 둘은 이 숏에 허용되지 않습니다.",
        "entities": "찰리는 흰 마스크형 얼굴, 베이지 장갑판, 육중한 금속 팔을 갖췄고 이전 숏처럼 모자와 코트가 없습니다. 흉부 고리는 밝게 작동하지 않으며 하체에는 구속 띠가 보입니다. 검은 머리의 젊은 동아시아계 남성이 흰옷을 입고 있으나 현우가 중앙의 상반신과 오른쪽의 거대한 얼굴로 중복됩니다. 오른쪽 얼굴에는 눈물이 있고 금속 봉을 잡은 손도 보입니다. 배경에는 별도의 흰옷 인물 두 명이 추가되었습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "현우가 중앙의 정상 크기 상반신과 오른쪽의 크게 합성된 얼굴로 중복되어 나타납니다.",
         "연구실 배경에 허용되지 않은 인물 두 명이 추가되었습니다."
        ],
        "physics": "찰리의 어깨와 몸통, 하체는 침대 상판에 받쳐져 있고, 뻗은 손은 팔꿈치와 손목 관절로 이어집니다. 손을 드는 동작 자체는 해제된 손으로 뺨을 만지라는 지정 동작에 맞습니다. 오른쪽 금속 봉에는 손의 파지가 보입니다. 다만 크게 삽입된 현우 얼굴은 중앙의 별도 몸과 연속된 단일 인체로 읽히지 않아 정상적인 두 인물의 접촉 장면이 성립하지 않습니다."
       },
       {
        "label": "A",
        "direction": "찰리는 오른쪽 현우 쪽으로 얼굴을 돌리고, 들어 올린 금속 검지 끝을 현우의 눈물 흐르는 뺨에 직접 댑니다. 현우의 시선은 왼쪽 아래 찰리의 얼굴 방향을 향합니다. 시선과 접촉 대상이 모두 지정된 상대에게 맞습니다.",
        "built_space": "스테인리스 침대 한 대의 상판과 가까운 가장자리가 찰리의 어깨 아래 왼쪽 하단에 좁게 보입니다. 현우는 침대 옆에서 상체를 가까이 기울이고 있습니다. 배경에는 유리 칸막이와 흐릿한 작업대, 장비 및 모니터들이 있으며, 오른쪽에는 현우가 잡은 금속 봉 하나가 있습니다. 상체 중심 미디엄 숏 안에서 침대가 두 인물의 공통 공간 기준으로 기능하고, 차가운 연구실 조명도 유지됩니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리와 눈물 자국, 슬픈 표정이 보이며 참조 인물의 외형에 가깝습니다. 다만 상의가 참조의 흰색이 아니라 검은색입니다. 찰리는 흰 마스크형 얼굴, 각진 샌드 베이지 장갑판과 긴 금속 팔을 유지하며 이전 숏에 없던 모자나 코트를 추가하지 않았습니다. 손목에 채워진 구속구는 보이지 않고 몸통 아래쪽에는 띠가 남아 있습니다. 흉부 고리는 발광하지 않습니다. 현우의 손에는 금속 봉이 들려 있으며 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "찰리의 어깨와 등은 침대에 기대어 지지되고, 들어 올린 팔은 어깨·팔꿈치·손목으로 자연스럽게 연결됩니다. 이는 명시된 뺨 접촉 동작이며 허공에 떠 있는 신체가 아닙니다. 금속 검지 끝은 현우의 뺨 표면에 닿고, 현우의 손은 수직 금속 봉을 실제로 감싸 잡습니다. 현우의 하체와 봉의 밑부분은 프레임 밖이므로 보이지 않는 지지 상태를 결함으로 판단할 근거는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.554
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.304
   },
   "violations": {
    "B": [
     "[gemini-pro] invented people (배경의 엑스트라 2명)",
     "[gemini-pro] duplicated or extra bodies (현우의 얼굴 중복)",
     "[gemini-pro] physically impossible staging (불가능한 반사 및 신체 분리)",
     "[gpt-high] 현우가 중앙의 정상 크기 상반신과 오른쪽의 크게 합성된 얼굴로 중복되어 나타납니다.",
     "[gpt-high] 연구실 배경에 허용되지 않은 인물 두 명이 추가되었습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 304
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 구도와 슬픈 감정선, 로봇의 포즈를 잘 구현했으나 현우의 의상 색상이 캐릭터 레퍼런스와 일치하지 않습니다."
   },
   {
    "label": "B",
    "score": 304,
    "verdict_ko": "프롬프트에 없는 인물들이 배경에 나타났으며, 현우의 얼굴이 중복되고 기형적으로 겹쳐져 치명적인 구조적 오류를 범했습니다.  ★위반: [gemini-pro] invented people (배경의 엑스트라 2명) / [gemini-pro] duplicated or extra bodies (현우의 얼굴 중복) / [gemini-pro] physically impossible staging (불가능한 반사 및 신체 분리) / [gpt-high] 현우가 중앙의 정상 크기 상반신과 오른쪽의 크게 합성된 얼굴로 중복되어 나타납니다. / [gpt-high] 연구실 배경에 허용되지 않은 인물 두 명이 추가되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S88sh7_sel.png",
    "asset_id": "9eb7a825-7736-410c-b47a-e74318934714",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-1da7-7a31-927f-074d0d021a5c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S88sh7"
  }
 },
 "S88sh24::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:16:11.669406+00:00",
  "fingerprint": "08612094e655a5862499c0d8d8edfce3121fb2475cf52b09aaa1411015cb7f8a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S88sh24_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S88sh24_sel.png",
  "source_sha256": "da6692e6ae88397a49e23c8c10011710b17e5befe65242fb6b15134f1879cae2",
  "file": "S88sh24_cine.png",
  "staged_sha256": "a804b50e8fb1b9dc66be7551c71779f1cea98a9d6b753ec015cbb708c67c65e7",
  "latency_ms": 9493
 },
 "S88sh29::signage": {
  "fp": "5dcae9eff65f5a79",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S88sh29": {
  "input_fingerprint": "bb74ebffab0806c2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 갑작스러운 붉은 사이렌 불빛 아래서 당황하여 주변의 한 곳으로 시선을 홱 돌린 채 굳어버린 지소영과 현우의 상체.\n\nLOCATION (lock): Inside the glass-walled laboratory's bedside and observation area, suddenly washed in red emergency warning light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Present beside the interrupted exchange) — A narrow portion of the stainless-steel side runs diagonally along the lower edge; used as Maintains the established bedside axis without including Charlie in this reaction frame; 지소영's chair (Occupied) — Only the portion supporting her seated body is visible behind her; used as Makes her lower position and interrupted seated movement legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Sudden red alarm light interrupts the existing laboratory illumination while preserving readable expressions and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, and fixed laboratory equipment. Exclude the earlier unalarmed lighting state and fully secured hand restraints; apply the newly activated red emergency illumination.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The laboratory is now under a red alert. Charlie remains on the stainless-steel bed with the hands released but the remaining restraints not yet removed. 현우: He remains tearful beside the laboratory bed, still in possession of the stand taken during the confrontation. 지소영: She remains by the chair in her research coat, abruptly turning her attention to the alarm.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 갑작스러운 붉은 사이렌 불빛 아래서 당황하여 주변의 한 곳으로 시선을 홱 돌린 채 굳어버린 지소영과 현우의 상체.\n\nLOCATION (lock): Inside the glass-walled laboratory's bedside and observation area, suddenly washed in red emergency warning light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Present beside the interrupted exchange) — A narrow portion of the stainless-steel side runs diagonally along the lower edge; used as Maintains the established bedside axis without including Charlie in this reaction frame; 지소영's chair (Occupied) — Only the portion supporting her seated body is visible behind her; used as Makes her lower position and interrupted seated movement legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Sudden red alarm light interrupts the existing laboratory illumination while preserving readable expressions and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, and fixed laboratory equipment. Exclude the earlier unalarmed lighting state and fully secured hand restraints; apply the newly activated red emergency illumination.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The laboratory is now under a red alert. Charlie remains on the stainless-steel bed with the hands released but the remaining restraints not yet removed. 현우: He remains tearful beside the laboratory bed, still in possession of the stand taken during the confrontation. 지소영: She remains by the chair in her research coat, abruptly turning her attention to the alarm.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 갑작스러운 붉은 사이렌 불빛 아래서 당황하여 주변의 한 곳으로 시선을 홱 돌린 채 굳어버린 지소영과 현우의 상체.\n\nLOCATION (lock): Inside the glass-walled laboratory's bedside and observation area, suddenly washed in red emergency warning light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Present beside the interrupted exchange) — A narrow portion of the stainless-steel side runs diagonally along the lower edge; used as Maintains the established bedside axis without including Charlie in this reaction frame; 지소영's chair (Occupied) — Only the portion supporting her seated body is visible behind her; used as Makes her lower position and interrupted seated movement legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Sudden red alarm light interrupts the existing laboratory illumination while preserving readable expressions and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, and fixed laboratory equipment. Exclude the earlier unalarmed lighting state and fully secured hand restraints; apply the newly activated red emergency illumination.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The laboratory is now under a red alert. Charlie remains on the stainless-steel bed with the hands released but the remaining restraints not yet removed. 현우: He remains tearful beside the laboratory bed, still in possession of the stand taken during the confrontation. 지소영: She remains by the chair in her research coat, abruptly turning her attention to the alarm.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물 모두 당황한 표정으로 프레임 밖 왼쪽을 강하게 응시함.",
    "built_space": "화면 하단에 침대 모서리가 대각선으로 위치하고 지소영 뒤에 의자가 보임. 붉은 조명이 공간 전체를 덮음.",
    "entities": "현우는 눈물을 흘리며 스탠드를 잡고 있고, 지소영은 레퍼런스와 동일한 특수 자켓을 착용함. 찰리는 화면에 없음.",
    "hard_violations": [],
    "physics": "지소영은 손으로 바닥을 짚고 몸을 일으키려는 자세이며, 현우는 금속 스탠드에 의지해 안정적으로 서 있음."
   },
   {
    "label": "B",
    "direction": "두 인물 모두 화면 밖 왼쪽을 바라보고 있음.",
    "built_space": "우측 하단에 침대가 대각선으로 놓여있고 붉은 비상조명이 켜져 있음.",
    "entities": "현우는 레퍼런스와 유사하나 지소영은 지정된 의상이 아닌 일반 가운을 입음. 찰리가 배경에 서서 등장함.",
    "hard_violations": [
     "[gemini-pro] 화면에서 제외되어야 할 찰리가 포함됨",
     "[gemini-pro] 침대에 누워있어야 할 찰리가 임의로 기립한 상태로 변경됨",
     "[gpt-high] 두 사람만 보여야 하는 반응 장면에 제3의 로봇형 인물을 추가했다. 이를 찰리로 보더라도 침대에 남아 화면에서 제외되어야 할 인물을 배경에 세운 것이므로 인원 및 위치 고정을 위반한다."
    ],
    "physics": "지소영은 의자를 짚고 앉아있고 현우는 몸을 굽힌 채 두 발로 지탱하며 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프레임 구성, 붉은 조명, 인물 외형을 완벽히 재현했으며 찰리를 화면에서 제외하라는 지시를 정확히 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면에서 제외되어야 할 찰리가 배경에 서서 등장하며, 지소영의 의상이 레퍼런스와 다름."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 당황한 표정으로 프레임 밖 왼쪽을 강하게 응시함.",
        "built_space": "화면 하단에 침대 모서리가 대각선으로 위치하고 지소영 뒤에 의자가 보임. 붉은 조명이 공간 전체를 덮음.",
        "entities": "현우는 눈물을 흘리며 스탠드를 잡고 있고, 지소영은 레퍼런스와 동일한 특수 자켓을 착용함. 찰리는 화면에 없음.",
        "hard_violations": [],
        "physics": "지소영은 손으로 바닥을 짚고 몸을 일으키려는 자세이며, 현우는 금속 스탠드에 의지해 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 화면 밖 왼쪽을 바라보고 있음.",
        "built_space": "우측 하단에 침대가 대각선으로 놓여있고 붉은 비상조명이 켜져 있음.",
        "entities": "현우는 레퍼런스와 유사하나 지소영은 지정된 의상이 아닌 일반 가운을 입음. 찰리가 배경에 서서 등장함.",
        "hard_violations": [
         "화면에서 제외되어야 할 찰리가 포함됨",
         "침대에 누워있어야 할 찰리가 임의로 기립한 상태로 변경됨"
        ],
        "physics": "지소영은 의자를 짚고 앉아있고 현우는 몸을 굽힌 채 두 발로 지탱하며 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프레임 구성, 붉은 조명, 인물 외형을 완벽히 재현했으며 찰리를 화면에서 제외하라는 지시를 정확히 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면에서 제외되어야 할 찰리가 배경에 서서 등장하며, 지소영의 의상이 레퍼런스와 다름."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 당황한 표정으로 프레임 밖 왼쪽을 강하게 응시함.",
        "built_space": "화면 하단에 침대 모서리가 대각선으로 위치하고 지소영 뒤에 의자가 보임. 붉은 조명이 공간 전체를 덮음.",
        "entities": "현우는 눈물을 흘리며 스탠드를 잡고 있고, 지소영은 레퍼런스와 동일한 특수 자켓을 착용함. 찰리는 화면에 없음.",
        "hard_violations": [],
        "physics": "지소영은 손으로 바닥을 짚고 몸을 일으키려는 자세이며, 현우는 금속 스탠드에 의지해 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 화면 밖 왼쪽을 바라보고 있음.",
        "built_space": "우측 하단에 침대가 대각선으로 놓여있고 붉은 비상조명이 켜져 있음.",
        "entities": "현우는 레퍼런스와 유사하나 지소영은 지정된 의상이 아닌 일반 가운을 입음. 찰리가 배경에 서서 등장함.",
        "hard_violations": [
         "화면에서 제외되어야 할 찰리가 포함됨",
         "침대에 누워있어야 할 찰리가 임의로 기립한 상태로 변경됨"
        ],
        "physics": "지소영은 의자를 짚고 앉아있고 현우는 몸을 굽힌 채 두 발로 지탱하며 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 제3의 로봇형 인물을 배경에 세웠으며, 침대가 지나치게 크게 노출되고 현우의 스탠드 소지 상태도 이어지지 않는다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 사람만의 상체 반응, 같은 방향으로 돌아간 시선, 착석한 지소영과 스탠드를 쥔 현우를 충실히 구현하지만 침대 상판 노출이 지정된 좁은 가장자리보다 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 몸을 비튼 채 화면 왼쪽 밖을 보고, 현우도 왼쪽 전방의 화면 밖을 바라본다. 경보를 알아차린 반응은 보이지만 두 시선이 정확히 같은 지점에 모이는지는 확인하기 어렵다. 뒤의 로봇형 인물은 카메라 쪽을 향한다.",
        "built_space": "유리 칸막이와 금속 프레임, 오른쪽 작업대의 화면 네 개, 왼쪽과 중앙의 장비 카트가 보인다. 지소영은 왼쪽 의자 한 개에 낮게 앉아 있고 현우는 오른쪽 침대 옆에 있다. 뒤쪽에는 별도의 작업 의자 한 개가 보인다. 침대 한 개의 넓은 상판이 오른쪽 아래를 크게 차지해, 하단에 좁은 금속 측면만 보이라는 구도와 다르다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "지소영은 정돈된 검은 단발의 중년 동아시아계 여성으로 흰 연구복을 입었지만, 안쪽 회색 셔츠는 참조의 높은 칼라 의상과 다르다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 검은 상의, 눈물 자국이 이전 장면과 부합한다. 보이는 양손은 스탠드를 잡고 있지 않으며 스탠드도 확인되지 않는다. 두 사람 외에 흰 장갑의 인간형 로봇이 배경에 추가되어 있다.",
        "hard_violations": [
         "두 사람만 보여야 하는 반응 장면에 제3의 로봇형 인물을 추가했다. 이를 찰리로 보더라도 침대에 남아 화면에서 제외되어야 할 인물을 배경에 세운 것이므로 인원 및 위치 고정을 위반한다."
        ],
        "physics": "지소영은 의자에 골반을 둔 채 손으로 팔걸이를 잡아 몸을 돌리고 있어 지지가 자연스럽다. 현우는 몸을 앞으로 기울였으며 하체와 발은 프레임 밖이지만 부유하는 자세는 아니다. 침대와 장비는 통상적인 바닥 설치 구조로 보인다. 배경 로봇의 발은 가려져 있으므로 부유한다고 단정할 수는 없지만, 그 존재와 배치는 장면 조건에 어긋난다."
       },
       {
        "label": "B",
        "direction": "지소영과 현우 모두 화면 왼쪽 밖의 경보가 난 쪽을 바라본다. 지소영은 몸통보다 머리를 더 왼쪽으로 돌렸고 머리카락도 옆으로 흩어져 급히 돌아본 순간이 읽힌다. 현우는 눈물을 흘린 채 입을 조금 벌리고 같은 쪽을 응시한다. 실제 경보 대상은 화면 밖이다.",
        "built_space": "유리 칸막이와 수직 금속 프레임, 뒤쪽 장비 작업대, 식별 가능한 화면 세 개와 오른쪽 블라인드가 보인다. 지소영 뒤에는 그녀를 받치는 의자 한 개의 등받이 일부만 노출된다. 현우는 오른쪽 침대 옆에서 스탠드 한 개를 쥐고 있다. 침대 한 개가 하단을 대각선으로 지나가지만 좁은 측면뿐 아니라 넓은 상판까지 드러난다. 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "등장인물은 지소영과 현우 두 명뿐이며 찰리의 신체는 보이지 않는다. 지소영은 중년 동아시아계 여성으로 참조와 유사한 얼굴, 검은 단발, 높은 칼라의 밝은 의상 위에 연구복을 입었다. 현우는 참조와 유사한 앳된 얼굴과 헝클어진 검은 머리, 이전 장면의 검은 상의 및 눈물 자국을 유지한다. 금속 스탠드도 손에 들려 있다. 붉은 경보광은 뚜렷하지만 기존 실험실 조명보다 다소 강하게 화면 전체를 물들인다. 읽을 수 있는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "지소영은 등 뒤의 의자에 골반을 둔 낮은 자세에서 상체를 비틀고 있어 착석 지지가 성립한다. 현우의 손가락은 수직 스탠드 봉을 실제로 감싸 쥐고 있으며, 봉은 하단 프레임 밖으로 이어진다. 두 사람의 하체는 가려져 있지만 지지 없이 떠 있는 몸이나 물체는 보이지 않는다. 머리카락의 작은 움직임도 급히 고개를 돌린 행동으로 설명 가능하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 제3의 로봇형 인물을 배경에 세웠으며, 침대가 지나치게 크게 노출되고 현우의 스탠드 소지 상태도 이어지지 않는다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 사람만의 상체 반응, 같은 방향으로 돌아간 시선, 착석한 지소영과 스탠드를 쥔 현우를 충실히 구현하지만 침대 상판 노출이 지정된 좁은 가장자리보다 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "지소영은 몸을 비튼 채 화면 왼쪽 밖을 보고, 현우도 왼쪽 전방의 화면 밖을 바라본다. 경보를 알아차린 반응은 보이지만 두 시선이 정확히 같은 지점에 모이는지는 확인하기 어렵다. 뒤의 로봇형 인물은 카메라 쪽을 향한다.",
        "built_space": "유리 칸막이와 금속 프레임, 오른쪽 작업대의 화면 네 개, 왼쪽과 중앙의 장비 카트가 보인다. 지소영은 왼쪽 의자 한 개에 낮게 앉아 있고 현우는 오른쪽 침대 옆에 있다. 뒤쪽에는 별도의 작업 의자 한 개가 보인다. 침대 한 개의 넓은 상판이 오른쪽 아래를 크게 차지해, 하단에 좁은 금속 측면만 보이라는 구도와 다르다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "지소영은 정돈된 검은 단발의 중년 동아시아계 여성으로 흰 연구복을 입었지만, 안쪽 회색 셔츠는 참조의 높은 칼라 의상과 다르다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 검은 상의, 눈물 자국이 이전 장면과 부합한다. 보이는 양손은 스탠드를 잡고 있지 않으며 스탠드도 확인되지 않는다. 두 사람 외에 흰 장갑의 인간형 로봇이 배경에 추가되어 있다.",
        "hard_violations": [
         "두 사람만 보여야 하는 반응 장면에 제3의 로봇형 인물을 추가했다. 이를 찰리로 보더라도 침대에 남아 화면에서 제외되어야 할 인물을 배경에 세운 것이므로 인원 및 위치 고정을 위반한다."
        ],
        "physics": "지소영은 의자에 골반을 둔 채 손으로 팔걸이를 잡아 몸을 돌리고 있어 지지가 자연스럽다. 현우는 몸을 앞으로 기울였으며 하체와 발은 프레임 밖이지만 부유하는 자세는 아니다. 침대와 장비는 통상적인 바닥 설치 구조로 보인다. 배경 로봇의 발은 가려져 있으므로 부유한다고 단정할 수는 없지만, 그 존재와 배치는 장면 조건에 어긋난다."
       },
       {
        "label": "A",
        "direction": "지소영과 현우 모두 화면 왼쪽 밖의 경보가 난 쪽을 바라본다. 지소영은 몸통보다 머리를 더 왼쪽으로 돌렸고 머리카락도 옆으로 흩어져 급히 돌아본 순간이 읽힌다. 현우는 눈물을 흘린 채 입을 조금 벌리고 같은 쪽을 응시한다. 실제 경보 대상은 화면 밖이다.",
        "built_space": "유리 칸막이와 수직 금속 프레임, 뒤쪽 장비 작업대, 식별 가능한 화면 세 개와 오른쪽 블라인드가 보인다. 지소영 뒤에는 그녀를 받치는 의자 한 개의 등받이 일부만 노출된다. 현우는 오른쪽 침대 옆에서 스탠드 한 개를 쥐고 있다. 침대 한 개가 하단을 대각선으로 지나가지만 좁은 측면뿐 아니라 넓은 상판까지 드러난다. 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "등장인물은 지소영과 현우 두 명뿐이며 찰리의 신체는 보이지 않는다. 지소영은 중년 동아시아계 여성으로 참조와 유사한 얼굴, 검은 단발, 높은 칼라의 밝은 의상 위에 연구복을 입었다. 현우는 참조와 유사한 앳된 얼굴과 헝클어진 검은 머리, 이전 장면의 검은 상의 및 눈물 자국을 유지한다. 금속 스탠드도 손에 들려 있다. 붉은 경보광은 뚜렷하지만 기존 실험실 조명보다 다소 강하게 화면 전체를 물들인다. 읽을 수 있는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "지소영은 등 뒤의 의자에 골반을 둔 낮은 자세에서 상체를 비틀고 있어 착석 지지가 성립한다. 현우의 손가락은 수직 스탠드 봉을 실제로 감싸 쥐고 있으며, 봉은 하단 프레임 밖으로 이어진다. 두 사람의 하체는 가려져 있지만 지지 없이 떠 있는 몸이나 물체는 보이지 않는다. 머리카락의 작은 움직임도 급히 고개를 돌린 행동으로 설명 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면에서 제외되어야 할 찰리가 포함됨",
     "[gemini-pro] 침대에 누워있어야 할 찰리가 임의로 기립한 상태로 변경됨",
     "[gpt-high] 두 사람만 보여야 하는 반응 장면에 제3의 로봇형 인물을 추가했다. 이를 찰리로 보더라도 침대에 남아 화면에서 제외되어야 할 인물을 배경에 세운 것이므로 인원 및 위치 고정을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프레임 구성, 붉은 조명, 인물 외형을 완벽히 재현했으며 찰리를 화면에서 제외하라는 지시를 정확히 따름."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "화면에서 제외되어야 할 찰리가 배경에 서서 등장하며, 지소영의 의상이 레퍼런스와 다름.  ★위반: [gemini-pro] 화면에서 제외되어야 할 찰리가 포함됨 / [gemini-pro] 침대에 누워있어야 할 찰리가 임의로 기립한 상태로 변경됨 / [gpt-high] 두 사람만 보여야 하는 반응 장면에 제3의 로봇형 인물을 추가했다. 이를 찰리로 보더라도 침대에 남아 화면에서 제외되어야 할 인물을 배경에 세운 것이므로 인원 및 위치 고정을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S88sh24_sel.png",
    "asset_id": "c8f78a99-642e-4ac5-95f3-e00e4f92384f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:826250>",
    "asset_id": "40e6bc8e-8b09-437a-b2fe-233de2e35fc1",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:703308>",
    "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-1f4c-7dfa-8b21-251f2573ff53",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S88sh24"
  }
 },
 "S88sh29::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:15:16.357589+00:00",
  "fingerprint": "d54a46a66ecd49604a55cb462271d9e222f27fc0de36839ee527d44baabcc825",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S88sh29_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S88sh29_sel.png",
  "source_sha256": "092840ad3abeac1ecc444ddcf7473d6832dcc35810599c108b1a2d2d4001984c",
  "file": "S88sh29_cine.png",
  "staged_sha256": "acfa082a9fe7ebbcb698e8c9b0b5f7a775c7e21c697b0915d6ccf878b077b6b1",
  "latency_ms": 11014
 },
 "S89sh42::signage": {
  "fp": "6f15d9bcb8905787",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9b5ca709d4a5ecc0": {
  "subjects": [],
  "subject_text": "제주도 연구소 옥상과 안테나 구역\n잔디밭을 갖춘 옥상 공원. 거대한 안테나와 사방을 향하는 스피커가 설치된 개방형 공간이다.",
  "identity": "canonical",
  "scope_id": "L278",
  "scope_role": "location_exterior",
  "scope_sha": "0f3607503a5f57e9"
 },
 "S89sh42::bgfirst_bg": {
  "input_fingerprint": "5101d39a2219e53a",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42__bgfirst_bg.png",
  "asset_id": "99219a2e-c866-4ccd-9da5-2b903d22263a",
  "input_asset_ids": [
   "2a41fdeb-7ce7-4a06-960e-c2b1ec36de8b",
   "6eab84a7-fe74-47c0-bd42-78a5d724b6a9",
   "ae92432a-4956-46ca-9845-8c2d71d637c5"
  ]
 },
 "S89sh42": {
  "input_fingerprint": "1eeaba1a83931203",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rooftop garden retains its large antenna and all-direction speakers, now amid bombardment damage and collapsed combat robots. A Yubik command helicopter has landed on the roof while other armed aircraft surround the facility; Charlie is free of the bed restraints and still carries the experimental sensors. 현우: He is on the rooftop with a wireless receiver attached to his body and fresh shrapnel wounds from the aerial attack. 윤성찬: He is disembarking from the command helicopter on the rooftop, smiling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rooftop garden retains its large antenna and all-direction speakers, now amid bombardment damage and collapsed combat robots. A Yubik command helicopter has landed on the roof while other armed aircraft surround the facility; Charlie is free of the bed restraints and still carries the experimental sensors. 현우: He is on the rooftop with a wireless receiver attached to his body and fresh shrapnel wounds from the aerial attack. 윤성찬: He is disembarking from the command helicopter on the rooftop, smiling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rooftop garden retains its large antenna and all-direction speakers, now amid bombardment damage and collapsed combat robots. A Yubik command helicopter has landed on the roof while other armed aircraft surround the facility; Charlie is free of the bed restraints and still carries the experimental sensors. 현우: He is on the rooftop with a wireless receiver attached to his body and fresh shrapnel wounds from the aerial attack. 윤성찬: He is disembarking from the command helicopter on the rooftop, smiling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42__bgfirst_bg.png",
     "asset_id": "99219a2e-c866-4ccd-9da5-2b903d22263a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S89sh42.png",
     "asset_id": "2a41fdeb-7ce7-4a06-960e-c2b1ec36de8b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:886743>",
     "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1337888>",
     "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B01.png",
     "asset_id": "6eab84a7-fe74-47c0-bd42-78a5d724b6a9",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_island_research_complex_sel.png",
     "asset_id": "ae92432a-4956-46ca-9845-8c2d71d637c5",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:886743>",
     "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1337888>",
     "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "윤성찬은 헬기 스텝을 딛고 아래의 현우를 향해 시선을 던지고 있으며, 현우 역시 위를 올려다보며 윤성찬과 시선을 정확히 맞추고 있다.",
    "built_space": "레퍼런스에 주어진 옥상 공간(콘크리트 바닥, 잔디, 통신 안테나, 외벽 문)이 정확한 위치와 비례로 구현되었으며, 헬기의 배치 또한 지시된 바와 같이 우측에 적절히 자리잡고 있다.",
    "entities": "윤성찬은 참조 이미지의 네이비 오버코트가 누락되고 핀스트라이프 정장만 입었으나 얼굴의 노화 정도와 표정은 지시문과 일치한다. 현우는 어깨의 찢어진 상처, 붕대, 환자복의 로고(JR), 허리의 무선 수신기까지 완벽히 일치한다. 프롬프트에 명시된 파괴된 로봇과 주변의 무장 헬기들까지 충실하게 묘사되었다.",
    "hard_violations": [
     "[gpt-high] 현우의 가슴에 판독 가능한 파란색 글자 로고가 노출되어, 읽을 수 있는 문자와 로고를 모두 금지한 조건을 위반한다."
    ],
    "physics": "윤성찬은 한 발을 헬기 스텝에, 다른 한 발을 바닥에 디딘 상태로 체중이 자연스럽게 분산되어 있으며, 현우는 바닥에 웅크려 안정적으로 지지받고 있다. 모든 인물과 사물의 물리적 접촉이 타당하다."
   },
   {
    "label": "B",
    "direction": "윤성찬은 헬기 스텝 위에서 현우를 내려다보고 있고, 현우는 선 채로 고개를 들어 윤성찬을 쳐다보고 있다.",
    "built_space": "옥상의 잔디, 안테나, 건물의 문과 벽면 등 레퍼런스의 공간적 특징이 올바르게 반영되었다.",
    "entities": "윤성찬은 오버코트를 포함한 복장과 외모가 참조 이미지와 잘 일치한다. 현우는 피를 흘리며 무선 수신기를 착용하고 있으나, 옷의 디테일(로고 등)이 부족하다. 옥상의 파괴된 잔해들이 일부 있으나 전투 로봇의 형태가 불분명하다.",
    "hard_violations": [],
    "physics": "윤성찬과 현우 모두 헬기 스텝과 바닥에 정상적으로 지지된 채 서 있으나, 윤성찬이 스텝에 두 발을 모두 올리고 있어 '한 발을 딛고'라는 프롬프트의 물리적 동작 지시를 따르지 않았다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "윤성찬의 오버코트가 누락되었으나, '한 발을 딛고 선 채'라는 핵심 동작을 완벽하게 구현하였고 주변 항공기와 파괴된 로봇 등 디테일의 재현도가 매우 뛰어남."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "윤성찬의 복장은 참조와 일치하나, 지시된 '한 발을 딛고' 있는 포즈를 무시하고 두 발로 스텝에 올라서 있으며 주변 환경 묘사가 상대적으로 단조로움."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬은 헬기 스텝을 딛고 아래의 현우를 향해 시선을 던지고 있으며, 현우 역시 위를 올려다보며 윤성찬과 시선을 정확히 맞추고 있다.",
        "built_space": "레퍼런스에 주어진 옥상 공간(콘크리트 바닥, 잔디, 통신 안테나, 외벽 문)이 정확한 위치와 비례로 구현되었으며, 헬기의 배치 또한 지시된 바와 같이 우측에 적절히 자리잡고 있다.",
        "entities": "윤성찬은 참조 이미지의 네이비 오버코트가 누락되고 핀스트라이프 정장만 입었으나 얼굴의 노화 정도와 표정은 지시문과 일치한다. 현우는 어깨의 찢어진 상처, 붕대, 환자복의 로고(JR), 허리의 무선 수신기까지 완벽히 일치한다. 프롬프트에 명시된 파괴된 로봇과 주변의 무장 헬기들까지 충실하게 묘사되었다.",
        "hard_violations": [],
        "physics": "윤성찬은 한 발을 헬기 스텝에, 다른 한 발을 바닥에 디딘 상태로 체중이 자연스럽게 분산되어 있으며, 현우는 바닥에 웅크려 안정적으로 지지받고 있다. 모든 인물과 사물의 물리적 접촉이 타당하다."
       },
       {
        "label": "B",
        "direction": "윤성찬은 헬기 스텝 위에서 현우를 내려다보고 있고, 현우는 선 채로 고개를 들어 윤성찬을 쳐다보고 있다.",
        "built_space": "옥상의 잔디, 안테나, 건물의 문과 벽면 등 레퍼런스의 공간적 특징이 올바르게 반영되었다.",
        "entities": "윤성찬은 오버코트를 포함한 복장과 외모가 참조 이미지와 잘 일치한다. 현우는 피를 흘리며 무선 수신기를 착용하고 있으나, 옷의 디테일(로고 등)이 부족하다. 옥상의 파괴된 잔해들이 일부 있으나 전투 로봇의 형태가 불분명하다.",
        "hard_violations": [],
        "physics": "윤성찬과 현우 모두 헬기 스텝과 바닥에 정상적으로 지지된 채 서 있으나, 윤성찬이 스텝에 두 발을 모두 올리고 있어 '한 발을 딛고'라는 프롬프트의 물리적 동작 지시를 따르지 않았다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "윤성찬의 오버코트가 누락되었으나, '한 발을 딛고 선 채'라는 핵심 동작을 완벽하게 구현하였고 주변 항공기와 파괴된 로봇 등 디테일의 재현도가 매우 뛰어남."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "윤성찬의 복장은 참조와 일치하나, 지시된 '한 발을 딛고' 있는 포즈를 무시하고 두 발로 스텝에 올라서 있으며 주변 환경 묘사가 상대적으로 단조로움."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬은 헬기 스텝을 딛고 아래의 현우를 향해 시선을 던지고 있으며, 현우 역시 위를 올려다보며 윤성찬과 시선을 정확히 맞추고 있다.",
        "built_space": "레퍼런스에 주어진 옥상 공간(콘크리트 바닥, 잔디, 통신 안테나, 외벽 문)이 정확한 위치와 비례로 구현되었으며, 헬기의 배치 또한 지시된 바와 같이 우측에 적절히 자리잡고 있다.",
        "entities": "윤성찬은 참조 이미지의 네이비 오버코트가 누락되고 핀스트라이프 정장만 입었으나 얼굴의 노화 정도와 표정은 지시문과 일치한다. 현우는 어깨의 찢어진 상처, 붕대, 환자복의 로고(JR), 허리의 무선 수신기까지 완벽히 일치한다. 프롬프트에 명시된 파괴된 로봇과 주변의 무장 헬기들까지 충실하게 묘사되었다.",
        "hard_violations": [],
        "physics": "윤성찬은 한 발을 헬기 스텝에, 다른 한 발을 바닥에 디딘 상태로 체중이 자연스럽게 분산되어 있으며, 현우는 바닥에 웅크려 안정적으로 지지받고 있다. 모든 인물과 사물의 물리적 접촉이 타당하다."
       },
       {
        "label": "B",
        "direction": "윤성찬은 헬기 스텝 위에서 현우를 내려다보고 있고, 현우는 선 채로 고개를 들어 윤성찬을 쳐다보고 있다.",
        "built_space": "옥상의 잔디, 안테나, 건물의 문과 벽면 등 레퍼런스의 공간적 특징이 올바르게 반영되었다.",
        "entities": "윤성찬은 오버코트를 포함한 복장과 외모가 참조 이미지와 잘 일치한다. 현우는 피를 흘리며 무선 수신기를 착용하고 있으나, 옷의 디테일(로고 등)이 부족하다. 옥상의 파괴된 잔해들이 일부 있으나 전투 로봇의 형태가 불분명하다.",
        "hard_violations": [],
        "physics": "윤성찬과 현우 모두 헬기 스텝과 바닥에 정상적으로 지지된 채 서 있으나, 윤성찬이 스텝에 두 발을 모두 올리고 있어 '한 발을 딛고'라는 프롬프트의 물리적 동작 지시를 따르지 않았다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "윤성찬의 전신과 현우를 내려다보는 시선은 맞지만, 두 발을 발판에 모은 자세와 구조 참조에 맞지 않는 안테나가 아쉽다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "한 발을 헬기 스텝에 딛는 동작과 두 사람의 고저차는 더 정확하지만, 현우 가슴의 판독 가능한 파란 글자 로고가 명시적인 문자·로고 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬은 고개와 눈을 왼쪽 아래의 현우에게 향하고 웃는다. 현우는 오른쪽 위의 윤성찬 얼굴을 올려다본다. 두 사람의 시선 대상은 맞으며, 조준하는 무기나 손에 든 지시 물체는 없다.",
        "built_space": "왼쪽 옥탑 벽에 문 하나, 문 환기구 하나, 벽등 하나가 보인다. 중앙에는 원형 접시 하나와 좌우 혼 스피커 두 개가 달린 철탑이 있다. 잔디와 파손된 포장면이 전경의 현우에서 오른쪽 헬기 발판까지 이어지며, 헬기 대부분은 오른쪽 밖으로 잘렸다. 다만 중앙 설비는 장소 사진을 따랐고, 우선권이 있는 구조 사진의 3단 사각 스피커 배열 및 주변 독립 혼 스피커 배치를 재현하지 않았다.",
        "entities": "사람은 두 명뿐이다. 윤성찬은 회색 머리, 안경, 주름진 고령의 동아시아계 남성으로 보이며 짙은 긴 코트와 정장이 참조에 가깝다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로, 찢어진 흰옷과 어깨 붕대, 얼굴과 팔의 피, 몸에 부착된 수신기가 보인다. 현우의 얼굴은 옆모습이라 정확한 동일인 여부 확인은 제한된다. 잔디에는 쓰러진 로봇 잔해가 세 무리 보이며, 낮의 옥상과 착륙 헬기가 표현되어 있다.",
        "hard_violations": [],
        "physics": "윤성찬의 두 구두는 모두 헬기 외부 발판 위에 놓여 있어 발판이 체중을 지지한다. 따라서 떠 있는 인물은 아니지만, 한 발을 딛고 하차하는 순간보다는 두 발을 모아 선 모습이다. 현우의 발은 화면 밖이므로 접지 상태는 확인할 수 없지만 서 있는 몸통과 다리는 자연스럽다. 헬기는 보이는 착륙 바퀴로 옥상에 접지하고, 로봇 잔해는 잔디에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "윤성찬의 얼굴과 시선은 왼쪽 아래에 웅크린 현우를 향한다. 현우도 오른쪽 위의 윤성찬 쪽으로 얼굴을 든다. 상대를 내려다보는 관계는 명확하지만, 윤성찬의 표정은 비스듬히 올린 입꼬리보다는 치아가 드러나는 넓은 웃음에 가깝다.",
        "built_space": "왼쪽에는 문 하나, 환기구 하나, 벽등 하나가 있는 옥탑 벽이 보인다. 중앙 뒤에는 접시 안테나 하나와 양옆 혼 스피커 두 개가 있다. 오른쪽 헬기의 열린 출입구, 손잡이, 발판과 착륙 바퀴가 윤성찬의 자세에 맞게 배치되어 있다. 낮은 전경의 현우와 윤성찬 사이에 연속된 옥상 바닥이 보이고 헬기 대부분은 오른쪽 밖에 있다. 그러나 중앙 설비는 이 후보 역시 우선 구조 참조의 다층 사각 스피커 설비와 다르다.",
        "entities": "보이는 사람은 윤성찬과 현우 두 명이다. 윤성찬의 회색 머리, 안경과 고령의 얼굴은 참조에 가깝지만 긴 코트 없이 줄무늬 정장과 조끼를 입고 있다. 현우는 검은 머리의 앳된 동아시아계 남성으로, 흰 환자복과 붕대, 얼굴 및 어깨의 출혈, 몸에 부착된 수신기가 보인다. 가슴에는 파란색 'R' 형태의 판독 가능한 글자 로고가 남아 있다. 지상 헬기 외에 하늘에 헬기 세 대가 있고 잔디에는 쓰러진 전투 로봇들이 보인다.",
        "hard_violations": [
         "현우의 가슴에 판독 가능한 파란색 글자 로고가 노출되어, 읽을 수 있는 문자와 로고를 모두 금지한 조건을 위반한다."
        ],
        "physics": "윤성찬은 한쪽 구두를 옥상 바닥에, 다른 구두를 헬기 발판에 놓고 한 손으로 출입구 손잡이를 잡는다. 지지점과 무릎 굽힘이 명확하여 하차 중 멈춘 자세로 성립한다. 현우는 몸을 낮추고 무릎을 굽힌 모습이며 하단의 접지 부분은 프레임에 잘려 있다. 착륙 헬기는 바퀴로 지지되고, 공중의 세 헬기는 회전익 비행으로 설명되며, 로봇들은 바닥에 누워 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "윤성찬의 전신과 현우를 내려다보는 시선은 맞지만, 두 발을 발판에 모은 자세와 구조 참조에 맞지 않는 안테나가 아쉽다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "한 발을 헬기 스텝에 딛는 동작과 두 사람의 고저차는 더 정확하지만, 현우 가슴의 판독 가능한 파란 글자 로고가 명시적인 문자·로고 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "윤성찬은 고개와 눈을 왼쪽 아래의 현우에게 향하고 웃는다. 현우는 오른쪽 위의 윤성찬 얼굴을 올려다본다. 두 사람의 시선 대상은 맞으며, 조준하는 무기나 손에 든 지시 물체는 없다.",
        "built_space": "왼쪽 옥탑 벽에 문 하나, 문 환기구 하나, 벽등 하나가 보인다. 중앙에는 원형 접시 하나와 좌우 혼 스피커 두 개가 달린 철탑이 있다. 잔디와 파손된 포장면이 전경의 현우에서 오른쪽 헬기 발판까지 이어지며, 헬기 대부분은 오른쪽 밖으로 잘렸다. 다만 중앙 설비는 장소 사진을 따랐고, 우선권이 있는 구조 사진의 3단 사각 스피커 배열 및 주변 독립 혼 스피커 배치를 재현하지 않았다.",
        "entities": "사람은 두 명뿐이다. 윤성찬은 회색 머리, 안경, 주름진 고령의 동아시아계 남성으로 보이며 짙은 긴 코트와 정장이 참조에 가깝다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로, 찢어진 흰옷과 어깨 붕대, 얼굴과 팔의 피, 몸에 부착된 수신기가 보인다. 현우의 얼굴은 옆모습이라 정확한 동일인 여부 확인은 제한된다. 잔디에는 쓰러진 로봇 잔해가 세 무리 보이며, 낮의 옥상과 착륙 헬기가 표현되어 있다.",
        "hard_violations": [],
        "physics": "윤성찬의 두 구두는 모두 헬기 외부 발판 위에 놓여 있어 발판이 체중을 지지한다. 따라서 떠 있는 인물은 아니지만, 한 발을 딛고 하차하는 순간보다는 두 발을 모아 선 모습이다. 현우의 발은 화면 밖이므로 접지 상태는 확인할 수 없지만 서 있는 몸통과 다리는 자연스럽다. 헬기는 보이는 착륙 바퀴로 옥상에 접지하고, 로봇 잔해는 잔디에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "윤성찬의 얼굴과 시선은 왼쪽 아래에 웅크린 현우를 향한다. 현우도 오른쪽 위의 윤성찬 쪽으로 얼굴을 든다. 상대를 내려다보는 관계는 명확하지만, 윤성찬의 표정은 비스듬히 올린 입꼬리보다는 치아가 드러나는 넓은 웃음에 가깝다.",
        "built_space": "왼쪽에는 문 하나, 환기구 하나, 벽등 하나가 있는 옥탑 벽이 보인다. 중앙 뒤에는 접시 안테나 하나와 양옆 혼 스피커 두 개가 있다. 오른쪽 헬기의 열린 출입구, 손잡이, 발판과 착륙 바퀴가 윤성찬의 자세에 맞게 배치되어 있다. 낮은 전경의 현우와 윤성찬 사이에 연속된 옥상 바닥이 보이고 헬기 대부분은 오른쪽 밖에 있다. 그러나 중앙 설비는 이 후보 역시 우선 구조 참조의 다층 사각 스피커 설비와 다르다.",
        "entities": "보이는 사람은 윤성찬과 현우 두 명이다. 윤성찬의 회색 머리, 안경과 고령의 얼굴은 참조에 가깝지만 긴 코트 없이 줄무늬 정장과 조끼를 입고 있다. 현우는 검은 머리의 앳된 동아시아계 남성으로, 흰 환자복과 붕대, 얼굴 및 어깨의 출혈, 몸에 부착된 수신기가 보인다. 가슴에는 파란색 'R' 형태의 판독 가능한 글자 로고가 남아 있다. 지상 헬기 외에 하늘에 헬기 세 대가 있고 잔디에는 쓰러진 전투 로봇들이 보인다.",
        "hard_violations": [
         "현우의 가슴에 판독 가능한 파란색 글자 로고가 노출되어, 읽을 수 있는 문자와 로고를 모두 금지한 조건을 위반한다."
        ],
        "physics": "윤성찬은 한쪽 구두를 옥상 바닥에, 다른 구두를 헬기 발판에 놓고 한 손으로 출입구 손잡이를 잡는다. 지지점과 무릎 굽힘이 명확하여 하차 중 멈춘 자세로 성립한다. 현우는 몸을 낮추고 무릎을 굽힌 모습이며 하단의 접지 부분은 프레임에 잘려 있다. 착륙 헬기는 바퀴로 지지되고, 공중의 세 헬기는 회전익 비행으로 설명되며, 로봇들은 바닥에 누워 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.667
   },
   "violations": {
    "A": [
     "[gpt-high] 현우의 가슴에 판독 가능한 파란색 글자 로고가 노출되어, 읽을 수 있는 문자와 로고를 모두 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1667
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "윤성찬의 오버코트가 누락되었으나, '한 발을 딛고 선 채'라는 핵심 동작을 완벽하게 구현하였고 주변 항공기와 파괴된 로봇 등 디테일의 재현도가 매우 뛰어남.  ★위반: [gpt-high] 현우의 가슴에 판독 가능한 파란색 글자 로고가 노출되어, 읽을 수 있는 문자와 로고를 모두 금지한 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "윤성찬의 복장은 참조와 일치하나, 지시된 '한 발을 딛고' 있는 포즈를 무시하고 두 발로 스텝에 올라서 있으며 주변 환경 묘사가 상대적으로 단조로움."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B01.png",
    "asset_id": "6eab84a7-fe74-47c0-bd42-78a5d724b6a9",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_island_research_complex_sel.png",
    "asset_id": "ae92432a-4956-46ca-9845-8c2d71d637c5",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:886743>",
    "asset_id": "6e844efd-130d-4b1c-82d7-4625e44c215b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1337888>",
    "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-20fd-7f8f-aec8-901819bbb181",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42__bgfirst_bg.png",
   "bg_asset_id": "99219a2e-c866-4ccd-9da5-2b903d22263a",
   "bg_record_key": "S89sh42::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready",
  "staged_characters_added": [
   "C01"
  ]
 },
 "S89sh42::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:16:42.241819+00:00",
  "fingerprint": "eb39c65e223222d58de4d46e7afd77e67d280db8ebca54f6e4db78afe03e2e09",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S89sh42_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S89sh42_sel.png",
  "source_sha256": "d444caf3df73c5ee7cbefb5a16757e3a3e7d4659d6254f112df1c80e024cd4d9",
  "file": "S89sh42_cine.png",
  "staged_sha256": "b3027ff90b781edbdd7f79b70bf0cf0181e718ccc70613bcd7faa91d985e0ea7",
  "latency_ms": 10441
 },
 "S89sh53::signage": {
  "fp": "4cc60383d8a181c7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S89sh53": {
  "input_fingerprint": "1f866d2dde3e3650",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 링에서 하늘을 향해 거대하고 눈부신 새하얀 광선 기둥이 수직으로 팽팽하게 뻗어 나간 압도적인 폭발 찰나.\n\nLOCATION (lock): On the research facility's exposed rooftop lawn near its large antenna and all-direction speakers, beneath the vertical energy beam. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Vertical white beam continuing directly upward from Charlie's chest ring in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Rooftop park lawn (Beneath the figures during the beam's eruption) — A shallow strip remains visible across the bottom of the composition; used as Anchors the extraordinary vertical event to the established rooftop location; Sky above the rooftop (Visible around the ascending beam); used as Provides lateral negative space that makes the beam's vertical extension legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant white chest beam overwhelms the daylight locally, with controlled highlight bloom retaining its origin and the figures at its base.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's chest ring is spinning at extreme speed and projecting a brilliant, nearly star-white beam vertically into the sky; the experimental sensors remain attached. The battered rooftop, landed command helicopter and surrounding aircraft remain in place as lights surge and dim and the control-room energy graph shoots upward. 현우: He remains wounded on the rooftop with the wireless receiver attached to his body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 링에서 하늘을 향해 거대하고 눈부신 새하얀 광선 기둥이 수직으로 팽팽하게 뻗어 나간 압도적인 폭발 찰나.\n\nLOCATION (lock): On the research facility's exposed rooftop lawn near its large antenna and all-direction speakers, beneath the vertical energy beam. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Vertical white beam continuing directly upward from Charlie's chest ring in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Rooftop park lawn (Beneath the figures during the beam's eruption) — A shallow strip remains visible across the bottom of the composition; used as Anchors the extraordinary vertical event to the established rooftop location; Sky above the rooftop (Visible around the ascending beam); used as Provides lateral negative space that makes the beam's vertical extension legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant white chest beam overwhelms the daylight locally, with controlled highlight bloom retaining its origin and the figures at its base.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's chest ring is spinning at extreme speed and projecting a brilliant, nearly star-white beam vertically into the sky; the experimental sensors remain attached. The battered rooftop, landed command helicopter and surrounding aircraft remain in place as lights surge and dim and the control-room energy graph shoots upward. 현우: He remains wounded on the rooftop with the wireless receiver attached to his body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 링에서 하늘을 향해 거대하고 눈부신 새하얀 광선 기둥이 수직으로 팽팽하게 뻗어 나간 압도적인 폭발 찰나.\n\nLOCATION (lock): On the research facility's exposed rooftop lawn near its large antenna and all-direction speakers, beneath the vertical energy beam. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Vertical white beam continuing directly upward from Charlie's chest ring in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Rooftop park lawn (Beneath the figures during the beam's eruption) — A shallow strip remains visible across the bottom of the composition; used as Anchors the extraordinary vertical event to the established rooftop location; Sky above the rooftop (Visible around the ascending beam); used as Provides lateral negative space that makes the beam's vertical extension legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant white chest beam overwhelms the daylight locally, with controlled highlight bloom retaining its origin and the figures at its base.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's chest ring is spinning at extreme speed and projecting a brilliant, nearly star-white beam vertically into the sky; the experimental sensors remain attached. The battered rooftop, landed command helicopter and surrounding aircraft remain in place as lights surge and dim and the control-room energy graph shoots upward. 현우: He remains wounded on the rooftop with the wireless receiver attached to his body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "거대한 백색 광선이 찰리의 가슴 중심에서 수직으로 하늘을 향해 뻗어 나감.",
    "built_space": "잔디가 깔린 옥상과 콘크리트 외벽이 보이나, 프롬프트와 이전 샷에 명시된 대형 안테나와 스피커 구조물은 프레임 내에 보이지 않음.",
    "entities": "찰리는 로봇 형태를 유지하나 지시된 코트와 모자를 착용하지 않은 맨몸 상태임. 현우는 이전 샷의 상태와 일관되게 서서 하늘을 응시하고 있음.",
    "hard_violations": [],
    "physics": "광선은 찰리의 가슴 링에서 정상적으로 발사되고 있으며, 두 인물 모두 지면에 안정적으로 발을 딛고 서 있음."
   },
   {
    "label": "B",
    "direction": "거대한 백색 광선이 찰리의 가슴이 아닌 그 앞의 지면에서 수직으로 솟아오름.",
    "built_space": "옥상 잔디밭 배경에 이전 샷에 등장했던 대형 안테나, 방향성 스피커, 그리고 바닥의 장비 파편들이 정확하게 위치함.",
    "entities": "찰리는 코트를 입고 가슴에 센서를 부착했으나 모자가 없음. 현우는 이전 샷과 달리 캐릭터 레퍼런스 이미지의 주저앉은 포즈를 유사하게 취하고 있음.",
    "hard_violations": [],
    "physics": "인물들은 지면의 지지를 받고 있으나, 광선의 발사 지점이 찰리의 몸과 완전히 분리되어 바닥의 흙을 튀기며 솟아오르는 물리적 불일치를 보임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "찰리의 의상과 배경의 안테나 구조물이 누락되었으나, 가슴 링에서 수직으로 뻗어나가는 광선의 핵심 액션과 구도를 최우선적으로 정확히 구현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경의 구조물과 코트의 질감 등 디테일은 훌륭하나, 프롬프트의 핵심인 광선이 가슴이 아닌 지면에서 발생하는 치명적인 연출 오류를 범함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "거대한 백색 광선이 찰리의 가슴 중심에서 수직으로 하늘을 향해 뻗어 나감.",
        "built_space": "잔디가 깔린 옥상과 콘크리트 외벽이 보이나, 프롬프트와 이전 샷에 명시된 대형 안테나와 스피커 구조물은 프레임 내에 보이지 않음.",
        "entities": "찰리는 로봇 형태를 유지하나 지시된 코트와 모자를 착용하지 않은 맨몸 상태임. 현우는 이전 샷의 상태와 일관되게 서서 하늘을 응시하고 있음.",
        "hard_violations": [],
        "physics": "광선은 찰리의 가슴 링에서 정상적으로 발사되고 있으며, 두 인물 모두 지면에 안정적으로 발을 딛고 서 있음."
       },
       {
        "label": "B",
        "direction": "거대한 백색 광선이 찰리의 가슴이 아닌 그 앞의 지면에서 수직으로 솟아오름.",
        "built_space": "옥상 잔디밭 배경에 이전 샷에 등장했던 대형 안테나, 방향성 스피커, 그리고 바닥의 장비 파편들이 정확하게 위치함.",
        "entities": "찰리는 코트를 입고 가슴에 센서를 부착했으나 모자가 없음. 현우는 이전 샷과 달리 캐릭터 레퍼런스 이미지의 주저앉은 포즈를 유사하게 취하고 있음.",
        "hard_violations": [],
        "physics": "인물들은 지면의 지지를 받고 있으나, 광선의 발사 지점이 찰리의 몸과 완전히 분리되어 바닥의 흙을 튀기며 솟아오르는 물리적 불일치를 보임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "찰리의 의상과 배경의 안테나 구조물이 누락되었으나, 가슴 링에서 수직으로 뻗어나가는 광선의 핵심 액션과 구도를 최우선적으로 정확히 구현함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경의 구조물과 코트의 질감 등 디테일은 훌륭하나, 프롬프트의 핵심인 광선이 가슴이 아닌 지면에서 발생하는 치명적인 연출 오류를 범함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "거대한 백색 광선이 찰리의 가슴 중심에서 수직으로 하늘을 향해 뻗어 나감.",
        "built_space": "잔디가 깔린 옥상과 콘크리트 외벽이 보이나, 프롬프트와 이전 샷에 명시된 대형 안테나와 스피커 구조물은 프레임 내에 보이지 않음.",
        "entities": "찰리는 로봇 형태를 유지하나 지시된 코트와 모자를 착용하지 않은 맨몸 상태임. 현우는 이전 샷의 상태와 일관되게 서서 하늘을 응시하고 있음.",
        "hard_violations": [],
        "physics": "광선은 찰리의 가슴 링에서 정상적으로 발사되고 있으며, 두 인물 모두 지면에 안정적으로 발을 딛고 서 있음."
       },
       {
        "label": "B",
        "direction": "거대한 백색 광선이 찰리의 가슴이 아닌 그 앞의 지면에서 수직으로 솟아오름.",
        "built_space": "옥상 잔디밭 배경에 이전 샷에 등장했던 대형 안테나, 방향성 스피커, 그리고 바닥의 장비 파편들이 정확하게 위치함.",
        "entities": "찰리는 코트를 입고 가슴에 센서를 부착했으나 모자가 없음. 현우는 이전 샷과 달리 캐릭터 레퍼런스 이미지의 주저앉은 포즈를 유사하게 취하고 있음.",
        "hard_violations": [],
        "physics": "인물들은 지면의 지지를 받고 있으나, 광선의 발사 지점이 찰리의 몸과 완전히 분리되어 바닥의 흙을 튀기며 솟아오르는 물리적 불일치를 보임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "옥상 시설과 센서는 잘 보이지만, 광선이 찰리의 가슴이 아니라 오른쪽 잔디에서 솟아 핵심 발사 방향과 기점을 어긴다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "가슴 링에서 상단 중앙으로 이어지는 수직 백색광과 얕은 잔디 띠가 핵심 구도에 부합하지만, 모자·긴 코트가 빠지고 얼굴 주변의 광량 번짐이 과하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "광선은 화면 중앙에서 하늘로 수직 상승하지만, 하단 기점은 찰리의 오른쪽 잔디다. 왼쪽에 따로 빛나는 가슴 링과 광선 기둥이 연결되지 않는다. 찰리는 고개를 오른쪽 아래로 숙이고, 현우는 왼쪽 위의 찰리와 발광부 쪽을 바라본다.",
        "built_space": "파손된 옥상 잔디와 콘크리트 난간이 있으며, 오른쪽에는 접시 안테나 1개, 철탑 1개, 좌우로 향한 혼 스피커 2개, 하단 장비함 1개가 보인다. 참고 장소의 주요 시설 형식은 유지된다. 찰리는 왼쪽 잔디에 서고 현우는 오른쪽 잔디에 앉아 있다. 잔디가 화면 하단의 상당 부분을 차지하고 안테나도 크게 보여, 요구된 얕은 바닥 띠와 작은 배경 시설보다 비중이 크다. 헬리콥터는 보이지 않는다.",
        "entities": "찰리 1명과 현우 1명이 보이며 다른 사람은 없다. 찰리는 긴 팔, 짧은 다리, 베이지 장갑, 흰 기계 마스크, 낡은 긴 코트를 갖췄지만 참고의 모자는 없다. 가슴 링 주변에는 부착 센서와 배선이 보인다. 현우는 젊은 동아시아계 남성으로 검은 머리, 피 묻은 흰옷, 몸에 부착된 검은 수신기가 보여 설정과 대체로 맞는다. 링의 극고속 회전은 발광 때문에 명확히 판별하기 어렵다.",
        "hard_violations": [],
        "physics": "찰리의 발은 잔디에 닿아 있고, 현우는 엉덩이와 뒤로 짚은 손으로 몸을 지탱한다. 공중의 흙과 잔해는 바로 아래 잔디 폭발에서 튀어 오른 것으로 읽힌다. 지지 없이 떠 있는 몸은 없지만, 이 지면 폭발은 요구된 가슴 링 발사와 다른 물리적 기점을 만든다."
       },
       {
        "label": "B",
        "direction": "백색 광선은 찰리의 가슴 링 중심에서 화면 상단 중앙으로 곧게 이어져 하늘을 향한다. 요구된 발사 기점과 수직 방향이 맞는다. 현우는 턱을 들고 위쪽 발광부를 바라보며, 찰리의 정확한 시선은 얼굴을 덮는 광량 번짐 때문에 확인하기 어렵다.",
        "built_space": "콘크리트 난간 뒤로 낮은 도시 윤곽과 산이 보이고, 옥상 잔디는 화면 맨 아래의 얕은 띠로 남는다. 찰리는 중앙 중경에 서고 현우는 그 앞 오른쪽에 서 있다. 오른쪽 가장자리에는 헬리콥터 로터 끝으로 보이는 부분만 걸린다. 안테나와 스피커는 프레임 안에 없어 참고 시설의 세부 일치 여부를 확인할 수 없지만, 중복 시설이나 불가능한 공간 배치는 보이지 않는다.",
        "entities": "찰리 1명과 현우 1명만 보인다. 찰리의 육중한 몸체, 긴 팔, 베이지 장갑과 원형 가슴 링은 맞지만, 참고의 모자와 긴 코트 자락이 없고 얼굴도 빛에 상당히 가려진다. 링 주변 부착물이 보이나 실험 센서의 형태는 뚜렷하지 않다. 현우는 젊은 동아시아계 남성, 헝클어진 검은 머리, 상처와 피 묻은 흰옷으로 대체로 일치한다. 수신기는 이 각도와 크기에서 확실히 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 두 발과 현우의 두 발이 잔디에 닿아 각각 몸을 지탱한다. 찰리의 긴 팔은 어깨에서 자연스럽게 내려오며, 현우도 지면 위에 서 있다. 광선은 설정된 에너지원인 가슴 링에 연결되어 있고, 근거 없이 떠 있는 사람이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "옥상 시설과 센서는 잘 보이지만, 광선이 찰리의 가슴이 아니라 오른쪽 잔디에서 솟아 핵심 발사 방향과 기점을 어긴다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴 링에서 상단 중앙으로 이어지는 수직 백색광과 얕은 잔디 띠가 핵심 구도에 부합하지만, 모자·긴 코트가 빠지고 얼굴 주변의 광량 번짐이 과하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "광선은 화면 중앙에서 하늘로 수직 상승하지만, 하단 기점은 찰리의 오른쪽 잔디다. 왼쪽에 따로 빛나는 가슴 링과 광선 기둥이 연결되지 않는다. 찰리는 고개를 오른쪽 아래로 숙이고, 현우는 왼쪽 위의 찰리와 발광부 쪽을 바라본다.",
        "built_space": "파손된 옥상 잔디와 콘크리트 난간이 있으며, 오른쪽에는 접시 안테나 1개, 철탑 1개, 좌우로 향한 혼 스피커 2개, 하단 장비함 1개가 보인다. 참고 장소의 주요 시설 형식은 유지된다. 찰리는 왼쪽 잔디에 서고 현우는 오른쪽 잔디에 앉아 있다. 잔디가 화면 하단의 상당 부분을 차지하고 안테나도 크게 보여, 요구된 얕은 바닥 띠와 작은 배경 시설보다 비중이 크다. 헬리콥터는 보이지 않는다.",
        "entities": "찰리 1명과 현우 1명이 보이며 다른 사람은 없다. 찰리는 긴 팔, 짧은 다리, 베이지 장갑, 흰 기계 마스크, 낡은 긴 코트를 갖췄지만 참고의 모자는 없다. 가슴 링 주변에는 부착 센서와 배선이 보인다. 현우는 젊은 동아시아계 남성으로 검은 머리, 피 묻은 흰옷, 몸에 부착된 검은 수신기가 보여 설정과 대체로 맞는다. 링의 극고속 회전은 발광 때문에 명확히 판별하기 어렵다.",
        "hard_violations": [],
        "physics": "찰리의 발은 잔디에 닿아 있고, 현우는 엉덩이와 뒤로 짚은 손으로 몸을 지탱한다. 공중의 흙과 잔해는 바로 아래 잔디 폭발에서 튀어 오른 것으로 읽힌다. 지지 없이 떠 있는 몸은 없지만, 이 지면 폭발은 요구된 가슴 링 발사와 다른 물리적 기점을 만든다."
       },
       {
        "label": "A",
        "direction": "백색 광선은 찰리의 가슴 링 중심에서 화면 상단 중앙으로 곧게 이어져 하늘을 향한다. 요구된 발사 기점과 수직 방향이 맞는다. 현우는 턱을 들고 위쪽 발광부를 바라보며, 찰리의 정확한 시선은 얼굴을 덮는 광량 번짐 때문에 확인하기 어렵다.",
        "built_space": "콘크리트 난간 뒤로 낮은 도시 윤곽과 산이 보이고, 옥상 잔디는 화면 맨 아래의 얕은 띠로 남는다. 찰리는 중앙 중경에 서고 현우는 그 앞 오른쪽에 서 있다. 오른쪽 가장자리에는 헬리콥터 로터 끝으로 보이는 부분만 걸린다. 안테나와 스피커는 프레임 안에 없어 참고 시설의 세부 일치 여부를 확인할 수 없지만, 중복 시설이나 불가능한 공간 배치는 보이지 않는다.",
        "entities": "찰리 1명과 현우 1명만 보인다. 찰리의 육중한 몸체, 긴 팔, 베이지 장갑과 원형 가슴 링은 맞지만, 참고의 모자와 긴 코트 자락이 없고 얼굴도 빛에 상당히 가려진다. 링 주변 부착물이 보이나 실험 센서의 형태는 뚜렷하지 않다. 현우는 젊은 동아시아계 남성, 헝클어진 검은 머리, 상처와 피 묻은 흰옷으로 대체로 일치한다. 수신기는 이 각도와 크기에서 확실히 식별되지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 두 발과 현우의 두 발이 잔디에 닿아 각각 몸을 지탱한다. 찰리의 긴 팔은 어깨에서 자연스럽게 내려오며, 현우도 지면 위에 서 있다. 광선은 설정된 에너지원인 가슴 링에 연결되어 있고, 근거 없이 떠 있는 사람이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.238
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.238
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1238
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "찰리의 의상과 배경의 안테나 구조물이 누락되었으나, 가슴 링에서 수직으로 뻗어나가는 광선의 핵심 액션과 구도를 최우선적으로 정확히 구현함."
   },
   {
    "label": "B",
    "score": 1238,
    "verdict_ko": "배경의 구조물과 코트의 질감 등 디테일은 훌륭하나, 프롬프트의 핵심인 광선이 가슴이 아닌 지면에서 발생하는 치명적인 연출 오류를 범함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42_sel.png",
    "asset_id": "6a613566-2709-42d4-8578-97299e41c781",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1337888>",
    "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-246e-7514-addc-4282a5fb858f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S89sh42"
  },
  "staged_characters_added": [
   "C01"
  ]
 },
 "S89sh53::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:17:50.971096+00:00",
  "fingerprint": "3d7f215b350eac0202c3fff6b69b91230295ccfb2853056bb6b41980f351fe72",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S89sh53_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S89sh53_sel.png",
  "source_sha256": "0109ded33cfa7a41aa7c6e7f68208d0eba24608c95552e31f599a1a9732c780b",
  "file": "S89sh53_cine.png",
  "staged_sha256": "402851ce880c9aebd1d95341f600c2fc78e73b2303efa0bafa6c3f5803d89009",
  "latency_ms": 10820
 },
 "S89sh66::signage": {
  "fp": "fa4dd617f446ced4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S89sh66::bgfirst_bg": {
  "input_fingerprint": "902b8a5a7deb272f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh66__bgfirst_bg.png",
  "asset_id": "bf0ba719-2d49-4122-ad48-c1530b3f79c4",
  "input_asset_ids": [
   "7273645e-cb70-462d-a06c-54beb3d21fa8",
   "76fc6513-8b50-4f33-b05a-840ac5bb52b7"
  ]
 },
 "S89sh66": {
  "input_fingerprint": "05d05a720b91f1fe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie's powered-off remains are held tightly against Hyunwoo's body in both of his arms. Only part of Charlie's destroyed body remains, and the source does not specify its orientation within the embrace or the arrangement of any surviving appendages.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The beam has vanished, leaving only a partially surviving torso and charred remnants of Charlie outside the research facility at night. Charlie's remaining power has shut down completely. 현우: He remains wounded from the attack and is sobbing with both arms tightly closed around the remains in his grasp.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie's powered-off remains are held tightly against Hyunwoo's body in both of his arms. Only part of Charlie's destroyed body remains, and the source does not specify its orientation within the embrace or the arrangement of any surviving appendages.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The beam has vanished, leaving only a partially surviving torso and charred remnants of Charlie outside the research facility at night. Charlie's remaining power has shut down completely. 현우: He remains wounded from the attack and is sobbing with both arms tightly closed around the remains in his grasp.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie's powered-off remains are held tightly against Hyunwoo's body in both of his arms. Only part of Charlie's destroyed body remains, and the source does not specify its orientation within the embrace or the arrangement of any surviving appendages.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The beam has vanished, leaving only a partially surviving torso and charred remnants of Charlie outside the research facility at night. Charlie's remaining power has shut down completely. 현우: He remains wounded from the attack and is sobbing with both arms tightly closed around the remains in his grasp.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh66__bgfirst_bg.png",
     "asset_id": "bf0ba719-2d49-4122-ad48-c1530b3f79c4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S89sh66.png",
     "asset_id": "7273645e-cb70-462d-a06c-54beb3d21fa8",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1337888>",
     "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1299563>",
     "asset_id": "2948628a-764c-4354-a3e9-c21d7ef85260",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B02.png",
     "asset_id": "76fc6513-8b50-4f33-b05a-840ac5bb52b7",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1337888>",
     "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1299563>",
     "asset_id": "2948628a-764c-4354-a3e9-c21d7ef85260",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 고개와 시선이 하늘을 향해 위로 꺾여 있음.",
    "built_space": "건물, 통신탑, 헬리콥터가 있으나, 레퍼런스의 핵심 지형인 단이 돋아난 잔디 화단 구역이 사라지고 완전히 평평한 콘크리트 바닥으로 변형됨.",
    "entities": "현우는 참고 이미지와 닮았으나 반소매를 입고 있음. 찰리의 잔해는 머리와 상체가 남아있으나 어린아이 크기 이하로 심각하게 축소됨.",
    "hard_violations": [],
    "physics": "현우는 바닥에 무릎을 꿇고 하반신을 지탱하며, 축소된 로봇 잔해를 양팔로 안고 있음. 떠 있는 물체는 없음."
   },
   {
    "label": "B",
    "direction": "현우의 고개와 시선이 하늘을 향해 완전히 젖혀진 채 위를 향하고 있음.",
    "built_space": "건물, 단이 돋아난 잔디 화단, 통신탑, 헬리콥터 등 레퍼런스의 아키텍처와 구조물 배치가 야간 환경 속에 정확히 구현됨.",
    "entities": "현우는 피 묻은 긴팔 의상을 입고 레퍼런스와 일치함. 찰리의 잔해는 가슴 장갑과 전선이 드러난 채 원래의 거대한 비율을 유지하고 있음.",
    "hard_violations": [],
    "physics": "현우는 바닥에 두 발로 굳게 서서 체중을 지탱하며, 무거운 로봇 흉부 잔해를 두 손으로 꽉 쥐어 가슴에 밀착시키고 있음. 허공에 뜬 물체나 인체 오류 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "돋아난 잔디 화단을 포함한 장소의 아키텍처를 정확히 구현(Priority 3)했으며, 로봇 잔해의 거대한 스케일을 현실적으로 유지한 채 오열하는 전신의 감정 표현을 완벽하게 담아냈습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "레퍼런스의 주요 건축 지형인 잔디 화단을 평지로 왜곡하여 장소 일치(Priority 3) 지침을 위반했으며, 거대한 로봇 잔해를 장난감처럼 비현실적으로 작게 축소해 개체 묘사가 부실합니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 고개와 시선이 하늘을 향해 완전히 젖혀진 채 위를 향하고 있음.",
        "built_space": "건물, 단이 돋아난 잔디 화단, 통신탑, 헬리콥터 등 레퍼런스의 아키텍처와 구조물 배치가 야간 환경 속에 정확히 구현됨.",
        "entities": "현우는 피 묻은 긴팔 의상을 입고 레퍼런스와 일치함. 찰리의 잔해는 가슴 장갑과 전선이 드러난 채 원래의 거대한 비율을 유지하고 있음.",
        "hard_violations": [],
        "physics": "현우는 바닥에 두 발로 굳게 서서 체중을 지탱하며, 무거운 로봇 흉부 잔해를 두 손으로 꽉 쥐어 가슴에 밀착시키고 있음. 허공에 뜬 물체나 인체 오류 없음."
       },
       {
        "label": "A",
        "direction": "현우의 고개와 시선이 하늘을 향해 위로 꺾여 있음.",
        "built_space": "건물, 통신탑, 헬리콥터가 있으나, 레퍼런스의 핵심 지형인 단이 돋아난 잔디 화단 구역이 사라지고 완전히 평평한 콘크리트 바닥으로 변형됨.",
        "entities": "현우는 참고 이미지와 닮았으나 반소매를 입고 있음. 찰리의 잔해는 머리와 상체가 남아있으나 어린아이 크기 이하로 심각하게 축소됨.",
        "hard_violations": [],
        "physics": "현우는 바닥에 무릎을 꿇고 하반신을 지탱하며, 축소된 로봇 잔해를 양팔로 안고 있음. 떠 있는 물체는 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "돋아난 잔디 화단을 포함한 장소의 아키텍처를 정확히 구현(Priority 3)했으며, 로봇 잔해의 거대한 스케일을 현실적으로 유지한 채 오열하는 전신의 감정 표현을 완벽하게 담아냈습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "레퍼런스의 주요 건축 지형인 잔디 화단을 평지로 왜곡하여 장소 일치(Priority 3) 지침을 위반했으며, 거대한 로봇 잔해를 장난감처럼 비현실적으로 작게 축소해 개체 묘사가 부실합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 고개와 시선이 하늘을 향해 완전히 젖혀진 채 위를 향하고 있음.",
        "built_space": "건물, 단이 돋아난 잔디 화단, 통신탑, 헬리콥터 등 레퍼런스의 아키텍처와 구조물 배치가 야간 환경 속에 정확히 구현됨.",
        "entities": "현우는 피 묻은 긴팔 의상을 입고 레퍼런스와 일치함. 찰리의 잔해는 가슴 장갑과 전선이 드러난 채 원래의 거대한 비율을 유지하고 있음.",
        "hard_violations": [],
        "physics": "현우는 바닥에 두 발로 굳게 서서 체중을 지탱하며, 무거운 로봇 흉부 잔해를 두 손으로 꽉 쥐어 가슴에 밀착시키고 있음. 허공에 뜬 물체나 인체 오류 없음."
       },
       {
        "label": "A",
        "direction": "현우의 고개와 시선이 하늘을 향해 위로 꺾여 있음.",
        "built_space": "건물, 통신탑, 헬리콥터가 있으나, 레퍼런스의 핵심 지형인 단이 돋아난 잔디 화단 구역이 사라지고 완전히 평평한 콘크리트 바닥으로 변형됨.",
        "entities": "현우는 참고 이미지와 닮았으나 반소매를 입고 있음. 찰리의 잔해는 머리와 상체가 남아있으나 어린아이 크기 이하로 심각하게 축소됨.",
        "hard_violations": [],
        "physics": "현우는 바닥에 무릎을 꿇고 하반신을 지탱하며, 축소된 로봇 잔해를 양팔로 안고 있음. 떠 있는 물체는 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "전신 와이드 구도에서 하늘을 향한 오열과 양팔의 밀착 포옹을 구현하고, 파괴된 잔디 지면과 부상·의상도 더 충실하지만 명시된 낮 시간에는 어긋납니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "전신과 주변 여백, 하늘을 향한 오열은 구현했으나 장소의 잔디·폭발 흔적을 넓은 포장면으로 바꾸고 현우의 의상도 달라졌으며 낮 시간 역시 지키지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 턱과 얼굴을 위로 들어 하늘을 향하고 입을 크게 벌리고 있습니다. 양손과 팔은 앞가슴의 찰리 몸통 잔해를 감싸며, 한 손은 파손된 가슴 부분에, 다른 손은 아래쪽 장갑판에 닿습니다. 찰리의 머리는 보이지 않아 시선 방향을 판단할 대상이 없습니다.",
        "built_space": "왼쪽에 환기 루버가 달린 문 1개와 그 위 벽등 1개, 뒤쪽에 접시 안테나 1개가 달린 통신탑 1개, 오른쪽에 헬리콥터 1대가 있습니다. 통신탑 확성기는 한쪽이 뚜렷하고 나머지는 인물에 가려집니다. 낮은 외곽벽, 먼 도시, 잔디와 폭발로 뒤집힌 지면, 젖은 전경 포장이 참고 장소와 잘 대응합니다. 현우는 포장면에 서 있으며 머리부터 발끝까지 보이고 좌우에 빈 공간이 남습니다. 젖은 바닥의 반사에 명백한 광학적 모순은 없습니다. 다만 도시 불빛과 어두운 환경은 야간으로 읽혀, 상충하는 지시 중 명시적인 낮 시간 잠금은 충족하지 못합니다.",
        "entities": "사람은 현우로 보이는 젊은 동아시아계 남성 1명뿐입니다. 헝클어진 검은 머리, 앳된 얼굴, 얼굴과 팔의 상처, 피 묻고 찢어진 흰 긴소매 상의와 흰 바지가 참고와 대체로 맞습니다. 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없습니다. 품에는 샌드 베이지색 각진 장갑판과 검게 파손된 내부가 있는 찰리의 부분 몸통이 있고, 머리와 온전한 팔다리는 없습니다. 이는 일부만 살아남은 잔해라는 조건에 맞습니다. 주변에는 바닥에 놓인 장갑 파편이 있으며 추가 인물, 광선, 판독 가능한 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "현우는 벌린 두 발로 바닥을 딛고 몸통을 세우고 있습니다. 찰리의 잔해는 현우의 가슴에 밀착되어 두 손과 굽힌 팔로 지지됩니다. 잔해가 스스로 자세를 유지하거나 작동하는 모습은 없고, 분리된 파편은 지면에 놓여 있습니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 뒤로 젖혀 화면 오른쪽 위 하늘을 향하고 입을 벌려 울고 있습니다. 두 팔은 찰리의 몸통을 가로질러 감싸고 양손이 잔해에 닿습니다. 찰리의 흰 얼굴은 카메라 쪽으로 비스듬히 향하지만 눈의 발광이나 능동적인 시선은 보이지 않습니다.",
        "built_space": "왼쪽에 루버가 달린 문 1개와 벽등 1개, 뒤쪽에 접시 안테나 1개와 좌우 확성기 2개를 갖춘 통신탑 1개, 오른쪽에 헬리콥터 1대가 있습니다. 외곽벽과 도시 배경도 유지됩니다. 다만 참고의 넓은 잔디와 뒤집힌 폭발 지면 대부분이 평평한 포장면으로 바뀌었습니다. 현우는 전경의 작은 잔해 구역에 무릎을 꿇고 있으며 전신과 넉넉한 주변 공간이 보입니다. 바닥 반사에 명백한 모순은 없습니다. 어두운 하늘과 켜진 도시 조명 때문에 야간으로 읽혀 명시적인 낮 시간 잠금에는 맞지 않습니다.",
        "entities": "젊고 마른 동아시아계 남성 1명이 현우로 표현되며 검은 머리와 얼굴·팔의 부상이 보입니다. 흰 바지는 참고와 유사하지만 상의는 짧게 찢어진 소매 형태이고 참고의 긴소매 및 대각선 천 띠가 재현되지 않습니다. 찰리는 흰 마스크형 머리와 심하게 탄 부분 몸통으로 남아 있으며 온전한 긴 팔과 다리는 없습니다. 장갑판은 참고보다 검게 소실되어 있으나 파괴된 상태라는 설명에는 부합합니다. 추가 사람이나 광선, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "현우의 양 무릎과 접힌 하체가 지면에 닿아 체중을 받습니다. 찰리의 몸통은 양팔과 가슴 사이에 끼워져 지지되고 아래쪽은 현우의 허벅지 사이로 내려옵니다. 머리는 남은 목 구조에 연결되어 있으며 전선은 아래로 늘어집니다. 주변 파편도 지면에 놓여 있어 지지 없는 부유나 명백히 불가능한 자세는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "전신 와이드 구도에서 하늘을 향한 오열과 양팔의 밀착 포옹을 구현하고, 파괴된 잔디 지면과 부상·의상도 더 충실하지만 명시된 낮 시간에는 어긋납니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "전신과 주변 여백, 하늘을 향한 오열은 구현했으나 장소의 잔디·폭발 흔적을 넓은 포장면으로 바꾸고 현우의 의상도 달라졌으며 낮 시간 역시 지키지 못했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 턱과 얼굴을 위로 들어 하늘을 향하고 입을 크게 벌리고 있습니다. 양손과 팔은 앞가슴의 찰리 몸통 잔해를 감싸며, 한 손은 파손된 가슴 부분에, 다른 손은 아래쪽 장갑판에 닿습니다. 찰리의 머리는 보이지 않아 시선 방향을 판단할 대상이 없습니다.",
        "built_space": "왼쪽에 환기 루버가 달린 문 1개와 그 위 벽등 1개, 뒤쪽에 접시 안테나 1개가 달린 통신탑 1개, 오른쪽에 헬리콥터 1대가 있습니다. 통신탑 확성기는 한쪽이 뚜렷하고 나머지는 인물에 가려집니다. 낮은 외곽벽, 먼 도시, 잔디와 폭발로 뒤집힌 지면, 젖은 전경 포장이 참고 장소와 잘 대응합니다. 현우는 포장면에 서 있으며 머리부터 발끝까지 보이고 좌우에 빈 공간이 남습니다. 젖은 바닥의 반사에 명백한 광학적 모순은 없습니다. 다만 도시 불빛과 어두운 환경은 야간으로 읽혀, 상충하는 지시 중 명시적인 낮 시간 잠금은 충족하지 못합니다.",
        "entities": "사람은 현우로 보이는 젊은 동아시아계 남성 1명뿐입니다. 헝클어진 검은 머리, 앳된 얼굴, 얼굴과 팔의 상처, 피 묻고 찢어진 흰 긴소매 상의와 흰 바지가 참고와 대체로 맞습니다. 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없습니다. 품에는 샌드 베이지색 각진 장갑판과 검게 파손된 내부가 있는 찰리의 부분 몸통이 있고, 머리와 온전한 팔다리는 없습니다. 이는 일부만 살아남은 잔해라는 조건에 맞습니다. 주변에는 바닥에 놓인 장갑 파편이 있으며 추가 인물, 광선, 판독 가능한 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "현우는 벌린 두 발로 바닥을 딛고 몸통을 세우고 있습니다. 찰리의 잔해는 현우의 가슴에 밀착되어 두 손과 굽힌 팔로 지지됩니다. 잔해가 스스로 자세를 유지하거나 작동하는 모습은 없고, 분리된 파편은 지면에 놓여 있습니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "현우는 고개를 뒤로 젖혀 화면 오른쪽 위 하늘을 향하고 입을 벌려 울고 있습니다. 두 팔은 찰리의 몸통을 가로질러 감싸고 양손이 잔해에 닿습니다. 찰리의 흰 얼굴은 카메라 쪽으로 비스듬히 향하지만 눈의 발광이나 능동적인 시선은 보이지 않습니다.",
        "built_space": "왼쪽에 루버가 달린 문 1개와 벽등 1개, 뒤쪽에 접시 안테나 1개와 좌우 확성기 2개를 갖춘 통신탑 1개, 오른쪽에 헬리콥터 1대가 있습니다. 외곽벽과 도시 배경도 유지됩니다. 다만 참고의 넓은 잔디와 뒤집힌 폭발 지면 대부분이 평평한 포장면으로 바뀌었습니다. 현우는 전경의 작은 잔해 구역에 무릎을 꿇고 있으며 전신과 넉넉한 주변 공간이 보입니다. 바닥 반사에 명백한 모순은 없습니다. 어두운 하늘과 켜진 도시 조명 때문에 야간으로 읽혀 명시적인 낮 시간 잠금에는 맞지 않습니다.",
        "entities": "젊고 마른 동아시아계 남성 1명이 현우로 표현되며 검은 머리와 얼굴·팔의 부상이 보입니다. 흰 바지는 참고와 유사하지만 상의는 짧게 찢어진 소매 형태이고 참고의 긴소매 및 대각선 천 띠가 재현되지 않습니다. 찰리는 흰 마스크형 머리와 심하게 탄 부분 몸통으로 남아 있으며 온전한 긴 팔과 다리는 없습니다. 장갑판은 참고보다 검게 소실되어 있으나 파괴된 상태라는 설명에는 부합합니다. 추가 사람이나 광선, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "현우의 양 무릎과 접힌 하체가 지면에 닿아 체중을 받습니다. 찰리의 몸통은 양팔과 가슴 사이에 끼워져 지지되고 아래쪽은 현우의 허벅지 사이로 내려옵니다. 머리는 남은 목 구조에 연결되어 있으며 전선은 아래로 늘어집니다. 주변 파편도 지면에 놓여 있어 지지 없는 부유나 명백히 불가능한 자세는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.194,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.194,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1194
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "돋아난 잔디 화단을 포함한 장소의 아키텍처를 정확히 구현(Priority 3)했으며, 로봇 잔해의 거대한 스케일을 현실적으로 유지한 채 오열하는 전신의 감정 표현을 완벽하게 담아냈습니다."
   },
   {
    "label": "A",
    "score": 1194,
    "verdict_ko": "레퍼런스의 주요 건축 지형인 잔디 화단을 평지로 왜곡하여 장소 일치(Priority 3) 지침을 위반했으며, 거대한 로봇 잔해를 장난감처럼 비현실적으로 작게 축소해 개체 묘사가 부실합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B02.png",
    "asset_id": "76fc6513-8b50-4f33-b05a-840ac5bb52b7",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1337888>",
    "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1299563>",
    "asset_id": "2948628a-764c-4354-a3e9-c21d7ef85260",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-2621-73c2-936e-6f3ab34e5ff5",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh66__bgfirst_bg.png",
   "bg_asset_id": "bf0ba719-2d49-4122-ad48-c1530b3f79c4",
   "bg_record_key": "S89sh66::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S89sh66::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:19:24.016655+00:00",
  "fingerprint": "a90f54a939e4e36f55dd94152541825cdeff392e7b832eaf6e0028aad337031a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S89sh66_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S89sh66_sel.png",
  "source_sha256": "d4884951b02cf287921bce381371adcdd52955f50b7dc29e465cc314bffeb70d",
  "file": "S89sh66_cine.png",
  "staged_sha256": "c5bf92db6dae7c8b1f6053106ef9092cb14531cb40a82003b25f4780735210d1",
  "latency_ms": 10181
 },
 "S90sh4::signage": {
  "fp": "f04fc0d37e0e4402",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::fad8ad38683873bf": {
  "subjects": [],
  "subject_text": "찰리의 부품 보관과 복원이 이루어지는 제주도 연구실\n전체적으로 어두운 넓은 연구실. 대형 모니터 아래 제단처럼 생긴 단상 테이블과 수술대가 있고, 책상에는 조작 버튼이 달려 있다.",
  "identity": "canonical",
  "scope_id": "L282",
  "scope_role": "location_interior",
  "scope_sha": "c219c6865b7d60f4"
 },
 "S90sh4::bgfirst_bg": {
  "input_fingerprint": "7dba88376e5c1b9e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4__bgfirst_bg.png",
  "asset_id": "d20cc5fd-c3cf-45a0-bbb5-940c2dade504",
  "input_asset_ids": [
   "9d655ae0-d34a-4dc7-bb3a-f3c76225f96b",
   "020c6829-306f-4816-9550-7878ddde52fa"
  ]
 },
 "S90sh4": {
  "input_fingerprint": "0e021faacb003d38",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large room is dim, with disassembled Charlie components arranged on an altar-like table. The large monitor shows Charlie energy units supplied to locations including Dubai, Europe, China and Africa. 현우: He is seated behind the display table, still covered in wounds. 지소영: She is beside the seated area, holding out the small data chip containing Charlie's stored memories.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 지소영 right now, so 지소영's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 지소영: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large room is dim, with disassembled Charlie components arranged on an altar-like table. The large monitor shows Charlie energy units supplied to locations including Dubai, Europe, China and Africa. 현우: He is seated behind the display table, still covered in wounds. 지소영: She is beside the seated area, holding out the small data chip containing Charlie's stored memories.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 지소영 right now, so 지소영's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 지소영: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large room is dim, with disassembled Charlie components arranged on an altar-like table. The large monitor shows Charlie energy units supplied to locations including Dubai, Europe, China and Africa. 현우: He is seated behind the display table, still covered in wounds. 지소영: She is beside the seated area, holding out the small data chip containing Charlie's stored memories.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 지소영 right now, so 지소영's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 지소영: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4__bgfirst_bg.png",
     "asset_id": "d20cc5fd-c3cf-45a0-bbb5-940c2dade504",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S90sh4.png",
     "asset_id": "9d655ae0-d34a-4dc7-bb3a-f3c76225f96b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1337888>",
     "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:703308>",
     "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L282B02.png",
     "asset_id": "020c6829-306f-4816-9550-7878ddde52fa",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1337888>",
     "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:703308>",
     "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "지소영의 시선은 현우를 향하고, 오른손은 현우의 왼쪽 어깨를 감싸며, 왼손으로 데이터 칩을 카메라와 현우 방향으로 내밀고 있습니다. 현우는 지소영 쪽을 약간 올려다봅니다.",
    "built_space": "어두운 연구실 내부로, 배경에 글씨가 뭉개져 판독할 수 없는 대형 디스플레이가 위치합니다. 우측에는 로봇 부품이 놓인 금속 테이블이 있으며, 현우는 등받이가 보이는 의자에 앉아 있습니다.",
    "entities": "지소영은 참조 이미지의 얼굴과 흰색 재킷을 잘 반영했습니다. 현우는 상처와 헤어스타일이 일치하지만, 참조 이미지의 찢어진 환자복 대신 평범한 흰색 티셔츠를 입고 있습니다. 데이터 칩은 형태가 잘 유지된 채 쥐어져 있습니다.",
    "hard_violations": [
     "[gpt-high] 기준 장소의 중앙 부품 작업대가 뒤쪽 작업대와 오른쪽 전경 작업대로 증설되고 유사한 로봇 머리도 두 개로 중복되어, 고정된 장소의 설비와 소품 구성을 바꾼다."
    ],
    "physics": "지소영은 서서 오른손을 현우의 어깨에 지탱하고 있으며, 현우는 의자에 앉아 체중을 싣고 있습니다. 데이터 칩은 지소영의 왼쪽 손가락들 사이에 단단히 고정되어 공중에 떠 있지 않습니다."
   },
   {
    "label": "B",
    "direction": "지소영이 현우를 내려다보며 오른손을 현우의 오른쪽 어깨에 얹고, 왼손으로 데이터 칩을 현우의 시선 쪽으로 내밀고 있습니다. 현우는 지소영의 얼굴을 응시합니다.",
    "built_space": "연구실 내부 배경의 대형 디스플레이에 'GLOBAL ENERGY SUPPLY'를 비롯한 영문 텍스트가 명확하게 보입니다. 테이블 위에 부품들이 있고, 현우는 의자에 앉아 있습니다.",
    "entities": "지소영은 스마트워치와 재킷 등 참조 이미지와 완벽히 일치합니다. 현우 역시 찢어진 옷과 얼굴 상처 등 참조 이미지의 특성을 정확히 재현했습니다. 데이터 칩도 명확히 묘사되었습니다.",
    "hard_violations": [
     "[gemini-pro] 화면에 명확히 읽히는 텍스트(GLOBAL ENERGY SUPPLY 등) 생성으로 인한 부정 프롬프트(No readable writing) 위반",
     "[gpt-high] 배경 대형 화면 상단의 영문 제목을 읽을 수 있어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "두 인물 모두 바닥과 의자에 물리적으로 자연스럽게 지탱하여 자세를 취하고 있으며, 들고 있는 데이터 칩 역시 손가락에 올바르게 쥐어져 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지소영이 현우의 어깨를 감싸고 칩을 내미는 앵글과 텍스트 비식별화 지침을 훌륭히 준수했으나, 현우의 찢어진 의상 디테일이 생략된 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "인물들의 외형과 찢어진 복장은 참조와 매우 일치하지만, 배경 디스플레이에 텍스트가 선명하게 생성되어 글씨 배제 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영의 시선은 현우를 향하고, 오른손은 현우의 왼쪽 어깨를 감싸며, 왼손으로 데이터 칩을 카메라와 현우 방향으로 내밀고 있습니다. 현우는 지소영 쪽을 약간 올려다봅니다.",
        "built_space": "어두운 연구실 내부로, 배경에 글씨가 뭉개져 판독할 수 없는 대형 디스플레이가 위치합니다. 우측에는 로봇 부품이 놓인 금속 테이블이 있으며, 현우는 등받이가 보이는 의자에 앉아 있습니다.",
        "entities": "지소영은 참조 이미지의 얼굴과 흰색 재킷을 잘 반영했습니다. 현우는 상처와 헤어스타일이 일치하지만, 참조 이미지의 찢어진 환자복 대신 평범한 흰색 티셔츠를 입고 있습니다. 데이터 칩은 형태가 잘 유지된 채 쥐어져 있습니다.",
        "hard_violations": [],
        "physics": "지소영은 서서 오른손을 현우의 어깨에 지탱하고 있으며, 현우는 의자에 앉아 체중을 싣고 있습니다. 데이터 칩은 지소영의 왼쪽 손가락들 사이에 단단히 고정되어 공중에 떠 있지 않습니다."
       },
       {
        "label": "B",
        "direction": "지소영이 현우를 내려다보며 오른손을 현우의 오른쪽 어깨에 얹고, 왼손으로 데이터 칩을 현우의 시선 쪽으로 내밀고 있습니다. 현우는 지소영의 얼굴을 응시합니다.",
        "built_space": "연구실 내부 배경의 대형 디스플레이에 'GLOBAL ENERGY SUPPLY'를 비롯한 영문 텍스트가 명확하게 보입니다. 테이블 위에 부품들이 있고, 현우는 의자에 앉아 있습니다.",
        "entities": "지소영은 스마트워치와 재킷 등 참조 이미지와 완벽히 일치합니다. 현우 역시 찢어진 옷과 얼굴 상처 등 참조 이미지의 특성을 정확히 재현했습니다. 데이터 칩도 명확히 묘사되었습니다.",
        "hard_violations": [
         "화면에 명확히 읽히는 텍스트(GLOBAL ENERGY SUPPLY 등) 생성으로 인한 부정 프롬프트(No readable writing) 위반"
        ],
        "physics": "두 인물 모두 바닥과 의자에 물리적으로 자연스럽게 지탱하여 자세를 취하고 있으며, 들고 있는 데이터 칩 역시 손가락에 올바르게 쥐어져 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지소영이 현우의 어깨를 감싸고 칩을 내미는 앵글과 텍스트 비식별화 지침을 훌륭히 준수했으나, 현우의 찢어진 의상 디테일이 생략된 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "인물들의 외형과 찢어진 복장은 참조와 매우 일치하지만, 배경 디스플레이에 텍스트가 선명하게 생성되어 글씨 배제 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "지소영의 시선은 현우를 향하고, 오른손은 현우의 왼쪽 어깨를 감싸며, 왼손으로 데이터 칩을 카메라와 현우 방향으로 내밀고 있습니다. 현우는 지소영 쪽을 약간 올려다봅니다.",
        "built_space": "어두운 연구실 내부로, 배경에 글씨가 뭉개져 판독할 수 없는 대형 디스플레이가 위치합니다. 우측에는 로봇 부품이 놓인 금속 테이블이 있으며, 현우는 등받이가 보이는 의자에 앉아 있습니다.",
        "entities": "지소영은 참조 이미지의 얼굴과 흰색 재킷을 잘 반영했습니다. 현우는 상처와 헤어스타일이 일치하지만, 참조 이미지의 찢어진 환자복 대신 평범한 흰색 티셔츠를 입고 있습니다. 데이터 칩은 형태가 잘 유지된 채 쥐어져 있습니다.",
        "hard_violations": [],
        "physics": "지소영은 서서 오른손을 현우의 어깨에 지탱하고 있으며, 현우는 의자에 앉아 체중을 싣고 있습니다. 데이터 칩은 지소영의 왼쪽 손가락들 사이에 단단히 고정되어 공중에 떠 있지 않습니다."
       },
       {
        "label": "B",
        "direction": "지소영이 현우를 내려다보며 오른손을 현우의 오른쪽 어깨에 얹고, 왼손으로 데이터 칩을 현우의 시선 쪽으로 내밀고 있습니다. 현우는 지소영의 얼굴을 응시합니다.",
        "built_space": "연구실 내부 배경의 대형 디스플레이에 'GLOBAL ENERGY SUPPLY'를 비롯한 영문 텍스트가 명확하게 보입니다. 테이블 위에 부품들이 있고, 현우는 의자에 앉아 있습니다.",
        "entities": "지소영은 스마트워치와 재킷 등 참조 이미지와 완벽히 일치합니다. 현우 역시 찢어진 옷과 얼굴 상처 등 참조 이미지의 특성을 정확히 재현했습니다. 데이터 칩도 명확히 묘사되었습니다.",
        "hard_violations": [
         "화면에 명확히 읽히는 텍스트(GLOBAL ENERGY SUPPLY 등) 생성으로 인한 부정 프롬프트(No readable writing) 위반"
        ],
        "physics": "두 인물 모두 바닥과 의자에 물리적으로 자연스럽게 지탱하여 자세를 취하고 있으며, 들고 있는 데이터 칩 역시 손가락에 올바르게 쥐어져 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "어깨 접촉과 상체 중심 구도는 맞지만, 칩을 현우보다 카메라 쪽으로 내밀고 배경 화면에 읽을 수 있는 영문 제목이 노출된다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우를 내려다보며 어깨를 감싸고 작은 칩을 건네기 직전인 동작은 더 정확하지만, 기준 장소에 없는 두 번째 부품 작업대와 중복된 로봇 머리가 장소 고정을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 왼쪽 아래의 현우 얼굴을 보고, 현우는 지소영을 올려다본다. 한 손은 현우의 어깨에 닿지만 칩을 든 다른 손은 현우가 있는 왼쪽보다 카메라 앞으로 뻗어 있다. 칩의 넓은 면도 카메라 쪽으로 드러나, 두 사람 사이에서 비스듬히 제시되는 전달 동작과 차이가 있다.",
        "built_space": "중앙 금속 부품 작업대 한 개, 뒤쪽 대형 벽면 화면 한 개, 양쪽 수납장과 천장 배관이 보인다. 현우 몸통 뒤 왼쪽에 검은 의자의 등받이 일부가 있고, 지소영은 그 옆에 서서 몸을 기울인다. 기준 장소의 주요 설비 관계는 대체로 유지된다. 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 두 명뿐이다. 현우는 헝클어진 검은 머리의 앳된 동아시아계 남성으로, 얼굴과 어깨의 상처 및 찢어진 흰옷이 기준에 부합한다. 지소영은 정돈된 짧은 검은 머리의 중년 동아시아계 여성이고 흰색 높은 깃 의상도 기준과 가깝다. 작은 직사각형 칩은 지소영의 손에 있다. 작업대에는 금속 관절과 로봇 머리가 놓여 있다. 화면에는 세계 지도와 그래프가 있으나 배치된 로봇 자체는 뚜렷하지 않으며, 상단 영문 제목은 읽을 수 있다.",
        "hard_violations": [
         "배경 대형 화면 상단의 영문 제목을 읽을 수 있어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "현우의 착석은 몸통 뒤 의자와 자연스럽게 연결된다. 지소영의 하체 지지는 화면 밖이지만 상체 기울기와 팔의 연결은 가능한 자세다. 어깨 위 손은 실제 접촉하고, 칩은 다른 손의 손가락 사이에 잡혀 있다. 로봇 부품은 작업대 위에 받쳐져 있으며 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "지소영은 앞에 앉은 현우를 내려다보고, 현우는 고개를 들어 지소영을 바라본다. 지소영의 한 손이 현우의 어깨를 감싸며, 다른 손은 두 사람 사이로 작은 칩을 내민다. 칩은 손가락 사이에서 비스듬하게 보여 전달 전 순간의 방향 관계가 A보다 정확하다.",
        "built_space": "왼쪽 출입문 한 개, 뒤쪽 대형 화면 한 개, 오른쪽 야경 창과 유리 수납장이 보인다. 현우 뒤와 아래에는 점유된 검은 의자의 일부가 자연스럽게 드러난다. 다만 뒤쪽 가로 작업대 외에 오른쪽 전경으로 뻗은 별도의 부품 작업대가 생겼고, 두 작업대에 각각 유사한 로봇 머리가 놓여 있어 기준 사진의 단일 중앙 작업대 구성이 바뀌었다.",
        "entities": "두 인물만 등장한다. 현우는 검은 헝클어진 머리, 젊은 남성의 체격, 옆얼굴 상처와 피 묻은 흰옷으로 묘사된다. 얼굴 대부분은 뒤돌아 있어 기준 인물과의 세부 일치는 제한적으로만 확인된다. 지소영의 중년 얼굴, 단정한 검은 단발과 흰색 제복은 기준에 가깝다. 칩은 작고 지소영 자신의 손에 잡혀 있으며 읽을 만한 표기는 없다. 배경 화면은 세계 지도와 그래프를 보여주지만 배치된 에너지 로봇은 뚜렷하지 않다. 로봇 머리는 뒤쪽과 오른쪽 전경에 하나씩, 총 두 개가 보인다.",
        "hard_violations": [
         "기준 장소의 중앙 부품 작업대가 뒤쪽 작업대와 오른쪽 전경 작업대로 증설되고 유사한 로봇 머리도 두 개로 중복되어, 고정된 장소의 설비와 소품 구성을 바꾼다."
        ],
        "physics": "현우의 등과 몸통은 검은 의자의 착석 방향에 맞고, 지소영은 의자 옆에서 서서 접근하는 자세다. 어깨를 감싼 손과 칩을 집은 손 모두 팔에 자연스럽게 연결된다. 칩은 손가락으로 지지되고 로봇 부품은 작업대에 놓여 있다. 공중에 지지 없이 뜬 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "어깨 접촉과 상체 중심 구도는 맞지만, 칩을 현우보다 카메라 쪽으로 내밀고 배경 화면에 읽을 수 있는 영문 제목이 노출된다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우를 내려다보며 어깨를 감싸고 작은 칩을 건네기 직전인 동작은 더 정확하지만, 기준 장소에 없는 두 번째 부품 작업대와 중복된 로봇 머리가 장소 고정을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "지소영은 왼쪽 아래의 현우 얼굴을 보고, 현우는 지소영을 올려다본다. 한 손은 현우의 어깨에 닿지만 칩을 든 다른 손은 현우가 있는 왼쪽보다 카메라 앞으로 뻗어 있다. 칩의 넓은 면도 카메라 쪽으로 드러나, 두 사람 사이에서 비스듬히 제시되는 전달 동작과 차이가 있다.",
        "built_space": "중앙 금속 부품 작업대 한 개, 뒤쪽 대형 벽면 화면 한 개, 양쪽 수납장과 천장 배관이 보인다. 현우 몸통 뒤 왼쪽에 검은 의자의 등받이 일부가 있고, 지소영은 그 옆에 서서 몸을 기울인다. 기준 장소의 주요 설비 관계는 대체로 유지된다. 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 두 명뿐이다. 현우는 헝클어진 검은 머리의 앳된 동아시아계 남성으로, 얼굴과 어깨의 상처 및 찢어진 흰옷이 기준에 부합한다. 지소영은 정돈된 짧은 검은 머리의 중년 동아시아계 여성이고 흰색 높은 깃 의상도 기준과 가깝다. 작은 직사각형 칩은 지소영의 손에 있다. 작업대에는 금속 관절과 로봇 머리가 놓여 있다. 화면에는 세계 지도와 그래프가 있으나 배치된 로봇 자체는 뚜렷하지 않으며, 상단 영문 제목은 읽을 수 있다.",
        "hard_violations": [
         "배경 대형 화면 상단의 영문 제목을 읽을 수 있어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "현우의 착석은 몸통 뒤 의자와 자연스럽게 연결된다. 지소영의 하체 지지는 화면 밖이지만 상체 기울기와 팔의 연결은 가능한 자세다. 어깨 위 손은 실제 접촉하고, 칩은 다른 손의 손가락 사이에 잡혀 있다. 로봇 부품은 작업대 위에 받쳐져 있으며 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "지소영은 앞에 앉은 현우를 내려다보고, 현우는 고개를 들어 지소영을 바라본다. 지소영의 한 손이 현우의 어깨를 감싸며, 다른 손은 두 사람 사이로 작은 칩을 내민다. 칩은 손가락 사이에서 비스듬하게 보여 전달 전 순간의 방향 관계가 A보다 정확하다.",
        "built_space": "왼쪽 출입문 한 개, 뒤쪽 대형 화면 한 개, 오른쪽 야경 창과 유리 수납장이 보인다. 현우 뒤와 아래에는 점유된 검은 의자의 일부가 자연스럽게 드러난다. 다만 뒤쪽 가로 작업대 외에 오른쪽 전경으로 뻗은 별도의 부품 작업대가 생겼고, 두 작업대에 각각 유사한 로봇 머리가 놓여 있어 기준 사진의 단일 중앙 작업대 구성이 바뀌었다.",
        "entities": "두 인물만 등장한다. 현우는 검은 헝클어진 머리, 젊은 남성의 체격, 옆얼굴 상처와 피 묻은 흰옷으로 묘사된다. 얼굴 대부분은 뒤돌아 있어 기준 인물과의 세부 일치는 제한적으로만 확인된다. 지소영의 중년 얼굴, 단정한 검은 단발과 흰색 제복은 기준에 가깝다. 칩은 작고 지소영 자신의 손에 잡혀 있으며 읽을 만한 표기는 없다. 배경 화면은 세계 지도와 그래프를 보여주지만 배치된 에너지 로봇은 뚜렷하지 않다. 로봇 머리는 뒤쪽과 오른쪽 전경에 하나씩, 총 두 개가 보인다.",
        "hard_violations": [
         "기준 장소의 중앙 부품 작업대가 뒤쪽 작업대와 오른쪽 전경 작업대로 증설되고 유사한 로봇 머리도 두 개로 중복되어, 고정된 장소의 설비와 소품 구성을 바꾼다."
        ],
        "physics": "현우의 등과 몸통은 검은 의자의 착석 방향에 맞고, 지소영은 의자 옆에서 서서 접근하는 자세다. 어깨를 감싼 손과 칩을 집은 손 모두 팔에 자연스럽게 연결된다. 칩은 손가락으로 지지되고 로봇 부품은 작업대에 놓여 있다. 공중에 지지 없이 뜬 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면에 명확히 읽히는 텍스트(GLOBAL ENERGY SUPPLY 등) 생성으로 인한 부정 프롬프트(No readable writing) 위반",
     "[gpt-high] 배경 대형 화면 상단의 영문 제목을 읽을 수 있어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 기준 장소의 중앙 부품 작업대가 뒤쪽 작업대와 오른쪽 전경 작업대로 증설되고 유사한 로봇 머리도 두 개로 중복되어, 고정된 장소의 설비와 소품 구성을 바꾼다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1125
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지소영이 현우의 어깨를 감싸고 칩을 내미는 앵글과 텍스트 비식별화 지침을 훌륭히 준수했으나, 현우의 찢어진 의상 디테일이 생략된 점이 아쉽습니다.  ★위반: [gpt-high] 기준 장소의 중앙 부품 작업대가 뒤쪽 작업대와 오른쪽 전경 작업대로 증설되고 유사한 로봇 머리도 두 개로 중복되어, 고정된 장소의 설비와 소품 구성을 바꾼다."
   },
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "인물들의 외형과 찢어진 복장은 참조와 매우 일치하지만, 배경 디스플레이에 텍스트가 선명하게 생성되어 글씨 배제 지침을 위반했습니다.  ★위반: [gemini-pro] 화면에 명확히 읽히는 텍스트(GLOBAL ENERGY SUPPLY 등) 생성으로 인한 부정 프롬프트(No readable writing) 위반 / [gpt-high] 배경 대형 화면 상단의 영문 제목을 읽을 수 있어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L282B02.png",
    "asset_id": "020c6829-306f-4816-9550-7878ddde52fa",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1337888>",
    "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:703308>",
    "asset_id": "b45e2bfc-4cdb-41ff-8548-4db79b6d994e",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-2976-7a0d-af9e-13538c910907",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4__bgfirst_bg.png",
   "bg_asset_id": "d20cc5fd-c3cf-45a0-bbb5-940c2dade504",
   "bg_record_key": "S90sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S90sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:21:00.124224+00:00",
  "fingerprint": "3ffdcff8d1f86b787f5a434b3cd0e9d67be3860ab98ff101cb5aafedabd71131",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S90sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S90sh4_sel.png",
  "source_sha256": "db4609bb61800c98c8518e82f6b1bc4b1a02cdf531ecea129c8d82ce74d6c73f",
  "file": "S90sh4_cine.png",
  "staged_sha256": "d78c1365ca2ab2eaff1fb2bb2f182c83476ceb82096cecfb7683ac00fb1ee920",
  "latency_ms": 10573
 },
 "S90sh9::signage": {
  "fp": "b7733e579bf3cbc0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S90sh9": {
  "input_fingerprint": "715609c35d67fa61",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 연구소 허공에 푸른 제주도 풍경과 함께 찰리의 거대한 홀로그램 영상이 띄워진 찰나.\n\nLOCATION (lock): In the open projection area of the dark restoration laboratory, illuminated by a large island-scene display and a robot hologram. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large laboratory screen (Displaying the blue Jeju landscape) — The image-bearing face is visible obliquely behind the hologram, with its boundaries retained; used as Supplies the island setting as displayed imagery rather than a literal change of location; the screen occupies less than two-fifths of the frame; 현우's seat (Still occupied before he rises) — A small rear portion appears beside his shoulder at the lower edge; used as Keeps the viewer anchored in the physical laboratory while the projected scene appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The blue Jeju imagery and luminous holographic appearance remain distinct against the generally dark laboratory, without implying a blue light source elsewhere in the room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The display has changed to a Jeju landscape and a holographic Charlie has appeared, while the real disassembled components remain on the table in the dim room. The holographic sequence includes Amber and Raul running toward Charlie. 현우: He is still covered in wounds and now has the data chip received during the conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (홀로그램) (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 연구소 허공에 푸른 제주도 풍경과 함께 찰리의 거대한 홀로그램 영상이 띄워진 찰나.\n\nLOCATION (lock): In the open projection area of the dark restoration laboratory, illuminated by a large island-scene display and a robot hologram. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large laboratory screen (Displaying the blue Jeju landscape) — The image-bearing face is visible obliquely behind the hologram, with its boundaries retained; used as Supplies the island setting as displayed imagery rather than a literal change of location; the screen occupies less than two-fifths of the frame; 현우's seat (Still occupied before he rises) — A small rear portion appears beside his shoulder at the lower edge; used as Keeps the viewer anchored in the physical laboratory while the projected scene appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The blue Jeju imagery and luminous holographic appearance remain distinct against the generally dark laboratory, without implying a blue light source elsewhere in the room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The display has changed to a Jeju landscape and a holographic Charlie has appeared, while the real disassembled components remain on the table in the dim room. The holographic sequence includes Amber and Raul running toward Charlie. 현우: He is still covered in wounds and now has the data chip received during the conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (홀로그램) (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 연구소 허공에 푸른 제주도 풍경과 함께 찰리의 거대한 홀로그램 영상이 띄워진 찰나.\n\nLOCATION (lock): In the open projection area of the dark restoration laboratory, illuminated by a large island-scene display and a robot hologram. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large laboratory screen (Displaying the blue Jeju landscape) — The image-bearing face is visible obliquely behind the hologram, with its boundaries retained; used as Supplies the island setting as displayed imagery rather than a literal change of location; the screen occupies less than two-fifths of the frame; 현우's seat (Still occupied before he rises) — A small rear portion appears beside his shoulder at the lower edge; used as Keeps the viewer anchored in the physical laboratory while the projected scene appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The blue Jeju imagery and luminous holographic appearance remain distinct against the generally dark laboratory, without implying a blue light source elsewhere in the room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The display has changed to a Jeju landscape and a holographic Charlie has appeared, while the real disassembled components remain on the table in the dim room. The holographic sequence includes Amber and Raul running toward Charlie. 현우: He is still covered in wounds and now has the data chip received during the conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (홀로그램) (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴); 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 현우의 등 뒤에서 테이블 위 홀로그램과 모니터를 정면으로 바라보고 있습니다.",
    "built_space": "이전 샷에 존재하던 벽면 고정형 대형 스크린이 사라지고, 그 자리에 문과 벽이 보이며 대신 이동식 스탠드 모니터가 공간 한가운데 임의로 추가되어 연구소의 기존 구조가 왜곡되었습니다.",
    "entities": "찰리 홀로그램과 그를 향해 달려가는 작은 두 인물이 묘사되었습니다. 현우는 뒷모습만 보이며 손이 프레임에 잡히지 않아 데이터 칩의 유무를 확인할 수 없습니다.",
    "hard_violations": [
     "[gemini-pro] 공간 구조 위반: 기준 이미지에 명시된 고정형 벽면 스크린을 삭제하고 임의의 이동식 스크린을 배치함",
     "[gpt-high] 이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다.",
     "[gpt-high] 이전 장면의 뒤쪽 벽면 대형 화면을 왼쪽 전경의 바퀴 달린 이동식 화면으로 바꾸어, 고정된 장소 설비의 연속성을 깨뜨렸다."
    ],
    "physics": "홀로그램이 테이블 위 공간에 떠서 영사되고 있으며 특별한 물리적 오류는 없습니다."
   },
   {
    "label": "B",
    "direction": "카메라가 현우의 좌측 후방에서 홀로그램, 벽면 스크린, 그리고 현우가 쥐고 있는 데이터 칩을 향해 시선을 둡니다.",
    "built_space": "벽면 대형 스크린, 금속 테이블, 부품들의 배치 등 연구소 내부의 고정 요소가 기준 이미지와 정확히 일치합니다. 다만 우측 창밖 야경에 프롬프트의 제주도 콘셉트가 일부 반영되어 산(성산일출봉)의 실루엣이 나타납니다.",
    "entities": "찰리 홀로그램과 테이블 위를 달리는 두 인물이 명확히 구현되었습니다. 현우는 목의 반창고와 피 묻은 셔츠 등 기준 이미지의 외형을 잘 유지하고 있으며, 요구된 데이터 칩을 쥐고 있습니다.",
    "hard_violations": [
     "[gpt-high] 이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다."
    ],
    "physics": "홀로그램이 부품이 놓인 테이블 위로 자연스럽게 투사되고 있으며, 현우의 오른손이 데이터 칩을 안정적으로 지지하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 넓은 화각, 홀로그램과 벽면 스크린의 정확한 배치, 현우가 데이터 칩을 들고 있는 디테일까지 완벽하게 구현했으나, 창밖 배경에 제주도의 지형이 일부 스며든 점이 유일한 흠입니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "기준 이미지의 고정 요소인 벽면 대형 스크린을 없애고 이동식 모니터를 임의로 추가하여 공간의 연속성을 훼손하는 구조적 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 현우의 등 뒤에서 테이블 위 홀로그램과 모니터를 정면으로 바라보고 있습니다.",
        "built_space": "이전 샷에 존재하던 벽면 고정형 대형 스크린이 사라지고, 그 자리에 문과 벽이 보이며 대신 이동식 스탠드 모니터가 공간 한가운데 임의로 추가되어 연구소의 기존 구조가 왜곡되었습니다.",
        "entities": "찰리 홀로그램과 그를 향해 달려가는 작은 두 인물이 묘사되었습니다. 현우는 뒷모습만 보이며 손이 프레임에 잡히지 않아 데이터 칩의 유무를 확인할 수 없습니다.",
        "hard_violations": [
         "공간 구조 위반: 기준 이미지에 명시된 고정형 벽면 스크린을 삭제하고 임의의 이동식 스크린을 배치함"
        ],
        "physics": "홀로그램이 테이블 위 공간에 떠서 영사되고 있으며 특별한 물리적 오류는 없습니다."
       },
       {
        "label": "B",
        "direction": "카메라가 현우의 좌측 후방에서 홀로그램, 벽면 스크린, 그리고 현우가 쥐고 있는 데이터 칩을 향해 시선을 둡니다.",
        "built_space": "벽면 대형 스크린, 금속 테이블, 부품들의 배치 등 연구소 내부의 고정 요소가 기준 이미지와 정확히 일치합니다. 다만 우측 창밖 야경에 프롬프트의 제주도 콘셉트가 일부 반영되어 산(성산일출봉)의 실루엣이 나타납니다.",
        "entities": "찰리 홀로그램과 테이블 위를 달리는 두 인물이 명확히 구현되었습니다. 현우는 목의 반창고와 피 묻은 셔츠 등 기준 이미지의 외형을 잘 유지하고 있으며, 요구된 데이터 칩을 쥐고 있습니다.",
        "hard_violations": [],
        "physics": "홀로그램이 부품이 놓인 테이블 위로 자연스럽게 투사되고 있으며, 현우의 오른손이 데이터 칩을 안정적으로 지지하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 넓은 화각, 홀로그램과 벽면 스크린의 정확한 배치, 현우가 데이터 칩을 들고 있는 디테일까지 완벽하게 구현했으나, 창밖 배경에 제주도의 지형이 일부 스며든 점이 유일한 흠입니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "기준 이미지의 고정 요소인 벽면 대형 스크린을 없애고 이동식 모니터를 임의로 추가하여 공간의 연속성을 훼손하는 구조적 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 현우의 등 뒤에서 테이블 위 홀로그램과 모니터를 정면으로 바라보고 있습니다.",
        "built_space": "이전 샷에 존재하던 벽면 고정형 대형 스크린이 사라지고, 그 자리에 문과 벽이 보이며 대신 이동식 스탠드 모니터가 공간 한가운데 임의로 추가되어 연구소의 기존 구조가 왜곡되었습니다.",
        "entities": "찰리 홀로그램과 그를 향해 달려가는 작은 두 인물이 묘사되었습니다. 현우는 뒷모습만 보이며 손이 프레임에 잡히지 않아 데이터 칩의 유무를 확인할 수 없습니다.",
        "hard_violations": [
         "공간 구조 위반: 기준 이미지에 명시된 고정형 벽면 스크린을 삭제하고 임의의 이동식 스크린을 배치함"
        ],
        "physics": "홀로그램이 테이블 위 공간에 떠서 영사되고 있으며 특별한 물리적 오류는 없습니다."
       },
       {
        "label": "B",
        "direction": "카메라가 현우의 좌측 후방에서 홀로그램, 벽면 스크린, 그리고 현우가 쥐고 있는 데이터 칩을 향해 시선을 둡니다.",
        "built_space": "벽면 대형 스크린, 금속 테이블, 부품들의 배치 등 연구소 내부의 고정 요소가 기준 이미지와 정확히 일치합니다. 다만 우측 창밖 야경에 프롬프트의 제주도 콘셉트가 일부 반영되어 산(성산일출봉)의 실루엣이 나타납니다.",
        "entities": "찰리 홀로그램과 테이블 위를 달리는 두 인물이 명확히 구현되었습니다. 현우는 목의 반창고와 피 묻은 셔츠 등 기준 이미지의 외형을 잘 유지하고 있으며, 요구된 데이터 칩을 쥐고 있습니다.",
        "hard_violations": [],
        "physics": "홀로그램이 부품이 놓인 테이블 위로 자연스럽게 투사되고 있으며, 현우의 오른손이 데이터 칩을 안정적으로 지지하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "금지된 추가 인물 두 명 때문에 부적격이지만, 기존 벽면 화면과 연구소 배치를 보존해 B보다 연속성이 높다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "거대한 찰리와 비스듬한 화면은 요구에 가깝지만, 추가 인물 두 명과 기존 화면을 이동식 화면으로 바꾼 공간 변경이 결정적이다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 앞쪽 위의 찰리를 바라보고, 찰리의 얼굴은 현우가 있는 왼쪽 아래를 향한다. 작은 투영 인물 두 명은 얼굴과 몸 앞면을 카메라 쪽으로 보이며 달려오므로, 찰리를 향해 달려가는 관계는 명확하지 않다. 제주 영상의 표시 면은 현우와 카메라를 향하지만, 요구한 사선보다는 거의 정면으로 보인다.",
        "built_space": "왼쪽 출입문 하나, 뒤쪽 벽면 대형 화면 하나, 오른쪽 야경 창, 유리 장비장과 금속 작업대가 이전 장면의 배치를 유지한다. 화면 경계가 남아 있고 면적은 프레임의 2/5 미만이다. 현우는 왼쪽 아래 의자에 앉아 있으나 어깨와 등받이가 상당히 크게 보여, 하단에 등받이 일부만 작게 보이라는 요구에는 덜 맞는다. 오른쪽 창에는 제주 화면의 반사로 보이는 영상이 있으며, 측벽 유리에 뒤쪽 화면이 반사되는 것 자체는 불가능하다고 단정할 수 없다.",
        "entities": "찰리는 긴 중량감 있는 팔, 짧은 다리, 각진 베이지 장갑과 흰 마스크형 얼굴을 갖춰 참조와 대체로 일치한다. 현우는 검은 헝클어진 머리, 상처, 이전 장면의 얼룩진 밝은 상의를 유지하며 젊은 동아시아계 남성으로 보인다. 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵다. 손에는 작은 칩이 있고, 작업대에는 분해된 구형 관절과 축 부품이 남아 있다. 제주 해안 영상과 밤의 연구소가 보인다. 다만 허용 인물 목록 밖의 작은 사람 형상 두 명이 추가되어 있다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다."
        ],
        "physics": "현우의 몸은 의자에 지지되고, 칩은 오른손 손가락 사이에 잡혀 있다. 작업대의 부품은 상판에 놓여 있다. 찰리와 작은 인물들은 발광하는 투영 영상이므로 공중에 보이는 것을 물리적 신체의 무지지 부유로 판정하지 않는다. 칩 부근에서 찰리 쪽으로 빛줄기가 퍼지지만, 칩이 투사 장치라는 설정은 명시되지 않았다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 중앙에서 찰리를 올려다본다. 찰리의 얼굴은 화면 오른쪽을 향해 있어 현우와 직접 시선을 맞추지는 않는다. 작은 투영 인물 두 명은 대체로 등을 보인 채 찰리 쪽으로 달리는 방향이다. 제주 화면은 찰리 뒤에서 비스듬히 표시 면을 드러내며 현우가 볼 수 있는 방향을 향한다.",
        "built_space": "왼쪽 출입문 하나, 오른쪽 야경 창과 장비장, 부품이 놓인 금속 작업대는 유지된다. 그러나 이전 장면의 뒤쪽 벽면 화면 대신, 왼쪽 앞에 바퀴 달린 받침대가 있는 대형 화면이 서 있다. 이는 단순한 카메라 각도 변화가 아니라 고정 설비의 위치와 설치 방식 변경이다. 화면 경계와 사선 방향은 분명하고 면적도 2/5 미만이다. 현우는 중앙 하단 의자에 앉아 있으나 등받이의 넓은 부분이 노출된다.",
        "entities": "찰리의 흰 마스크형 얼굴, 베이지 장갑, 큰 어깨와 긴 팔은 참조의 주요 특징에 부합한다. 현우는 뒤통수와 상체만 보여 얼굴이나 정확한 나이는 확인하기 어렵지만, 검은 머리와 상처 난 목, 얼룩진 밝은 상의는 이어진다. 칩은 손이 프레임에 드러나지 않아 확인할 수 없다. 제주도는 해안 풍경보다 섬 전체를 내려다본 영상으로 표현된다. 실제 분해 부품은 작업대에 남아 있다. 허용되지 않은 작은 인간형 투영 인물 두 명이 있으며, 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다.",
         "이전 장면의 뒤쪽 벽면 대형 화면을 왼쪽 전경의 바퀴 달린 이동식 화면으로 바꾸어, 고정된 장소 설비의 연속성을 깨뜨렸다."
        ],
        "physics": "현우는 등받이가 뒤에 있는 의자에 정상적으로 앉아 있다. 작업대의 분해 부품은 상판에 지지되고, 대형 화면은 바닥의 바퀴 달린 받침대에 지지된다. 찰리와 두 작은 인물은 푸른 발광 투영으로 표현되어 물리적 신체가 근거 없이 떠 있는 사례는 아니다. 찰리의 하체 일부는 현우에게 가려져 접지 여부를 평가할 필요가 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "금지된 추가 인물 두 명 때문에 부적격이지만, 기존 벽면 화면과 연구소 배치를 보존해 B보다 연속성이 높다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "거대한 찰리와 비스듬한 화면은 요구에 가깝지만, 추가 인물 두 명과 기존 화면을 이동식 화면으로 바꾼 공간 변경이 결정적이다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 앞쪽 위의 찰리를 바라보고, 찰리의 얼굴은 현우가 있는 왼쪽 아래를 향한다. 작은 투영 인물 두 명은 얼굴과 몸 앞면을 카메라 쪽으로 보이며 달려오므로, 찰리를 향해 달려가는 관계는 명확하지 않다. 제주 영상의 표시 면은 현우와 카메라를 향하지만, 요구한 사선보다는 거의 정면으로 보인다.",
        "built_space": "왼쪽 출입문 하나, 뒤쪽 벽면 대형 화면 하나, 오른쪽 야경 창, 유리 장비장과 금속 작업대가 이전 장면의 배치를 유지한다. 화면 경계가 남아 있고 면적은 프레임의 2/5 미만이다. 현우는 왼쪽 아래 의자에 앉아 있으나 어깨와 등받이가 상당히 크게 보여, 하단에 등받이 일부만 작게 보이라는 요구에는 덜 맞는다. 오른쪽 창에는 제주 화면의 반사로 보이는 영상이 있으며, 측벽 유리에 뒤쪽 화면이 반사되는 것 자체는 불가능하다고 단정할 수 없다.",
        "entities": "찰리는 긴 중량감 있는 팔, 짧은 다리, 각진 베이지 장갑과 흰 마스크형 얼굴을 갖춰 참조와 대체로 일치한다. 현우는 검은 헝클어진 머리, 상처, 이전 장면의 얼룩진 밝은 상의를 유지하며 젊은 동아시아계 남성으로 보인다. 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵다. 손에는 작은 칩이 있고, 작업대에는 분해된 구형 관절과 축 부품이 남아 있다. 제주 해안 영상과 밤의 연구소가 보인다. 다만 허용 인물 목록 밖의 작은 사람 형상 두 명이 추가되어 있다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다."
        ],
        "physics": "현우의 몸은 의자에 지지되고, 칩은 오른손 손가락 사이에 잡혀 있다. 작업대의 부품은 상판에 놓여 있다. 찰리와 작은 인물들은 발광하는 투영 영상이므로 공중에 보이는 것을 물리적 신체의 무지지 부유로 판정하지 않는다. 칩 부근에서 찰리 쪽으로 빛줄기가 퍼지지만, 칩이 투사 장치라는 설정은 명시되지 않았다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 중앙에서 찰리를 올려다본다. 찰리의 얼굴은 화면 오른쪽을 향해 있어 현우와 직접 시선을 맞추지는 않는다. 작은 투영 인물 두 명은 대체로 등을 보인 채 찰리 쪽으로 달리는 방향이다. 제주 화면은 찰리 뒤에서 비스듬히 표시 면을 드러내며 현우가 볼 수 있는 방향을 향한다.",
        "built_space": "왼쪽 출입문 하나, 오른쪽 야경 창과 장비장, 부품이 놓인 금속 작업대는 유지된다. 그러나 이전 장면의 뒤쪽 벽면 화면 대신, 왼쪽 앞에 바퀴 달린 받침대가 있는 대형 화면이 서 있다. 이는 단순한 카메라 각도 변화가 아니라 고정 설비의 위치와 설치 방식 변경이다. 화면 경계와 사선 방향은 분명하고 면적도 2/5 미만이다. 현우는 중앙 하단 의자에 앉아 있으나 등받이의 넓은 부분이 노출된다.",
        "entities": "찰리의 흰 마스크형 얼굴, 베이지 장갑, 큰 어깨와 긴 팔은 참조의 주요 특징에 부합한다. 현우는 뒤통수와 상체만 보여 얼굴이나 정확한 나이는 확인하기 어렵지만, 검은 머리와 상처 난 목, 얼룩진 밝은 상의는 이어진다. 칩은 손이 프레임에 드러나지 않아 확인할 수 없다. 제주도는 해안 풍경보다 섬 전체를 내려다본 영상으로 표현된다. 실제 분해 부품은 작업대에 남아 있다. 허용되지 않은 작은 인간형 투영 인물 두 명이 있으며, 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다.",
         "이전 장면의 뒤쪽 벽면 대형 화면을 왼쪽 전경의 바퀴 달린 이동식 화면으로 바꾸어, 고정된 장소 설비의 연속성을 깨뜨렸다."
        ],
        "physics": "현우는 등받이가 뒤에 있는 의자에 정상적으로 앉아 있다. 작업대의 분해 부품은 상판에 지지되고, 대형 화면은 바닥의 바퀴 달린 받침대에 지지된다. 찰리와 두 작은 인물은 푸른 발광 투영으로 표현되어 물리적 신체가 근거 없이 떠 있는 사례는 아니다. 찰리의 하체 일부는 현우에게 가려져 접지 여부를 평가할 필요가 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.194,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.944,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 공간 구조 위반: 기준 이미지에 명시된 고정형 벽면 스크린을 삭제하고 임의의 이동식 스크린을 배치함",
     "[gpt-high] 이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다.",
     "[gpt-high] 이전 장면의 뒤쪽 벽면 대형 화면을 왼쪽 전경의 바퀴 달린 이동식 화면으로 바꾸어, 고정된 장소 설비의 연속성을 깨뜨렸다."
    ],
    "B": [
     "[gpt-high] 이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 944
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "프롬프트가 요구한 넓은 화각, 홀로그램과 벽면 스크린의 정확한 배치, 현우가 데이터 칩을 들고 있는 디테일까지 완벽하게 구현했으나, 창밖 배경에 제주도의 지형이 일부 스며든 점이 유일한 흠입니다.  ★위반: [gpt-high] 이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다."
   },
   {
    "label": "A",
    "score": 944,
    "verdict_ko": "기준 이미지의 고정 요소인 벽면 대형 스크린을 없애고 이동식 모니터를 임의로 추가하여 공간의 연속성을 훼손하는 구조적 위반이 발생했습니다.  ★위반: [gemini-pro] 공간 구조 위반: 기준 이미지에 명시된 고정형 벽면 스크린을 삭제하고 임의의 이동식 스크린을 배치함 / [gpt-high] 이 숏에서 허용한 찰리와 현우 외에, 작은 인간형 투영 인물 두 명이 추가되어 있다. / [gpt-high] 이전 장면의 뒤쪽 벽면 대형 화면을 왼쪽 전경의 바퀴 달린 이동식 화면으로 바꾸어, 고정된 장소 설비의 연속성을 깨뜨렸다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4_sel.png",
    "asset_id": "7e616658-b999-4512-ac7f-7f57fea882c3",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리 (홀로그램): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1375583>",
    "asset_id": "4aa97496-776a-4ae1-9f9c-10a2d484db6f",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1337888>",
    "asset_id": "9092b1a1-0def-4012-8bce-0c7786d5e5c3",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-2cd5-71b4-83dc-a5e306a350bf",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S90sh4"
  },
  "staged_characters_added": [
   "C01"
  ]
 },
 "S90sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:50:01.752903+00:00",
  "fingerprint": "8cd74973d0dffb728df2f857d0d55f5b140a36ba2f1836c4404c9a25904468cf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S90sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S90sh9_sel.png",
  "source_sha256": "bf65b76a57907d1a378f4bdf8d1f2d8db7b266365e7873de8c900edfca28671f",
  "file": "S90sh9_cine.png",
  "staged_sha256": "773cc223a59cf0ccddb4a6e7d785b57f4863629ec91ac7cd28f934645d748a23",
  "latency_ms": 10959
 },
 "S90sh13::signage": {
  "fp": "3dfd262870e4321d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S90sh13": {
  "input_fingerprint": "422f4e0b10576dc2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차가운 수술대 위, 부서진 부품들이 서로 빈틈없이 맞물려 온전한 형태를 갖춘 찰리의 낡은 전신.\n\nLOCATION (lock): On the restoration table inside the spacious research room, with the surrounding displays illuminating the reassembled robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating table (Supporting Charlie's fully fitted body) — The upper supporting surface and foot-side corner are visible from the high oblique position; used as Provides a clear body-length reference while keeping the surrounding laboratory visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued illumination appropriate to the dark laboratory gently separates the fitted body sections without suggesting an activation flash or dream effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the operating table with his broken components fitted back together into a complete body and his eyes closed. The source does not specify his head's direction, which side of his torso faces upward, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's broken components are being fitted back together on the operating table, with the reconstructed eyes closed. The holographic Jeju presentation remains separate from the physical reconstruction; no return to active operation has yet been established.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차가운 수술대 위, 부서진 부품들이 서로 빈틈없이 맞물려 온전한 형태를 갖춘 찰리의 낡은 전신.\n\nLOCATION (lock): On the restoration table inside the spacious research room, with the surrounding displays illuminating the reassembled robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating table (Supporting Charlie's fully fitted body) — The upper supporting surface and foot-side corner are visible from the high oblique position; used as Provides a clear body-length reference while keeping the surrounding laboratory visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued illumination appropriate to the dark laboratory gently separates the fitted body sections without suggesting an activation flash or dream effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the operating table with his broken components fitted back together into a complete body and his eyes closed. The source does not specify his head's direction, which side of his torso faces upward, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's broken components are being fitted back together on the operating table, with the reconstructed eyes closed. The holographic Jeju presentation remains separate from the physical reconstruction; no return to active operation has yet been established.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차가운 수술대 위, 부서진 부품들이 서로 빈틈없이 맞물려 온전한 형태를 갖춘 찰리의 낡은 전신.\n\nLOCATION (lock): On the restoration table inside the spacious research room, with the surrounding displays illuminating the reassembled robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating table (Supporting Charlie's fully fitted body) — The upper supporting surface and foot-side corner are visible from the high oblique position; used as Provides a clear body-length reference while keeping the surrounding laboratory visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued illumination appropriate to the dark laboratory gently separates the fitted body sections without suggesting an activation flash or dream effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the operating table with his broken components fitted back together into a complete body and his eyes closed. The source does not specify his head's direction, which side of his torso faces upward, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's broken components are being fitted back together on the operating table, with the reconstructed eyes closed. The holographic Jeju presentation remains separate from the physical reconstruction; no return to active operation has yet been established.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 찰리가 수술대 위에 누워있는 모습을 높은 사각(high oblique)에서 내려다보고 있으며, 찰리의 머리는 화면 좌측 상단을 향하고 있다.",
    "built_space": "연구실 내부에 수술대와 각종 기기들이 배치되어 있으나, 레퍼런스의 핵심 배경 요소인 중앙의 대형 풍경 스크린과 오른쪽 창밖 풍경이 다른 형태의 모니터와 어두운 벽면으로 대체되어 있다.",
    "entities": "고릴라형 몸체와 마스크형 얼굴을 가진 찰리의 외형이 레퍼런스와 정확히 일치하며 조립된 채로 누워 있다. 프롬프트 지시대로 다른 인물은 없다.",
    "hard_violations": [],
    "physics": "찰리의 육중한 기계 몸체는 수술대 표면에 평평하게 밀착되어 중력에 맞게 완전히 지탱되고 있다."
   },
   {
    "label": "B",
    "direction": "카메라는 수술대 위에 누워있는 찰리를 거의 눈높이에 가까운 측면 수평 앵글에서 바라보고 있다.",
    "built_space": "왼쪽의 문, 중앙 벽면의 대형 풍경 스크린, 오른쪽의 창밖 풍경 등 레퍼런스에서 설정된 연구실의 고정된 공간 특징과 구조를 정확히 유지하고 있다.",
    "entities": "찰리의 몸체는 레퍼런스에 맞게 조립되어 누워 있으나, 그의 몸 위와 옆으로 지시되지 않은 푸른빛의 홀로그램 스크린들이 나타나 있다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트가 명시적으로 금지한 홀로그램 스크린과 다이어그램을 임의로 생성함 (Invented objects, leaked diagrams/text)."
    ],
    "physics": "찰리의 몸과 팔다리는 수술대 위에 닿아 구조적으로 잘 지탱되고 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스의 배경 요소(대형 스크린)가 일부 누락되었으나, 지시된 높은 사각(high oblique) 구도와 비활성 상태의 조명 분위기를 완벽하게 구현하여 샷 텍스트의 연출 의도를 가장 잘 따른 결과물입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경 공간은 레퍼런스와 정확히 일치하나, 요구된 카메라 앵글을 완전히 무시했으며 명시적으로 배제된 홀로그램 인터페이스를 추가하는 치명적인 규칙 위반을 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 찰리가 수술대 위에 누워있는 모습을 높은 사각(high oblique)에서 내려다보고 있으며, 찰리의 머리는 화면 좌측 상단을 향하고 있다.",
        "built_space": "연구실 내부에 수술대와 각종 기기들이 배치되어 있으나, 레퍼런스의 핵심 배경 요소인 중앙의 대형 풍경 스크린과 오른쪽 창밖 풍경이 다른 형태의 모니터와 어두운 벽면으로 대체되어 있다.",
        "entities": "고릴라형 몸체와 마스크형 얼굴을 가진 찰리의 외형이 레퍼런스와 정확히 일치하며 조립된 채로 누워 있다. 프롬프트 지시대로 다른 인물은 없다.",
        "hard_violations": [],
        "physics": "찰리의 육중한 기계 몸체는 수술대 표면에 평평하게 밀착되어 중력에 맞게 완전히 지탱되고 있다."
       },
       {
        "label": "B",
        "direction": "카메라는 수술대 위에 누워있는 찰리를 거의 눈높이에 가까운 측면 수평 앵글에서 바라보고 있다.",
        "built_space": "왼쪽의 문, 중앙 벽면의 대형 풍경 스크린, 오른쪽의 창밖 풍경 등 레퍼런스에서 설정된 연구실의 고정된 공간 특징과 구조를 정확히 유지하고 있다.",
        "entities": "찰리의 몸체는 레퍼런스에 맞게 조립되어 누워 있으나, 그의 몸 위와 옆으로 지시되지 않은 푸른빛의 홀로그램 스크린들이 나타나 있다.",
        "hard_violations": [
         "프롬프트가 명시적으로 금지한 홀로그램 스크린과 다이어그램을 임의로 생성함 (Invented objects, leaked diagrams/text)."
        ],
        "physics": "찰리의 몸과 팔다리는 수술대 위에 닿아 구조적으로 잘 지탱되고 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스의 배경 요소(대형 스크린)가 일부 누락되었으나, 지시된 높은 사각(high oblique) 구도와 비활성 상태의 조명 분위기를 완벽하게 구현하여 샷 텍스트의 연출 의도를 가장 잘 따른 결과물입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경 공간은 레퍼런스와 정확히 일치하나, 요구된 카메라 앵글을 완전히 무시했으며 명시적으로 배제된 홀로그램 인터페이스를 추가하는 치명적인 규칙 위반을 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 찰리가 수술대 위에 누워있는 모습을 높은 사각(high oblique)에서 내려다보고 있으며, 찰리의 머리는 화면 좌측 상단을 향하고 있다.",
        "built_space": "연구실 내부에 수술대와 각종 기기들이 배치되어 있으나, 레퍼런스의 핵심 배경 요소인 중앙의 대형 풍경 스크린과 오른쪽 창밖 풍경이 다른 형태의 모니터와 어두운 벽면으로 대체되어 있다.",
        "entities": "고릴라형 몸체와 마스크형 얼굴을 가진 찰리의 외형이 레퍼런스와 정확히 일치하며 조립된 채로 누워 있다. 프롬프트 지시대로 다른 인물은 없다.",
        "hard_violations": [],
        "physics": "찰리의 육중한 기계 몸체는 수술대 표면에 평평하게 밀착되어 중력에 맞게 완전히 지탱되고 있다."
       },
       {
        "label": "B",
        "direction": "카메라는 수술대 위에 누워있는 찰리를 거의 눈높이에 가까운 측면 수평 앵글에서 바라보고 있다.",
        "built_space": "왼쪽의 문, 중앙 벽면의 대형 풍경 스크린, 오른쪽의 창밖 풍경 등 레퍼런스에서 설정된 연구실의 고정된 공간 특징과 구조를 정확히 유지하고 있다.",
        "entities": "찰리의 몸체는 레퍼런스에 맞게 조립되어 누워 있으나, 그의 몸 위와 옆으로 지시되지 않은 푸른빛의 홀로그램 스크린들이 나타나 있다.",
        "hard_violations": [
         "프롬프트가 명시적으로 금지한 홀로그램 스크린과 다이어그램을 임의로 생성함 (Invented objects, leaked diagrams/text)."
        ],
        "physics": "찰리의 몸과 팔다리는 수술대 위에 닿아 구조적으로 잘 지탱되고 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "전신과 수술대 상판·발치 모서리, 기존 연구실을 함께 담은 높은 사선 와이드 구도가 가장 충실하지만, 눈을 감았다는 표현은 불분명하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "복원된 찰리의 외형과 누운 자세는 충실하나, 더 조밀한 구도로 발치 수술대 모서리가 잘리고 기존 연구실의 공간 관계가 덜 드러난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 머리를 화면 오른쪽, 발을 왼쪽에 두고 등을 대고 누워 있다. 얼굴은 천장 쪽을 향하며 특정 대상을 응시하지 않는다. 보이는 눈구멍은 어둡지만 닫힌 눈꺼풀인지는 명확하지 않다. 무기나 이동 동작은 없다. 오른쪽 진단 화면들은 수술대 안쪽을 향해 비스듬히 배치되어 있다.",
        "built_space": "전경에 수술대 한 개, 그 뒤에 금속 작업대 한 개, 후벽에 대형 해안 영상 화면 한 개와 하부 작업 공간이 보인다. 왼쪽 문 한 개와 장비 선반, 오른쪽 유리문 장비장들과 큰 야간 창이 있어 이전 장면의 배치를 잘 이어간다. 수술대 상판과 왼쪽 발치 모서리가 모두 보인다. 오른쪽 창에 보이는 해안 화면의 반사는 맞은편 후벽 화면의 반사로 해석 가능한 위치다. 수술대 오른쪽에는 반투명 진단 패널 여러 개가 추가되어 있다.",
        "entities": "등장 몸체는 찰리 한 대뿐이며 다른 사람은 없다. 흰 각진 마스크, 모래색의 긁히고 닳은 장갑판, 원형 가슴 부품, 육중한 긴 팔과 짧은 다리가 캐릭터 참조와 대체로 일치한다. 머리·몸통·양팔·양다리가 연결된 완전한 몸체다. 차가운 금속 수술대, 주변 화면, 별도로 남아 있는 제주 해안 영상도 보인다. 화면의 미세한 표시는 있으나 명확히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "몸통과 골반은 수술대 위에 놓이고, 전경의 팔뚝과 손은 검은 상판에 내려앉아 있다. 다리는 장갑의 두께 때문에 굽어 보이지만 하퇴와 뒤꿈치 쪽이 상판에 닿아 지지된다. 머리는 목 구조와 등·어깨 조립체에 연결되어 있으며 독립적으로 공중에 떠 있지 않다. 뒤쪽 팔의 접촉점 일부는 몸에 가려져 있다. 몸 전체를 받치는 수술대와 하부 지지 구조가 보이며 도약이나 활성화 동작은 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 화면 왼쪽 위, 발은 오른쪽 아래를 향하고 얼굴과 가슴은 위를 향한다. 눈은 어두운 안구 구획으로 보이며 명확하게 감긴 상태인지는 판단하기 어렵다. 양발의 발끝은 위로 향하지만 이동 동작은 없다. 배경 모니터들은 작업 공간 쪽을 향하고 카메라에서도 앞면이 보이는 각도다.",
        "built_space": "중앙의 금속 수술대 한 개를 높은 사선에서 내려다본다. 상판은 충분히 보이지만 화면 아래쪽에서 발치의 가까운 모서리가 잘린다. 후벽 대형 해안 화면 한 개, 그 아래 작업대와 소형 화면들, 오른쪽 장비장과 야간 창, 왼쪽 장비 선반이 보인다. 오른쪽 별도 작업대에는 둥근 기계 부품 두 개와 축 부품이 놓여 있다. 이전 장면의 재료와 설비 종류는 유지하지만 중앙 작업대와 주변 설비의 관계는 A보다 확인하기 어렵다. 창에는 실내의 어두운 반사가 보이며 명백한 광학적 모순은 없다.",
        "entities": "찰리 한 대만 등장한다. 흰 마스크형 얼굴, 닳은 베이지 장갑, 원형 가슴 장치, 큰 어깨와 긴 팔, 짧고 두꺼운 다리가 참조와 잘 맞는다. 몸체는 전신이 연결된 상태이며 여분의 인물이나 중복 신체는 없다. 수술대와 주변 진단 모니터, 별도 배경의 제주 영상이 포함된다. 모니터에 작은 정보 무늬가 있으나 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반은 금속 상판에 받쳐지고 가까운 팔뚝과 손도 상판에 놓여 있다. 무릎는 장갑 구조에 따라 굽혀져 있으며 하퇴와 뒤꿈치가 상판에 닿는다. 머리는 상판보다 높지만 목과 두꺼운 등·어깨 조립체에 연결되어 있어 지지 없이 떠 있는 몸체는 아니다. 먼 쪽 팔의 손과 일부 접촉점은 몸통에 가려진다. 수술대 아래에는 큰 받침 구조가 보여 무거운 로봇을 지탱하는 관계가 성립한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "전신과 수술대 상판·발치 모서리, 기존 연구실을 함께 담은 높은 사선 와이드 구도가 가장 충실하지만, 눈을 감았다는 표현은 불분명하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "복원된 찰리의 외형과 누운 자세는 충실하나, 더 조밀한 구도로 발치 수술대 모서리가 잘리고 기존 연구실의 공간 관계가 덜 드러난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 머리를 화면 오른쪽, 발을 왼쪽에 두고 등을 대고 누워 있다. 얼굴은 천장 쪽을 향하며 특정 대상을 응시하지 않는다. 보이는 눈구멍은 어둡지만 닫힌 눈꺼풀인지는 명확하지 않다. 무기나 이동 동작은 없다. 오른쪽 진단 화면들은 수술대 안쪽을 향해 비스듬히 배치되어 있다.",
        "built_space": "전경에 수술대 한 개, 그 뒤에 금속 작업대 한 개, 후벽에 대형 해안 영상 화면 한 개와 하부 작업 공간이 보인다. 왼쪽 문 한 개와 장비 선반, 오른쪽 유리문 장비장들과 큰 야간 창이 있어 이전 장면의 배치를 잘 이어간다. 수술대 상판과 왼쪽 발치 모서리가 모두 보인다. 오른쪽 창에 보이는 해안 화면의 반사는 맞은편 후벽 화면의 반사로 해석 가능한 위치다. 수술대 오른쪽에는 반투명 진단 패널 여러 개가 추가되어 있다.",
        "entities": "등장 몸체는 찰리 한 대뿐이며 다른 사람은 없다. 흰 각진 마스크, 모래색의 긁히고 닳은 장갑판, 원형 가슴 부품, 육중한 긴 팔과 짧은 다리가 캐릭터 참조와 대체로 일치한다. 머리·몸통·양팔·양다리가 연결된 완전한 몸체다. 차가운 금속 수술대, 주변 화면, 별도로 남아 있는 제주 해안 영상도 보인다. 화면의 미세한 표시는 있으나 명확히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "몸통과 골반은 수술대 위에 놓이고, 전경의 팔뚝과 손은 검은 상판에 내려앉아 있다. 다리는 장갑의 두께 때문에 굽어 보이지만 하퇴와 뒤꿈치 쪽이 상판에 닿아 지지된다. 머리는 목 구조와 등·어깨 조립체에 연결되어 있으며 독립적으로 공중에 떠 있지 않다. 뒤쪽 팔의 접촉점 일부는 몸에 가려져 있다. 몸 전체를 받치는 수술대와 하부 지지 구조가 보이며 도약이나 활성화 동작은 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 머리는 화면 왼쪽 위, 발은 오른쪽 아래를 향하고 얼굴과 가슴은 위를 향한다. 눈은 어두운 안구 구획으로 보이며 명확하게 감긴 상태인지는 판단하기 어렵다. 양발의 발끝은 위로 향하지만 이동 동작은 없다. 배경 모니터들은 작업 공간 쪽을 향하고 카메라에서도 앞면이 보이는 각도다.",
        "built_space": "중앙의 금속 수술대 한 개를 높은 사선에서 내려다본다. 상판은 충분히 보이지만 화면 아래쪽에서 발치의 가까운 모서리가 잘린다. 후벽 대형 해안 화면 한 개, 그 아래 작업대와 소형 화면들, 오른쪽 장비장과 야간 창, 왼쪽 장비 선반이 보인다. 오른쪽 별도 작업대에는 둥근 기계 부품 두 개와 축 부품이 놓여 있다. 이전 장면의 재료와 설비 종류는 유지하지만 중앙 작업대와 주변 설비의 관계는 A보다 확인하기 어렵다. 창에는 실내의 어두운 반사가 보이며 명백한 광학적 모순은 없다.",
        "entities": "찰리 한 대만 등장한다. 흰 마스크형 얼굴, 닳은 베이지 장갑, 원형 가슴 장치, 큰 어깨와 긴 팔, 짧고 두꺼운 다리가 참조와 잘 맞는다. 몸체는 전신이 연결된 상태이며 여분의 인물이나 중복 신체는 없다. 수술대와 주변 진단 모니터, 별도 배경의 제주 영상이 포함된다. 모니터에 작은 정보 무늬가 있으나 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반은 금속 상판에 받쳐지고 가까운 팔뚝과 손도 상판에 놓여 있다. 무릎는 장갑 구조에 따라 굽혀져 있으며 하퇴와 뒤꿈치가 상판에 닿는다. 머리는 상판보다 높지만 목과 두꺼운 등·어깨 조립체에 연결되어 있어 지지 없이 떠 있는 몸체는 아니다. 먼 쪽 팔의 손과 일부 접촉점은 몸통에 가려진다. 수술대 아래에는 큰 받침 구조가 보여 무거운 로봇을 지탱하는 관계가 성립한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.875,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.875,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트가 명시적으로 금지한 홀로그램 스크린과 다이어그램을 임의로 생성함 (Invented objects, leaked diagrams/text)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1875,
   "B": 1125
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1875,
    "verdict_ko": "레퍼런스의 배경 요소(대형 스크린)가 일부 누락되었으나, 지시된 높은 사각(high oblique) 구도와 비활성 상태의 조명 분위기를 완벽하게 구현하여 샷 텍스트의 연출 의도를 가장 잘 따른 결과물입니다."
   },
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "배경 공간은 레퍼런스와 정확히 일치하나, 요구된 카메라 앵글을 완전히 무시했으며 명시적으로 배제된 홀로그램 인터페이스를 추가하는 치명적인 규칙 위반을 범했습니다.  ★위반: [gemini-pro] 프롬프트가 명시적으로 금지한 홀로그램 스크린과 다이어그램을 임의로 생성함 (Invented objects, leaked diagrams/text)."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh9_sel.png",
    "asset_id": "4a0f2894-0e61-4616-9615-725d3291ef9c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1461211>",
    "asset_id": "24d1c172-ad7f-4d9b-9990-330c659e9ca6",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-2ebb-7aba-9b68-cb3d3f755c47",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S90sh9"
  }
 },
 "S90sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:23:43.726887+00:00",
  "fingerprint": "a2837f5932cd1eebee6d20132bf491e9b90ccdc0580521d04bfe88c0d26fbcf3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S90sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S90sh13_sel.png",
  "source_sha256": "c02e6a54553ed2b20c6f66a2cc0f9ed2d9cb7f052464cf9c77f6d6f268726025",
  "file": "S90sh13_cine.png",
  "staged_sha256": "19485a48455061f4db738c1bc9f7205517d9094701fc0c4be12ca939f45a0a11",
  "latency_ms": 10398
 },
 "S91sh4::signage": {
  "fp": "db4e4da97b9b5332",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::986d272d2d1a6fd3": {
  "subjects": [],
  "subject_text": "에필로그 바위섬과 임시 식탁\n바다 위로 드러난 바위섬의 야외 공간. 임시 식탁 위에 여러 음식이 차려져 있으며 밝은 낮빛이 비친다.",
  "identity": "canonical",
  "scope_id": "L284",
  "scope_role": "location_exterior",
  "scope_sha": "27853240e279399c"
 },
 "groupbg::islet_picnic_shore": {
  "input_fingerprint": "e0eac67a9dac1ad9",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "islet_picnic_shore",
    "tags": [
     "S91sh4",
     "S91sh5"
    ]
   },
   "context_sig": "638e73bc2f63dae3"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n에필로그 바위섬과 임시 식탁: 푸른 바다를 배경으로 거친 암석 위에 하얀 식탁이 근사하게 차려진 피크닉 풍경. (특징: 바다와 맞닿은 거칠고 넓은 잿빛 섬바위 표면; 바위 위에 놓인 깨끗한 테이블과 접시, 풍성한 음식들; 사람 어깨 위에 사뿐히 내려앉은 작은 새; 따뜻하고 밝게 빛나는 주간 야외 채광)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 91. 에필로그 / 어느 바다 섬바위 – D 자막: 한 달 후\n- 현우는 지민과 식탁에 근사하게 음식을 차리고 있고.\n- 어디선가 날아 온 아기 새, 찰리의 어깨 위로 내려앉는다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n에필로그 바위섬과 임시 식탁: 푸른 바다를 배경으로 거친 암석 위에 하얀 식탁이 근사하게 차려진 피크닉 풍경. (특징: 바다와 맞닿은 거칠고 넓은 잿빛 섬바위 표면; 바위 위에 놓인 깨끗한 테이블과 접시, 풍성한 음식들; 사람 어깨 위에 사뿐히 내려앉은 작은 새; 따뜻하고 밝게 빛나는 주간 야외 채광)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 91. 에필로그 / 어느 바다 섬바위 – D 자막: 한 달 후\n- 현우는 지민과 식탁에 근사하게 음식을 차리고 있고.\n- 어디선가 날아 온 아기 새, 찰리의 어깨 위로 내려앉는다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_islet_picnic_shore_1b9adf.png",
  "asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae",
  "input_asset_ids": [
   "8048621a-1847-4b10-b7c7-7f2985a29c65"
  ],
  "origin_tag": "S91sh4",
  "place_text": "At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.",
  "origin_inputs": {
   "place_text": "At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.",
   "time_of_day_en": "day",
   "conti_asset_id": "8048621a-1847-4b10-b7c7-7f2985a29c65"
  }
 },
 "S91sh4::bgfirst_bg": {
  "input_fingerprint": "92cfe0c383842d43",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4__bgfirst_bg.png",
  "asset_id": "410404cd-7317-4419-a279-f0a70de5bf97",
  "input_asset_ids": [
   "8048621a-1847-4b10-b7c7-7f2985a29c65",
   "080c8b58-a1b3-40f1-8b38-c1b52f00ffae"
  ]
 },
 "S91sh4": {
  "input_fingerprint": "8955b11647f14b67",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Clear seawater reveals fish beneath the rocky island, and an outdoor dining table is being laid with an ample meal. Charlie is fully reassembled and active again, trying to catch fish. 현우: He is at the dining table arranging food. 서지민: She is at the dining table arranging food.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Clear seawater reveals fish beneath the rocky island, and an outdoor dining table is being laid with an ample meal. Charlie is fully reassembled and active again, trying to catch fish. 현우: He is at the dining table arranging food. 서지민: She is at the dining table arranging food.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Clear seawater reveals fish beneath the rocky island, and an outdoor dining table is being laid with an ample meal. Charlie is fully reassembled and active again, trying to catch fish. 현우: He is at the dining table arranging food. 서지민: She is at the dining table arranging food.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4__bgfirst_bg.png",
     "asset_id": "410404cd-7317-4419-a279-f0a70de5bf97",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S91sh4.png",
     "asset_id": "8048621a-1847-4b10-b7c7-7f2985a29c65",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:930556>",
     "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_islet_picnic_shore_1b9adf.png",
     "asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:917012>",
     "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:930556>",
     "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우와 서지민은 서로를 바라보며 미소 짓고 있고 찰리는 물가를 향함.",
    "built_space": "바위섬 위 임시 식탁 양옆에 두 인물이 마주 서 있으며 식탁 다리 형태가 레퍼런스와 다소 다름.",
    "entities": "현우와 서지민의 외형 특징이 정확히 묘사되었고 지시된 로봇 개 찰리도 배치됨.",
    "hard_violations": [],
    "physics": "모든 인물과 찰리의 다리, 식탁 그리고 식기류가 바닥과 손에 의해 안정적으로 지지됨."
   },
   {
    "label": "B",
    "direction": "두 인물이 식탁을 사이에 두고 서로를 눈을 맞추며 웃고 있음.",
    "built_space": "바위섬 위 X자형 다리를 가진 식탁과 두 인물이 자연스러운 간격으로 배치됨.",
    "entities": "현우와 서지민의 인물 특징은 잘 반영되었으나 찰리가 묘사되지 않음.",
    "hard_violations": [],
    "physics": "인물의 양발, 식탁, 그리고 각자 들고 있는 식기가 어색함 없이 물리적으로 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 낚시 행동 지시까지 빠짐없이 반영하여 전반적인 프롬프트 구현도가 가장 높습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "인물과 공간의 묘사는 준수하나, 구체적으로 지시된 찰리가 완전히 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 서지민은 서로를 바라보며 미소 짓고 있고 찰리는 물가를 향함.",
        "built_space": "바위섬 위 임시 식탁 양옆에 두 인물이 마주 서 있으며 식탁 다리 형태가 레퍼런스와 다소 다름.",
        "entities": "현우와 서지민의 외형 특징이 정확히 묘사되었고 지시된 로봇 개 찰리도 배치됨.",
        "hard_violations": [],
        "physics": "모든 인물과 찰리의 다리, 식탁 그리고 식기류가 바닥과 손에 의해 안정적으로 지지됨."
       },
       {
        "label": "B",
        "direction": "두 인물이 식탁을 사이에 두고 서로를 눈을 맞추며 웃고 있음.",
        "built_space": "바위섬 위 X자형 다리를 가진 식탁과 두 인물이 자연스러운 간격으로 배치됨.",
        "entities": "현우와 서지민의 인물 특징은 잘 반영되었으나 찰리가 묘사되지 않음.",
        "hard_violations": [],
        "physics": "인물의 양발, 식탁, 그리고 각자 들고 있는 식기가 어색함 없이 물리적으로 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 낚시 행동 지시까지 빠짐없이 반영하여 전반적인 프롬프트 구현도가 가장 높습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "인물과 공간의 묘사는 준수하나, 구체적으로 지시된 찰리가 완전히 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 서지민은 서로를 바라보며 미소 짓고 있고 찰리는 물가를 향함.",
        "built_space": "바위섬 위 임시 식탁 양옆에 두 인물이 마주 서 있으며 식탁 다리 형태가 레퍼런스와 다소 다름.",
        "entities": "현우와 서지민의 외형 특징이 정확히 묘사되었고 지시된 로봇 개 찰리도 배치됨.",
        "hard_violations": [],
        "physics": "모든 인물과 찰리의 다리, 식탁 그리고 식기류가 바닥과 손에 의해 안정적으로 지지됨."
       },
       {
        "label": "B",
        "direction": "두 인물이 식탁을 사이에 두고 서로를 눈을 맞추며 웃고 있음.",
        "built_space": "바위섬 위 X자형 다리를 가진 식탁과 두 인물이 자연스러운 간격으로 배치됨.",
        "entities": "현우와 서지민의 인물 특징은 잘 반영되었으나 찰리가 묘사되지 않음.",
        "hard_violations": [],
        "physics": "인물의 양발, 식탁, 그리고 각자 들고 있는 식기가 어색함 없이 물리적으로 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "서로 웃으며 음식을 놓는 동작과 사선으로 보이는 식탁은 충실하지만, 인물이 크게 잡혀 화면 아래에서 신발 일부가 잘리므로 핵심인 전신 와이드 숏에 미달한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "두 사람의 발끝까지 포함한 와이드 숏으로 서로를 향한 환한 웃음과 상차림을 구현하고, 바위 지형과 임시 식탁도 참조에 더 가깝게 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 서지민의 얼굴을 보고 웃고, 서지민도 현우의 얼굴을 향해 웃는다. 현우가 든 음식 그릇과 서지민이 든 장식 접시는 식탁 위로 향하며, 접시 안쪽이 위를 향해 카메라에도 자연스럽게 보인다.",
        "built_space": "노출된 해안 바위 위에 흰 임시 식탁 한 개가 있고 의자나 별도 가구는 없다. 상판의 깊이와 가까운 테두리가 두 사람 사이에 보이며, 하부에는 교차형 지지대가 보이지만 바닥 접점 일부는 화면 밖이다. 현우는 식탁 왼쪽, 서지민은 오른쪽에 서 있다. 바다와 오른쪽 먼 바위섬은 장소 참조와 잘 맞지만 식탁 지지대의 형태와 배치는 참조와 다르다.",
        "entities": "사람은 젊은 동아시아계 남녀 두 명뿐이다. 현우의 헝클어진 검은 머리, 회색 셔츠, 녹색 카고 바지와 갈색 부츠는 참조에 부합한다. 서지민의 검은 중간 길이 머리와 흰 작업복, 허리 장비, 흰 신발도 대체로 맞지만 참조의 머리 위 고글과 장갑은 보이지 않는다. 음식 여러 접시와 샐러드, 잔, 작은 꽃병이 보이며, 서지민은 무늬 있는 접시를 들고 있다. 찰리와 수중 물고기는 이 구도에서 확인되지 않는다. 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 바위에 발을 딛고 상체를 조금 기울인다. 현우의 그릇은 두 손으로 받쳐지고 서지민의 접시도 양손으로 지지된다. 나머지 음식과 식기는 상판에 놓여 있다. 식탁은 교차 지지대로 받쳐지며, 화면 아래로 잘린 접점을 근거로 부유한다고 판단할 수는 없다. 명백하게 지지 없이 떠 있는 물체나 불가능한 자세는 없다."
       },
       {
        "label": "B",
        "direction": "현우와 서지민은 각각 상대의 얼굴을 향해 시선을 맞추고 환하게 웃는다. 두 사람의 손은 식탁 위 음식 그릇을 놓거나 정돈하는 방향이다. 오른쪽 물가의 작은 로봇은 몸과 앞쪽 팔다리를 수면 쪽으로 기울여 물속을 살피는 모습이다.",
        "built_space": "바위 위에 흰 직사각형 접이식 식탁 한 개가 있고, 좌우 두 쌍의 다리가 바위까지 이어진다. 추가 의자는 없다. 현우는 왼쪽, 서지민은 오른쪽에 서며 두 사람 모두 머리부터 신발까지 화면 안에 들어온다. 상판과 가까운 측면, 음식 내부가 내려다보이는 각도지만 식탁의 사선 정도는 A보다 약하다. 회색 바위와 바다, 오른쪽 먼 섬의 배치 및 식탁 형태는 장소 참조에 가깝다.",
        "entities": "젊은 동아시아계 남녀 두 명의 외모와 체격은 인물 설정에 대체로 부합한다. 현우는 검은 흐트러진 머리, 회색 셔츠, 녹색 카고 바지와 갈색 부츠를 착용한다. 서지민은 검은 머리, 흰 작업복과 장갑, 허리 장비 및 흰 신발을 착용하지만 머리 위 고글은 없다. 식탁에는 여러 음식 그릇, 생선 요리, 샐러드, 옅은 색 접시, 물병 한 개와 잔 두 개가 보인다. 오른쪽 로봇은 재조립 후 물고기를 잡으려는 찰리의 연속 상태에 부합하는 표현이다. 수중 물고기는 뚜렷하게 식별되지 않으며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 신발은 바위에 닿아 있고, 가볍게 앞으로 기울인 상차림 자세가 발의 지지와 연결된다. 음식 그릇에는 손이 닿아 있으며 다른 식기는 상판이 받친다. 식탁 다리도 바위에 닿는다. 물가 로봇은 뒤쪽 다리와 앞쪽 팔다리를 바위 및 물가 쪽으로 뻗어 몸을 지지하는 자세로, 공중에 매달린 상태로 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "서로 웃으며 음식을 놓는 동작과 사선으로 보이는 식탁은 충실하지만, 인물이 크게 잡혀 화면 아래에서 신발 일부가 잘리므로 핵심인 전신 와이드 숏에 미달한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "두 사람의 발끝까지 포함한 와이드 숏으로 서로를 향한 환한 웃음과 상차림을 구현하고, 바위 지형과 임시 식탁도 참조에 더 가깝게 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 서지민의 얼굴을 보고 웃고, 서지민도 현우의 얼굴을 향해 웃는다. 현우가 든 음식 그릇과 서지민이 든 장식 접시는 식탁 위로 향하며, 접시 안쪽이 위를 향해 카메라에도 자연스럽게 보인다.",
        "built_space": "노출된 해안 바위 위에 흰 임시 식탁 한 개가 있고 의자나 별도 가구는 없다. 상판의 깊이와 가까운 테두리가 두 사람 사이에 보이며, 하부에는 교차형 지지대가 보이지만 바닥 접점 일부는 화면 밖이다. 현우는 식탁 왼쪽, 서지민은 오른쪽에 서 있다. 바다와 오른쪽 먼 바위섬은 장소 참조와 잘 맞지만 식탁 지지대의 형태와 배치는 참조와 다르다.",
        "entities": "사람은 젊은 동아시아계 남녀 두 명뿐이다. 현우의 헝클어진 검은 머리, 회색 셔츠, 녹색 카고 바지와 갈색 부츠는 참조에 부합한다. 서지민의 검은 중간 길이 머리와 흰 작업복, 허리 장비, 흰 신발도 대체로 맞지만 참조의 머리 위 고글과 장갑은 보이지 않는다. 음식 여러 접시와 샐러드, 잔, 작은 꽃병이 보이며, 서지민은 무늬 있는 접시를 들고 있다. 찰리와 수중 물고기는 이 구도에서 확인되지 않는다. 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 바위에 발을 딛고 상체를 조금 기울인다. 현우의 그릇은 두 손으로 받쳐지고 서지민의 접시도 양손으로 지지된다. 나머지 음식과 식기는 상판에 놓여 있다. 식탁은 교차 지지대로 받쳐지며, 화면 아래로 잘린 접점을 근거로 부유한다고 판단할 수는 없다. 명백하게 지지 없이 떠 있는 물체나 불가능한 자세는 없다."
       },
       {
        "label": "A",
        "direction": "현우와 서지민은 각각 상대의 얼굴을 향해 시선을 맞추고 환하게 웃는다. 두 사람의 손은 식탁 위 음식 그릇을 놓거나 정돈하는 방향이다. 오른쪽 물가의 작은 로봇은 몸과 앞쪽 팔다리를 수면 쪽으로 기울여 물속을 살피는 모습이다.",
        "built_space": "바위 위에 흰 직사각형 접이식 식탁 한 개가 있고, 좌우 두 쌍의 다리가 바위까지 이어진다. 추가 의자는 없다. 현우는 왼쪽, 서지민은 오른쪽에 서며 두 사람 모두 머리부터 신발까지 화면 안에 들어온다. 상판과 가까운 측면, 음식 내부가 내려다보이는 각도지만 식탁의 사선 정도는 A보다 약하다. 회색 바위와 바다, 오른쪽 먼 섬의 배치 및 식탁 형태는 장소 참조에 가깝다.",
        "entities": "젊은 동아시아계 남녀 두 명의 외모와 체격은 인물 설정에 대체로 부합한다. 현우는 검은 흐트러진 머리, 회색 셔츠, 녹색 카고 바지와 갈색 부츠를 착용한다. 서지민은 검은 머리, 흰 작업복과 장갑, 허리 장비 및 흰 신발을 착용하지만 머리 위 고글은 없다. 식탁에는 여러 음식 그릇, 생선 요리, 샐러드, 옅은 색 접시, 물병 한 개와 잔 두 개가 보인다. 오른쪽 로봇은 재조립 후 물고기를 잡으려는 찰리의 연속 상태에 부합하는 표현이다. 수중 물고기는 뚜렷하게 식별되지 않으며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 신발은 바위에 닿아 있고, 가볍게 앞으로 기울인 상차림 자세가 발의 지지와 연결된다. 음식 그릇에는 손이 닿아 있으며 다른 식기는 상판이 받친다. 식탁 다리도 바위에 닿는다. 물가 로봇은 뒤쪽 다리와 앞쪽 팔다리를 바위 및 물가 쪽으로 뻗어 몸을 지지하는 자세로, 공중에 매달린 상태로 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.492
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.492
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1492
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "찰리의 낚시 행동 지시까지 빠짐없이 반영하여 전반적인 프롬프트 구현도가 가장 높습니다."
   },
   {
    "label": "B",
    "score": 1492,
    "verdict_ko": "인물과 공간의 묘사는 준수하나, 구체적으로 지시된 찰리가 완전히 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_islet_picnic_shore_1b9adf.png",
    "asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:930556>",
    "asset_id": "797e849f-7ceb-4e15-ae8a-f00466224dc8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-3065-7625-aebc-e4fa9a7dad3e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4__bgfirst_bg.png",
   "bg_asset_id": "410404cd-7317-4419-a279-f0a70de5bf97",
   "bg_record_key": "S91sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "islet_picnic_shore",
   "groupbg_asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S91sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:25:44.034283+00:00",
  "fingerprint": "a30e4140fa04eae79ecc93b752d69c097872eba2fe19433c09e4ee00ba8c52c0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S91sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S91sh4_sel.png",
  "source_sha256": "ffa44576961f67fdf64baa3a76d504844ff8df86f94ad344bd50608448af18ef",
  "file": "S91sh4_cine.png",
  "staged_sha256": "b4a9d4475fbb178ea672540140e09b4699156b8ca44d933f3238e6b9dec989c0",
  "latency_ms": 11641
 },
 "S91sh5::signage": {
  "fp": "b38697630282dd68",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S91sh5": {
  "input_fingerprint": "4a76ac8d3de8725d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 낡은 금속 어깨 위로 깃털이 부스스한 아기 새(B-200이 돌보던 새)가 가만히 내려앉은 클로즈업.\n\nLOCATION (lock): At the water's edge beside the small rocky islet, where the robot is fishing in clear daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sea below Charlie (Clear water containing fish) — Seen downward beyond the shoulder, with the water below rather than a reflected scene as the background; used as Provides softly resolved spatial context for Charlie's continued fishing attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight preserves fine feather and worn-metal detail with soft contrast, without introducing unsupported reflections or atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains fully reassembled and active, with the chest ring intact; a baby bird has settled on the metal shoulder. Fish remain visible in the clear seawater, and the meal is being arranged on the nearby table.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 낡은 금속 어깨 위로 깃털이 부스스한 아기 새(B-200이 돌보던 새)가 가만히 내려앉은 클로즈업.\n\nLOCATION (lock): At the water's edge beside the small rocky islet, where the robot is fishing in clear daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sea below Charlie (Clear water containing fish) — Seen downward beyond the shoulder, with the water below rather than a reflected scene as the background; used as Provides softly resolved spatial context for Charlie's continued fishing attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight preserves fine feather and worn-metal detail with soft contrast, without introducing unsupported reflections or atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains fully reassembled and active, with the chest ring intact; a baby bird has settled on the metal shoulder. Fish remain visible in the clear seawater, and the meal is being arranged on the nearby table.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 낡은 금속 어깨 위로 깃털이 부스스한 아기 새(B-200이 돌보던 새)가 가만히 내려앉은 클로즈업.\n\nLOCATION (lock): At the water's edge beside the small rocky islet, where the robot is fishing in clear daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sea below Charlie (Clear water containing fish) — Seen downward beyond the shoulder, with the water below rather than a reflected scene as the background; used as Provides softly resolved spatial context for Charlie's continued fishing attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight preserves fine feather and worn-metal detail with soft contrast, without introducing unsupported reflections or atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains fully reassembled and active, with the chest ring intact; a baby bird has settled on the metal shoulder. Fish remain visible in the clear seawater, and the meal is being arranged on the nearby table.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 금속 어깨 너머로 아래쪽 수면을 향해 있으며, 새는 화면 왼쪽을 향하고 있습니다.",
    "built_space": "바위 해안가 근처 맑은 바닷물이 보이며, 물속에 물고기들이 있습니다.",
    "entities": "레퍼런스와 일치하는 아기 새가 낡은 샌드 베이지색 금속 어깨 장갑판 위에 앉아 있으며, 하단에 낚싯줄을 쥔 로봇의 손이 보입니다.",
    "hard_violations": [],
    "physics": "새의 발이 금속 어깨 표면에 안정적으로 닿아 지지받고 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 로봇의 등 뒤에서 어깨와 바다를 수평적으로 바라보며, 새는 화면 오른쪽을 향하고 있습니다.",
    "built_space": "뒷배경으로 바위섬과 바다 수면이 넓게 펼쳐져 있습니다.",
    "entities": "아기 새는 어깨 장갑판 위에 있으나, 찰리의 머리와 목 뒷부분 디자인이 레퍼런스(모자, 흰 마스크형 얼굴 구조)와 다르게 임의의 기계 부품으로 묘사되었습니다.",
    "hard_violations": [
     "[gemini-pro] 지시된 카메라 뷰포인트(어깨 너머 아래쪽 수면을 내려다보는 구도) 위반"
    ],
    "physics": "새가 어깨 장갑판 위에 물리적으로 올바르게 서 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 하향식 클로즈업 앵글을 정확히 구현하였고, 어깨 너머 맑은 물속의 물고기와 낚싯줄을 쥐고 있는 로봇의 손을 보여주어 프롬프트의 상황을 완벽하게 포착했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "카메라 앵글이 지시된 하향식이 아닌 수평에 가깝고, 프레임에 포함된 로봇의 뒷통수와 목 구조가 캐릭터 레퍼런스와 일치하지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 금속 어깨 너머로 아래쪽 수면을 향해 있으며, 새는 화면 왼쪽을 향하고 있습니다.",
        "built_space": "바위 해안가 근처 맑은 바닷물이 보이며, 물속에 물고기들이 있습니다.",
        "entities": "레퍼런스와 일치하는 아기 새가 낡은 샌드 베이지색 금속 어깨 장갑판 위에 앉아 있으며, 하단에 낚싯줄을 쥔 로봇의 손이 보입니다.",
        "hard_violations": [],
        "physics": "새의 발이 금속 어깨 표면에 안정적으로 닿아 지지받고 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 로봇의 등 뒤에서 어깨와 바다를 수평적으로 바라보며, 새는 화면 오른쪽을 향하고 있습니다.",
        "built_space": "뒷배경으로 바위섬과 바다 수면이 넓게 펼쳐져 있습니다.",
        "entities": "아기 새는 어깨 장갑판 위에 있으나, 찰리의 머리와 목 뒷부분 디자인이 레퍼런스(모자, 흰 마스크형 얼굴 구조)와 다르게 임의의 기계 부품으로 묘사되었습니다.",
        "hard_violations": [
         "지시된 카메라 뷰포인트(어깨 너머 아래쪽 수면을 내려다보는 구도) 위반"
        ],
        "physics": "새가 어깨 장갑판 위에 물리적으로 올바르게 서 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 하향식 클로즈업 앵글을 정확히 구현하였고, 어깨 너머 맑은 물속의 물고기와 낚싯줄을 쥐고 있는 로봇의 손을 보여주어 프롬프트의 상황을 완벽하게 포착했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "카메라 앵글이 지시된 하향식이 아닌 수평에 가깝고, 프레임에 포함된 로봇의 뒷통수와 목 구조가 캐릭터 레퍼런스와 일치하지 않습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 금속 어깨 너머로 아래쪽 수면을 향해 있으며, 새는 화면 왼쪽을 향하고 있습니다.",
        "built_space": "바위 해안가 근처 맑은 바닷물이 보이며, 물속에 물고기들이 있습니다.",
        "entities": "레퍼런스와 일치하는 아기 새가 낡은 샌드 베이지색 금속 어깨 장갑판 위에 앉아 있으며, 하단에 낚싯줄을 쥔 로봇의 손이 보입니다.",
        "hard_violations": [],
        "physics": "새의 발이 금속 어깨 표면에 안정적으로 닿아 지지받고 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 로봇의 등 뒤에서 어깨와 바다를 수평적으로 바라보며, 새는 화면 오른쪽을 향하고 있습니다.",
        "built_space": "뒷배경으로 바위섬과 바다 수면이 넓게 펼쳐져 있습니다.",
        "entities": "아기 새는 어깨 장갑판 위에 있으나, 찰리의 머리와 목 뒷부분 디자인이 레퍼런스(모자, 흰 마스크형 얼굴 구조)와 다르게 임의의 기계 부품으로 묘사되었습니다.",
        "hard_violations": [
         "지시된 카메라 뷰포인트(어깨 너머 아래쪽 수면을 내려다보는 구도) 위반"
        ],
        "physics": "새가 어깨 장갑판 위에 물리적으로 올바르게 서 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "새가 낡은 어깨에 조용히 앉은 클로즈업과 어깨 너머 아래쪽의 맑은 바다를 충실히 구현하지만, 머리와 등 부품은 찰리 참조와 차이가 있다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "새의 착지와 금속 어깨 클로즈업은 맞지만, 수평선까지 펼친 배경과 높은 해안 절벽이 아래쪽 바다 중심의 구도 및 장소 연속성을 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "새의 부리와 눈은 화면 오른쪽 바다 쪽을 향한다. 찰리는 뒤통수가 보이며 얼굴은 물 쪽으로 돌아가 있지만, 실제 눈의 시선과 물고기를 내려다보는 각도는 확인되지 않는다. 카메라는 어깨 너머 아래의 물속을 내려다보며 수평선은 보이지 않는다.",
        "built_space": "인공 구조물이나 고정 설비는 보이지 않는다. 전경에 찰리의 등과 한쪽 어깨가 있고, 그 아래 배경에는 투명한 바닷물과 잠긴 돌들이 있으며 오른쪽 위에 물가 바위가 있다. 이전 장면의 바위 해안과 양립하는 구도다. 식탁은 클로즈업 밖이므로 누락으로 평가하지 않는다. 배경은 반사된 풍경이 아니라 실제 물속이다.",
        "entities": "찰리 한 대의 상체 일부와 아기 새 한 마리가 보이고 다른 인물은 없다. 새는 올리브갈색의 둥근 몸통, 작은 부리, 짧게 접힌 날개, 가는 발과 부스스한 솜털로 참조에 부합한다. 찰리의 샌드 베이지 어깨 장갑과 올리브색 외투는 맞지만 머리 외장과 등에 보이는 원형 부품은 참조에서 확인되는 형상과 차이가 있다. 얼굴과 가슴 고리는 보이지 않아 판정할 수 없다. 물속에는 흐릿한 물고기 형태들이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "새의 두 발과 발톱이 어깨 장갑의 윗면에 닿아 체중을 지탱하고, 날개를 접은 자세는 내려앉은 직후의 정지 순간으로 자연스럽다. 어깨는 관절과 몸통에 연결되어 있다. 찰리의 지면 접촉부는 화면 밖이며, 보이는 부분에 공중 부양이나 불가능한 지지는 없다."
       },
       {
        "label": "B",
        "direction": "새의 부리와 시선은 화면 왼쪽, 카메라 쪽으로 비스듬히 향한다. 찰리의 머리는 화면 밖이어서 낚시 대상에 대한 시선을 확인할 수 없다. 아래쪽에는 물속이 보이지만 위쪽에는 수평선과 하늘까지 들어와, 어깨 너머 아래 바다를 보는 지정보다 수평 방향의 해안 전망이 강조된다.",
        "built_space": "인공 구조물이나 고정 설비는 보이지 않는다. 전경 왼쪽을 어깨 장갑이 크게 차지하고, 오른쪽 아래에는 얕은 물과 돌, 위쪽에는 먼 바다와 수평선이 있다. 왼쪽 배경의 높은 절벽 해안은 이전 장면의 낮은 암반과 작은 바위섬에 비해 지형 연속성이 약하다. 식탁은 화면 밖이며 물속 배경에 불가능한 반사는 없다.",
        "entities": "아기 새 한 마리와 찰리의 어깨 및 손 일부가 보이며 추가 인물은 없다. 새의 색, 둥근 몸통, 짧은 날개, 가는 발과 솜털은 참조에 가깝다. 어깨는 마모된 샌드 베이지 금속이지만 겹친 장갑판과 긴 홈의 형상은 참조 어깨와 다르다. 얼굴, 모자, 가슴 고리는 프레임 밖이라 평가하지 않는다. 맑은 물속에는 여러 물고기가 보이고, 하단에는 가는 낚싯줄로 보이는 선과 로봇 손 일부가 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "새의 양발이 경사진 어깨판과 돌출된 홈 가장자리에 접촉해 몸을 지탱한다. 접힌 날개와 안정된 몸통은 조용히 내려앉은 상태로 가능하다. 어깨 아래에는 연결된 기계 관절이 보이고, 하단 손 부근을 가는 줄이 지나가지만 정확한 파지점은 프레임과 가림 때문에 불분명하다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "새가 낡은 어깨에 조용히 앉은 클로즈업과 어깨 너머 아래쪽의 맑은 바다를 충실히 구현하지만, 머리와 등 부품은 찰리 참조와 차이가 있다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "새의 착지와 금속 어깨 클로즈업은 맞지만, 수평선까지 펼친 배경과 높은 해안 절벽이 아래쪽 바다 중심의 구도 및 장소 연속성을 약화한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "새의 부리와 눈은 화면 오른쪽 바다 쪽을 향한다. 찰리는 뒤통수가 보이며 얼굴은 물 쪽으로 돌아가 있지만, 실제 눈의 시선과 물고기를 내려다보는 각도는 확인되지 않는다. 카메라는 어깨 너머 아래의 물속을 내려다보며 수평선은 보이지 않는다.",
        "built_space": "인공 구조물이나 고정 설비는 보이지 않는다. 전경에 찰리의 등과 한쪽 어깨가 있고, 그 아래 배경에는 투명한 바닷물과 잠긴 돌들이 있으며 오른쪽 위에 물가 바위가 있다. 이전 장면의 바위 해안과 양립하는 구도다. 식탁은 클로즈업 밖이므로 누락으로 평가하지 않는다. 배경은 반사된 풍경이 아니라 실제 물속이다.",
        "entities": "찰리 한 대의 상체 일부와 아기 새 한 마리가 보이고 다른 인물은 없다. 새는 올리브갈색의 둥근 몸통, 작은 부리, 짧게 접힌 날개, 가는 발과 부스스한 솜털로 참조에 부합한다. 찰리의 샌드 베이지 어깨 장갑과 올리브색 외투는 맞지만 머리 외장과 등에 보이는 원형 부품은 참조에서 확인되는 형상과 차이가 있다. 얼굴과 가슴 고리는 보이지 않아 판정할 수 없다. 물속에는 흐릿한 물고기 형태들이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "새의 두 발과 발톱이 어깨 장갑의 윗면에 닿아 체중을 지탱하고, 날개를 접은 자세는 내려앉은 직후의 정지 순간으로 자연스럽다. 어깨는 관절과 몸통에 연결되어 있다. 찰리의 지면 접촉부는 화면 밖이며, 보이는 부분에 공중 부양이나 불가능한 지지는 없다."
       },
       {
        "label": "A",
        "direction": "새의 부리와 시선은 화면 왼쪽, 카메라 쪽으로 비스듬히 향한다. 찰리의 머리는 화면 밖이어서 낚시 대상에 대한 시선을 확인할 수 없다. 아래쪽에는 물속이 보이지만 위쪽에는 수평선과 하늘까지 들어와, 어깨 너머 아래 바다를 보는 지정보다 수평 방향의 해안 전망이 강조된다.",
        "built_space": "인공 구조물이나 고정 설비는 보이지 않는다. 전경 왼쪽을 어깨 장갑이 크게 차지하고, 오른쪽 아래에는 얕은 물과 돌, 위쪽에는 먼 바다와 수평선이 있다. 왼쪽 배경의 높은 절벽 해안은 이전 장면의 낮은 암반과 작은 바위섬에 비해 지형 연속성이 약하다. 식탁은 화면 밖이며 물속 배경에 불가능한 반사는 없다.",
        "entities": "아기 새 한 마리와 찰리의 어깨 및 손 일부가 보이며 추가 인물은 없다. 새의 색, 둥근 몸통, 짧은 날개, 가는 발과 솜털은 참조에 가깝다. 어깨는 마모된 샌드 베이지 금속이지만 겹친 장갑판과 긴 홈의 형상은 참조 어깨와 다르다. 얼굴, 모자, 가슴 고리는 프레임 밖이라 평가하지 않는다. 맑은 물속에는 여러 물고기가 보이고, 하단에는 가는 낚싯줄로 보이는 선과 로봇 손 일부가 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "새의 양발이 경사진 어깨판과 돌출된 홈 가장자리에 접촉해 몸을 지탱한다. 접힌 날개와 안정된 몸통은 조용히 내려앉은 상태로 가능하다. 어깨 아래에는 연결된 기계 관절이 보이고, 하단 손 부근을 가는 줄이 지나가지만 정확한 파지점은 프레임과 가림 때문에 불분명하다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 지시된 카메라 뷰포인트(어깨 너머 아래쪽 수면을 내려다보는 구도) 위반"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 하향식 클로즈업 앵글을 정확히 구현하였고, 어깨 너머 맑은 물속의 물고기와 낚싯줄을 쥐고 있는 로봇의 손을 보여주어 프롬프트의 상황을 완벽하게 포착했습니다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "카메라 앵글이 지시된 하향식이 아닌 수평에 가깝고, 프레임에 포함된 로봇의 뒷통수와 목 구조가 캐릭터 레퍼런스와 일치하지 않습니다.  ★위반: [gemini-pro] 지시된 카메라 뷰포인트(어깨 너머 아래쪽 수면을 내려다보는 구도) 위반"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4_sel.png",
    "asset_id": "9495278b-8140-411c-a821-c73168b73c3b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163969>",
    "asset_id": "eb6423fc-e0d4-4710-92ed-f6aaac73e667",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:833571>",
    "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf411-3527-7392-9819-86e2197f14ac",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S91sh4"
  }
 },
 "S91sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T20:26:47.684697+00:00",
  "fingerprint": "e4b7079f4586f9bb29941039a7635c1fddf44d9cd203f346e51bd82921084e4a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S91sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S91sh5_sel.png",
  "source_sha256": "5bf0ba2da94a9e2dd2415c29edbcf50451868254516eca0223745e056ed87424",
  "file": "S91sh5_cine.png",
  "staged_sha256": "e42655c9e70d7ff908d2d37476ad59e28d6c60c524b738045b3057f2e688c030",
  "latency_ms": 9913
 },
 "S72sh43::cine": {
  "applied": false,
  "attempted_at": "2026-09-19T19:24:18.981200+00:00",
  "fingerprint": "918f5505a15ec70925c533347abe0ca8f6620762c6c21a56da2f80e3fda08be9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S72sh43_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "621e26e0f24113412ffeec0146852fa3a1cfc968e973887f9e3daf2682920c0c",
  "error": "RuntimeError: moderation blocked: {\"code\":\"imagine:content-moderated\",\"error\":\"Generated image rejected by content moderation.\",\"usage\":{\"cost_in_usd_ticks\":220000000}}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S72sh43::cine::fb_grok": {
  "applied": false,
  "attempted_at": "2026-09-19T19:24:25.793239+00:00",
  "fingerprint": "ae8d48103835412808303b9b90b3fe10a564df91878b278fd97a14de7c62f8d8",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-quality",
  "multiroll_tag": "still_S72sh43_cine_fb_grok",
  "slot": "fb_grok",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "621e26e0f24113412ffeec0146852fa3a1cfc968e973887f9e3daf2682920c0c",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"code\\\":\\\"imagine:content-moderated\\\",\\\"error\\\":\\\"Generated image rejected by content moderation.\\\",\\\"usage\\\":{\\\"cost_in_usd_ticks\\\":600000000}}\",\"provider_name\":\"xAI\",\"is_byok\":false}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S72sh43::cine::fb_mai": {
  "applied": false,
  "attempted_at": "2026-09-19T19:24:30.430610+00:00",
  "fingerprint": "880b49f9967df3c46a9d7697987db46da0d6cddff739ac56f80d9f34787bf0ff",
  "fingerprint_version": 2,
  "provider": "mai",
  "endpoint": "openrouter/chat-completions",
  "model": "microsoft/mai-image-2.6",
  "multiroll_tag": "still_S72sh43_cine_fb_mai",
  "slot": "fb_mai",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "621e26e0f24113412ffeec0146852fa3a1cfc968e973887f9e3daf2682920c0c",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"error\\\":{\\\"code\\\":\\\"content_safety_violation\\\",\\\"message\\\":\\\"Response content blocked by label 'MultiSeverity_ViolenceScore'.\\\",\\\"details\\\":\\\"Response content blocked by label 'MultiSeverity_ViolenceScore'.\\\"}}\",\"provider_name\":\"Azure\",\"is_byok\":false,\"provider_error_code\":\"content_safety_violation\"}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S72sh43::cine::fb_seedream": {
  "applied": true,
  "attempted_at": "2026-09-19T19:24:48.957760+00:00",
  "fingerprint": "b095fff494f76d8e825f4b6ea6d4f7f7df8356a42b724ed6bf8a00f56310d46d",
  "fingerprint_version": 2,
  "provider": "seedream",
  "endpoint": "openrouter/images+aspect_ratio",
  "model": "bytedance-seed/seedream-5-0-pro",
  "multiroll_tag": "still_S72sh43_cine_fb_seedream",
  "slot": "fb_seedream",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "621e26e0f24113412ffeec0146852fa3a1cfc968e973887f9e3daf2682920c0c",
  "file": "S72sh43_cine_fb_seedream.png",
  "staged_sha256": "6276a75a29f47c9b25d0e7b6622562c4b52e4bd5f6ab98ce224bcdf5522e084d",
  "latency_ms": 124888
 },
 "S11sh1::ab_noconti": {
  "input_fingerprint": "63762d0560cc85e5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 화면 우측 하단의 바닥을 향하고 있습니다.",
    "built_space": "카메라는 프롬프트가 지시한 입구 안쪽이 아닌, 레퍼런스와 동일한 바깥쪽에서 안을 바라보는 구도를 취하고 있습니다. 철문 끝단에 중앙 표지판이 반으로 잘린 채 왜곡되어 붙어 있으며, '난'이라는 글자가 명확히 보입니다.",
    "entities": "현우의 의상(회색 셔츠, 올리브색 바지, 부츠)과 헤어스타일, 전반적인 외형은 레퍼런스와 잘 일치합니다.",
    "hard_violations": [
     "[gemini-pro] 지정된 카메라 위치(안쪽) 위반 및 구조를 벗어난 바깥쪽 뷰 렌더링",
     "[gemini-pro] 레퍼런스의 구조를 무시한 철문 부착물(반쪽 표지판) 생성",
     "[gemini-pro] 읽을 수 있는 텍스트('난') 노출 (금지 사항 위반)",
     "[gpt-high] 카메라가 요구된 출입구 안쪽이 아니라 정착촌 바깥쪽에 놓여, 문 너머 외부 골목 대신 정착촌 내부를 보여 준다.",
     "[gpt-high] 철문 녹색 문패에 판독 가능한 한글이 있어 글자를 전혀 읽을 수 없어야 한다는 조건을 위반한다."
    ],
    "physics": "두 발은 바닥의 레일 부근을 딛고 있으며, 왼손으로 철문 모서리를 잡고 몸을 지탱하고 있어 물리적으로는 안정적입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '입구 안쪽' 카메라 위치를 어기고 바깥쪽 구도를 사용했으며, 철문 구조 변형과 금지된 텍스트 노출이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측 하단의 바닥을 향하고 있습니다.",
        "built_space": "카메라는 프롬프트가 지시한 입구 안쪽이 아닌, 레퍼런스와 동일한 바깥쪽에서 안을 바라보는 구도를 취하고 있습니다. 철문 끝단에 중앙 표지판이 반으로 잘린 채 왜곡되어 붙어 있으며, '난'이라는 글자가 명확히 보입니다.",
        "entities": "현우의 의상(회색 셔츠, 올리브색 바지, 부츠)과 헤어스타일, 전반적인 외형은 레퍼런스와 잘 일치합니다.",
        "hard_violations": [
         "지정된 카메라 위치(안쪽) 위반 및 구조를 벗어난 바깥쪽 뷰 렌더링",
         "레퍼런스의 구조를 무시한 철문 부착물(반쪽 표지판) 생성",
         "읽을 수 있는 텍스트('난') 노출 (금지 사항 위반)"
        ],
        "physics": "두 발은 바닥의 레일 부근을 딛고 있으며, 왼손으로 철문 모서리를 잡고 몸을 지탱하고 있어 물리적으로는 안정적입니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '입구 안쪽' 카메라 위치를 어기고 바깥쪽 구도를 사용했으며, 철문 구조 변형과 금지된 텍스트 노출이 발생했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측 하단의 바닥을 향하고 있습니다.",
        "built_space": "카메라는 프롬프트가 지시한 입구 안쪽이 아닌, 레퍼런스와 동일한 바깥쪽에서 안을 바라보는 구도를 취하고 있습니다. 철문 끝단에 중앙 표지판이 반으로 잘린 채 왜곡되어 붙어 있으며, '난'이라는 글자가 명확히 보입니다.",
        "entities": "현우의 의상(회색 셔츠, 올리브색 바지, 부츠)과 헤어스타일, 전반적인 외형은 레퍼런스와 잘 일치합니다.",
        "hard_violations": [
         "지정된 카메라 위치(안쪽) 위반 및 구조를 벗어난 바깥쪽 뷰 렌더링",
         "레퍼런스의 구조를 무시한 철문 부착물(반쪽 표지판) 생성",
         "읽을 수 있는 텍스트('난') 노출 (금지 사항 위반)"
        ],
        "physics": "두 발은 바닥의 레일 부근을 딛고 있으며, 왼손으로 철문 모서리를 잡고 몸을 지탱하고 있어 물리적으로는 안정적입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "야간 와이드 전신과 철문 재질은 맞지만, 정착촌 밖에서 안을 보는 역방향 구도이며 좁아지는 틈에 끼인 동작이 없고 읽히는 글자까지 남아 부적합합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 시선을 화면 왼쪽 철문 가장자리 쪽으로 돌리고 몸을 전경 쪽으로 기울인다. 문 뒤로 상점과 컨테이너 주거지가 펼쳐져 있어, 정착촌 안으로 들어오기보다 정착촌에서 카메라 쪽으로 나오는 동작으로 읽힌다. 요구된 내부에서 외부 골목을 바라보는 방향과 반대다.",
        "built_space": "화면 왼쪽에 마름모 장식 철문 한 조립체와 세로 구획 네 칸, 왼쪽 문기둥 하나, 구형 등 하나, 볼록거울 하나가 보인다. 문턱과 금속 레일은 하단을 비스듬히 가로지른다. 뒤쪽에는 왼편 상점과 오른편 컨테이너들, 중앙으로 굽어 이어지는 넓은 정착촌 도로가 있다. 참고 장소의 재료와 주요 요소는 유사하지만 외부에서 내부를 보는 배치다. 현우 오른쪽에 넓은 통행 공간이 남아 있고 몸을 압박할 반대쪽 문짝이나 문설주가 보이지 않아 좁아지는 통로가 성립하지 않는다. 볼록거울의 도로 반사는 눈에 띄게 불가능하다고 판단할 근거가 없다.",
        "entities": "인물은 한 명이며 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 회색 셔츠, 갈색 벨트, 올리브색 카고 바지와 어두운 부츠가 인물 참고와 대체로 일치한다. 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 무릎 부근에 손상된 천은 보이지만 미처 치료하지 않은 개 물림과 머리 부상은 명확하지 않다. 무거운 철창문과 컨테이너 주거지는 있으나 문 너머는 요구된 외부의 좁은 골목이 아니다. 녹색 문패에 읽을 수 있는 한글 '난'이 남아 있다.",
        "hard_violations": [
         "카메라가 요구된 출입구 안쪽이 아니라 정착촌 바깥쪽에 놓여, 문 너머 외부 골목 대신 정착촌 내부를 보여 준다.",
         "철문 녹색 문패에 판독 가능한 한글이 있어 글자를 전혀 읽을 수 없어야 한다는 조건을 위반한다."
        ],
        "physics": "두 부츠가 문턱과 노면에 닿아 체중을 지지하고, 보이는 손은 철문 세로 가장자리를 잡는다. 몸의 기울기와 벌린 다리는 문을 붙들고 이동하는 실제 동작으로 가능하며 공중에 뜬 신체나 물체는 없다. 다만 상체를 문 옆으로 기울인 자세일 뿐, 닫히는 두 경계 사이로 몸을 비틀어 구겨 넣어야 할 물리적 압박은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "야간 와이드 전신과 철문 재질은 맞지만, 정착촌 밖에서 안을 보는 역방향 구도이며 좁아지는 틈에 끼인 동작이 없고 읽히는 글자까지 남아 부적합합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 시선을 화면 왼쪽 철문 가장자리 쪽으로 돌리고 몸을 전경 쪽으로 기울인다. 문 뒤로 상점과 컨테이너 주거지가 펼쳐져 있어, 정착촌 안으로 들어오기보다 정착촌에서 카메라 쪽으로 나오는 동작으로 읽힌다. 요구된 내부에서 외부 골목을 바라보는 방향과 반대다.",
        "built_space": "화면 왼쪽에 마름모 장식 철문 한 조립체와 세로 구획 네 칸, 왼쪽 문기둥 하나, 구형 등 하나, 볼록거울 하나가 보인다. 문턱과 금속 레일은 하단을 비스듬히 가로지른다. 뒤쪽에는 왼편 상점과 오른편 컨테이너들, 중앙으로 굽어 이어지는 넓은 정착촌 도로가 있다. 참고 장소의 재료와 주요 요소는 유사하지만 외부에서 내부를 보는 배치다. 현우 오른쪽에 넓은 통행 공간이 남아 있고 몸을 압박할 반대쪽 문짝이나 문설주가 보이지 않아 좁아지는 통로가 성립하지 않는다. 볼록거울의 도로 반사는 눈에 띄게 불가능하다고 판단할 근거가 없다.",
        "entities": "인물은 한 명이며 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 회색 셔츠, 갈색 벨트, 올리브색 카고 바지와 어두운 부츠가 인물 참고와 대체로 일치한다. 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 무릎 부근에 손상된 천은 보이지만 미처 치료하지 않은 개 물림과 머리 부상은 명확하지 않다. 무거운 철창문과 컨테이너 주거지는 있으나 문 너머는 요구된 외부의 좁은 골목이 아니다. 녹색 문패에 읽을 수 있는 한글 '난'이 남아 있다.",
        "hard_violations": [
         "카메라가 요구된 출입구 안쪽이 아니라 정착촌 바깥쪽에 놓여, 문 너머 외부 골목 대신 정착촌 내부를 보여 준다.",
         "철문 녹색 문패에 판독 가능한 한글이 있어 글자를 전혀 읽을 수 없어야 한다는 조건을 위반한다."
        ],
        "physics": "두 부츠가 문턱과 노면에 닿아 체중을 지지하고, 보이는 손은 철문 세로 가장자리를 잡는다. 몸의 기울기와 벌린 다리는 문을 붙들고 이동하는 실제 동작으로 가능하며 공중에 뜬 신체나 물체는 없다. 다만 상체를 문 옆으로 기울인 자세일 뿐, 닫히는 두 경계 사이로 몸을 비틀어 구겨 넣어야 할 물리적 압박은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0
   },
   "adjusted": {
    "A": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 카메라 위치(안쪽) 위반 및 구조를 벗어난 바깥쪽 뷰 렌더링",
     "[gemini-pro] 레퍼런스의 구조를 무시한 철문 부착물(반쪽 표지판) 생성",
     "[gemini-pro] 읽을 수 있는 텍스트('난') 노출 (금지 사항 위반)",
     "[gpt-high] 카메라가 요구된 출입구 안쪽이 아니라 정착촌 바깥쪽에 놓여, 문 너머 외부 골목 대신 정착촌 내부를 보여 준다.",
     "[gpt-high] 철문 녹색 문패에 판독 가능한 한글이 있어 글자를 전혀 읽을 수 없어야 한다는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750
  },
  "selected": "A",
  "ranking": [
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 '입구 안쪽' 카메라 위치를 어기고 바깥쪽 구도를 사용했으며, 철문 구조 변형과 금지된 텍스트 노출이 발생했습니다.  ★위반: [gemini-pro] 지정된 카메라 위치(안쪽) 위반 및 구조를 벗어난 바깥쪽 뷰 렌더링 / [gemini-pro] 레퍼런스의 구조를 무시한 철문 부착물(반쪽 표지판) 생성 / [gemini-pro] 읽을 수 있는 텍스트('난') 노출 (금지 사항 위반) / [gpt-high] 카메라가 요구된 출입구 안쪽이 아니라 정착촌 바깥쪽에 놓여, 문 너머 외부 골목 대신 정착촌 내부를 보여 준다. / [gpt-high] 철문 녹색 문패에 판독 가능한 한글이 있어 글자를 전혀 읽을 수 없어야 한다는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": true,
  "shot_run_uid": "06aaeb5d-51ab-7a16-a642-773277a8b1a2"
 },
 "S11sh1::conti_ab_decision": {
  "fingerprint": "c9ebe21ebe30a044",
  "winner": "A",
  "outer": {
   "winner": "A",
   "verdicts": [
    {
     "order": "AB",
     "winner_label": "1",
     "winner_branch": "A",
     "reason_ko": "1번 후보는 좁은 골목길이라는 배경 설정과 18세 주인공의 앳된 외모를 잘 살렸으며, 2번에서 발견되는 읽을 수 있는 글자(표지판의 '난') 위반 사항이 없습니다."
    },
    {
     "order": "BA",
     "winner_label": "2",
     "winner_branch": "A",
     "reason_ko": "후보 2는 좁아지는 철창문 틈새로 몸을 비틀어 통과하는 다급한 자세와 프롬프트에 명시된 부상(머리와 팔의 상처)을 정확하게 묘사했습니다."
    }
   ]
  }
 },
 "S11sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T16:20:22.384204+00:00",
  "fingerprint": "3247142574b85b7d1ae2b7bc36f50ad5f62fafce13c6afc41d71eec4532dc5c9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S11sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S11sh1_sel.png",
  "source_sha256": "d8cec804b8a2cbc962b22e0945cd65a86ee611264eaf5f28dd84b97c190ccdb1",
  "file": "S11sh1_cine.png",
  "staged_sha256": "4c9ed196e01635b06d8b4f330a13f14b839b77c246831b9071bf0228a02ffd4d",
  "latency_ms": 12392
 },
 "S11sh1": {
  "input_fingerprint": "20f398e0fddffe58",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 화면 우측의 어두운 골목길을 향하고 있으며, 몸은 철문 틈새를 통과하는 방향으로 향해 있습니다.",
    "built_space": "화면 좌측에 마름모 장식이 있는 육중한 철창문이 닫히는 중이며, 우측에는 컨테이너 구조물들이 늘어선 골목길과 가로등, 그리고 우측 끝에 구형 조명과 확성기가 달린 기둥이 위치해 있습니다.",
    "entities": "현우(젊은 남성, 검은 머리, 회색 셔츠, 카고 바지)가 등장하며 얼굴과 팔다리에 상처와 핏자국이 묘사되어 레퍼런스와 일치합니다.",
    "hard_violations": [],
    "physics": "현우의 두 발은 바닥을 딛고 있으며, 오른손은 닫히는 철문의 끝부분을 짚고 있어 몸을 비틀어 빠져나가는 역동적인 자세를 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "주어진 레이아웃 스케치와 레퍼런스를 충실히 반영하여, 야간 배경 속 닫히는 철문 틈새로 몸을 구겨 넣는 현우의 전신과 부상 상태를 성공적으로 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측의 어두운 골목길을 향하고 있으며, 몸은 철문 틈새를 통과하는 방향으로 향해 있습니다.",
        "built_space": "화면 좌측에 마름모 장식이 있는 육중한 철창문이 닫히는 중이며, 우측에는 컨테이너 구조물들이 늘어선 골목길과 가로등, 그리고 우측 끝에 구형 조명과 확성기가 달린 기둥이 위치해 있습니다.",
        "entities": "현우(젊은 남성, 검은 머리, 회색 셔츠, 카고 바지)가 등장하며 얼굴과 팔다리에 상처와 핏자국이 묘사되어 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "현우의 두 발은 바닥을 딛고 있으며, 오른손은 닫히는 철문의 끝부분을 짚고 있어 몸을 비틀어 빠져나가는 역동적인 자세를 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "주어진 레이아웃 스케치와 레퍼런스를 충실히 반영하여, 야간 배경 속 닫히는 철문 틈새로 몸을 구겨 넣는 현우의 전신과 부상 상태를 성공적으로 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측의 어두운 골목길을 향하고 있으며, 몸은 철문 틈새를 통과하는 방향으로 향해 있습니다.",
        "built_space": "화면 좌측에 마름모 장식이 있는 육중한 철창문이 닫히는 중이며, 우측에는 컨테이너 구조물들이 늘어선 골목길과 가로등, 그리고 우측 끝에 구형 조명과 확성기가 달린 기둥이 위치해 있습니다.",
        "entities": "현우(젊은 남성, 검은 머리, 회색 셔츠, 카고 바지)가 등장하며 얼굴과 팔다리에 상처와 핏자국이 묘사되어 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "현우의 두 발은 바닥을 딛고 있으며, 오른손은 닫히는 철문의 끝부분을 짚고 있어 몸을 비틀어 빠져나가는 역동적인 자세를 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간 와이드숏과 문틈을 비집는 다급한 동작은 잘 구현했지만, 철문의 높이·장식·분할이 장소 사진과 다르고 정착촌 안으로 들어오는 이동 방향도 명확하지 않다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 골목을 향하고, 상체와 앞발도 오른쪽으로 나아간다. 뒤로 뻗은 손은 철문 세로 가장자리를 잡는다. 좁은 틈을 통과하는 행동은 분명하지만, 컨테이너 주거지가 오른쪽 골목에 보이므로 카메라가 정착촌 안에서 바깥을 본다는 조건과 현우의 안쪽 진입 경로를 동시에 명확히 확인하기는 어렵다.",
        "built_space": "왼쪽의 넓은 철문 면과 중앙의 좁은 철문 면 사이에 현우가 끼어 있으며, 하단에는 문 바퀴와 레일·문턱이 보인다. 오른쪽에는 구형 등 하나와 확성기 하나가 달린 콘크리트 기둥이 있고, 중앙 뒤쪽에도 작은 구형 등이 있는 기둥이 보인다. 오른쪽 배경에는 컨테이너 벽과 꺼진 가로등들이 이어진다. 배치 스케치의 큰 구도는 따르지만, 장소 사진의 낮고 여러 마름모 패널로 분할된 출입문 대신 사람보다 훨씬 높은 철창과 커다란 단일 마름모 장식을 보여 정확한 장소 구조 재현은 부족하다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 마른 체형은 인물 참고와 대체로 맞으며 국적은 외관만으로 확인할 수 없다. 회색 단추 셔츠, 올리브색 작업 바지와 어두운 부츠도 일치한다. 소매는 참고보다 걷혀 있고 팔과 바지에 상처 및 혈흔이 보이지만, 개에게 물린 상처인지와 기존 머리 부상의 상태는 명확하지 않다. 철문과 어두운 골목은 존재하며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 부츠가 문턱 안쪽 바닥에 닿아 체중을 받고, 뒤쪽 다리는 문틈 너머로 뻗어 있다. 뒤쪽 발의 접지는 문 가장자리와 어둠에 일부 가려져 있으나 앞발과 문을 잡은 손이 몸을 지지하므로 공중에 떠 있는 자세는 아니다. 굽힌 무릎과 기울여 비튼 상체는 급하게 좁은 틈을 통과하는 동작으로 가능하다. 철문은 하단 바퀴·레일과 기둥 구조로 지지되며, 정지 화면만으로 닫히는 운동 방향 자체는 확정하기 어렵다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간 와이드숏과 문틈을 비집는 다급한 동작은 잘 구현했지만, 철문의 높이·장식·분할이 장소 사진과 다르고 정착촌 안으로 들어오는 이동 방향도 명확하지 않다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 골목을 향하고, 상체와 앞발도 오른쪽으로 나아간다. 뒤로 뻗은 손은 철문 세로 가장자리를 잡는다. 좁은 틈을 통과하는 행동은 분명하지만, 컨테이너 주거지가 오른쪽 골목에 보이므로 카메라가 정착촌 안에서 바깥을 본다는 조건과 현우의 안쪽 진입 경로를 동시에 명확히 확인하기는 어렵다.",
        "built_space": "왼쪽의 넓은 철문 면과 중앙의 좁은 철문 면 사이에 현우가 끼어 있으며, 하단에는 문 바퀴와 레일·문턱이 보인다. 오른쪽에는 구형 등 하나와 확성기 하나가 달린 콘크리트 기둥이 있고, 중앙 뒤쪽에도 작은 구형 등이 있는 기둥이 보인다. 오른쪽 배경에는 컨테이너 벽과 꺼진 가로등들이 이어진다. 배치 스케치의 큰 구도는 따르지만, 장소 사진의 낮고 여러 마름모 패널로 분할된 출입문 대신 사람보다 훨씬 높은 철창과 커다란 단일 마름모 장식을 보여 정확한 장소 구조 재현은 부족하다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 마른 체형은 인물 참고와 대체로 맞으며 국적은 외관만으로 확인할 수 없다. 회색 단추 셔츠, 올리브색 작업 바지와 어두운 부츠도 일치한다. 소매는 참고보다 걷혀 있고 팔과 바지에 상처 및 혈흔이 보이지만, 개에게 물린 상처인지와 기존 머리 부상의 상태는 명확하지 않다. 철문과 어두운 골목은 존재하며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 부츠가 문턱 안쪽 바닥에 닿아 체중을 받고, 뒤쪽 다리는 문틈 너머로 뻗어 있다. 뒤쪽 발의 접지는 문 가장자리와 어둠에 일부 가려져 있으나 앞발과 문을 잡은 손이 몸을 지지하므로 공중에 떠 있는 자세는 아니다. 굽힌 무릎과 기울여 비튼 상체는 급하게 좁은 틈을 통과하는 동작으로 가능하다. 철문은 하단 바퀴·레일과 기둥 구조로 지지되며, 정지 화면만으로 닫히는 운동 방향 자체는 확정하기 어렵다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0
   },
   "adjusted": {
    "A": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000
  },
  "selected": "A",
  "ranking": [
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "주어진 레이아웃 스케치와 레퍼런스를 충실히 반영하여, 야간 배경 속 닫히는 철문 틈새로 몸을 구겨 넣는 현우의 전신과 부상 상태를 성공적으로 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "LAYOUT SKETCH — a bare thin-line layout guide, a REFERENCE ONLY: take from it ONLY the camera framing, figure placement, pose and size/depth order. It carries ZERO visual style — every texture, material, light and all realism come from the text and the photographic reference. Never let any line-drawing quality leak into the output.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S11sh1.png",
    "asset_id": "37dd0d43-2fdf-429b-82e4-045c520df701",
    "role": "conti_light"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:917012>",
    "asset_id": "20e84b30-7cc7-40b6-865a-3d45770aa9e8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40a-b844-7513-ab72-02f6c535a0f6",
  "conti_ab": {
   "winner": "A",
   "outer": {
    "winner": "A",
    "verdicts": [
     {
      "order": "AB",
      "winner_label": "1",
      "winner_branch": "A",
      "reason_ko": "1번 후보는 좁은 골목길이라는 배경 설정과 18세 주인공의 앳된 외모를 잘 살렸으며, 2번에서 발견되는 읽을 수 있는 글자(표지판의 '난') 위반 사항이 없습니다."
     },
     {
      "order": "BA",
      "winner_label": "2",
      "winner_branch": "A",
      "reason_ko": "후보 2는 좁아지는 철창문 틈새로 몸을 비틀어 통과하는 다급한 자세와 프롬프트에 명시된 부상(머리와 팔의 상처)을 정확하게 묘사했습니다."
     }
    ]
   },
   "outer_judged_this_run": false,
   "sel_conti": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh1__ab_conti_sel.png",
   "sel_noconti": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh1__ab_noconti_sel.png"
  },
  "ref_mode": "seed-bg+콘티+엔티티 (복잡구조물 A/B: 콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S48sh5::confined_fp": {
  "reads": {
   "controls": "The only primary control drawn is the steering wheel, attached to the front-left driver station immediately behind the windshield. It is not attached to Amber’s front-right passenger station. Both front seats face toward the windshield.",
   "mirrors": "No mirror or other explicitly reflective surface is marked. The windshield is identified as a window, with no indicated reflective face or reflection.",
   "camera": "The camera symbol sits immediately outboard of the front-right passenger seat and points inward, across the vehicle toward Amber. Its dashed viewing cone specifies a close-up of her face, rather than a forward view through the windshield. From this camera, the vehicle’s front is screen-right and its rear is screen-left.",
   "occupants": "Amber occupies the front-right passenger seat. Charlie occupies the rear-right seat and has a blanket over him. The driver seat and rear-left seat are not marked as occupied. Two police figures stand at the checkpoint outside, beyond the windshield; they are outside the indicated close-up."
  },
  "mismatches": [
   "The LOCATION specifies looking through the windshield toward the checkpoint, but the drawn camera points sideways inward at Amber from beside the passenger seat. The windshield and checkpoint are consequently outside this indicated close-up, toward screen-right."
  ],
  "scene_description_en": "The camera looks inward from beside the front-right passenger station, framing Amber’s face close and centrally rather than looking forward through the windshield. Amber occupies that passenger seat, which faces the vehicle’s front toward screen-right, making this a side-on view of her forward-facing position. Only a narrow portion of her passenger seat belongs beside and behind her in the close framing, toward screen-left. The unoccupied driver station lies farther across the cabin, with its steering wheel forward of the seat toward screen-right and outside the tight facial crop. Charlie occupies the rear-right seat under a blanket, facing forward toward screen-right but located off-screen to the left; the rear-left seat is also outside the frame and has no marked occupant. The windshield is forward and off-screen to the right, with the distant road barricades, two police figures, checkpoint booth, police vehicle, dead trees, and ruined houses beyond it, not visible in this sideways close-up. No mirror is drawn, so there is no mirror face or reflected view in the composition.",
  "readback_fallback": {
   "first_model": "gemini-pro",
   "first_error": "LLM returned empty response for step=confined_fp_readback_S48sh5_fix, model=gemini-pro, finish_reason='content_filter'",
   "model": "gpt-high",
   "physical_model": "gpt-6-astra"
  },
  "fixed": true,
  "input_fingerprint": "24c8c21baf352f38"
 },
 "era_assess::a0596258cfca1fe0": {
  "subjects": [],
  "subject_text": "화성의 황량한 도로\n황량한 들판을 가로지르는 도로. 주변에는 메마른 나무와 폐가가 드문드문 남아 있고 멀리 검문 시설이 보인다.",
  "identity": "canonical",
  "scope_id": "L203",
  "scope_role": "location_exterior",
  "scope_sha": "baf8d1a915ba6e43"
 },
 "S48sh5": {
  "input_fingerprint": "9974d58212cae45d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멀리 도로 끝에 세워진 바리케이드와 경찰 검문소를 바라보며 두 눈을 동그랗게 뜬 앰버의 놀란 얼굴.\n\nLOCATION (lock): Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Front passenger seat (Occupied by 앰버) — A narrow portion of the seat beside her shoulder is visible; used as Grounds the close reaction within the vehicle without revealing the other passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves clear eye detail and restrained contrast without exaggerating her surprise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The aging camper has a leaking roof and window frames, with the gathered supplies aboard; a police checkpoint stands farther along the road through barren fields, dead trees and abandoned houses. Charlie retains his worn metal body and blanket covering and is in the rear seat. 앰버: She occupies the passenger seat with the map and wears the replacement shoes obtained at the unmanned store.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera looks inward from beside the front-right passenger station, framing Amber’s face close and centrally rather than looking forward through the windshield. Amber occupies that passenger seat, which faces the vehicle’s front toward screen-right, making this a side-on view of her forward-facing position. Only a narrow portion of her passenger seat belongs beside and behind her in the close framing, toward screen-left. The unoccupied driver station lies farther across the cabin, with its steering wheel forward of the seat toward screen-right and outside the tight facial crop. Charlie occupies the rear-right seat under a blanket, facing forward toward screen-right but located off-screen to the left; the rear-left seat is also outside the frame and has no marked occupant. The windshield is forward and off-screen to the right, with the distant road barricades, two police figures, checkpoint booth, police vehicle, dead trees, and ruined houses beyond it, not visible in this sideways close-up. No mirror is drawn, so there is no mirror face or reflected view in the composition.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멀리 도로 끝에 세워진 바리케이드와 경찰 검문소를 바라보며 두 눈을 동그랗게 뜬 앰버의 놀란 얼굴.\n\nLOCATION (lock): Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves clear eye detail and restrained contrast without exaggerating her surprise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The aging camper has a leaking roof and window frames, with the gathered supplies aboard; a police checkpoint stands farther along the road through barren fields, dead trees and abandoned houses. Charlie retains his worn metal body and blanket covering and is in the rear seat. 앰버: She occupies the passenger seat with the map and wears the replacement shoes obtained at the unmanned store.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera looks inward from beside the front-right passenger station, framing Amber’s face close and centrally rather than looking forward through the windshield. Amber occupies that passenger seat, which faces the vehicle’s front toward screen-right, making this a side-on view of her forward-facing position. Only a narrow portion of her passenger seat belongs beside and behind her in the close framing, toward screen-left. The unoccupied driver station lies farther across the cabin, with its steering wheel forward of the seat toward screen-right and outside the tight facial crop. Charlie occupies the rear-right seat under a blanket, facing forward toward screen-right but located off-screen to the left; the rear-left seat is also outside the frame and has no marked occupant. The windshield is forward and off-screen to the right, with the distant road barricades, two police figures, checkpoint booth, police vehicle, dead trees, and ruined houses beyond it, not visible in this sideways close-up. No mirror is drawn, so there is no mirror face or reflected view in the composition.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멀리 도로 끝에 세워진 바리케이드와 경찰 검문소를 바라보며 두 눈을 동그랗게 뜬 앰버의 놀란 얼굴.\n\nLOCATION (lock): Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves clear eye detail and restrained contrast without exaggerating her surprise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The aging camper has a leaking roof and window frames, with the gathered supplies aboard; a police checkpoint stands farther along the road through barren fields, dead trees and abandoned houses. Charlie retains his worn metal body and blanket covering and is in the rear seat. 앰버: She occupies the passenger seat with the map and wears the replacement shoes obtained at the unmanned store.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh5_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "앰버",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh5_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "앰버",
     "path": "<bytes:934888>",
     "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 몸과 얼굴을 뒤로 돌려 화면 왼쪽의 차내 후방을 바라본다. 앞유리 밖 도로의 바리케이드와 검문소는 얼굴 뒤편에 있으므로 시선의 목표가 아니다. 멀리 선 경찰 두 명은 캠핑카 쪽을 향한다. 무기나 이동 중인 물체는 보이지 않는다.",
    "built_space": "후방 실내에서 앞을 보는 구도다. 왼쪽 운전석 등받이 하나와 그 앞 운전대 하나, 오른쪽 조수석 등받이 하나, 대시보드와 앞유리가 보인다. 앰버는 오른쪽 조수석에 있어 좌석 관계는 맞는다. 그러나 어깨 옆의 좁은 좌석 조각만 보여야 하는 지시와 달리 양쪽 좌석 및 실내 대부분이 노출된다. 도로에는 바리케이드 세 구간, 초소 하나, 경찰차 한 대가 보인다. 문제될 반사는 없다.",
    "entities": "금발과 둥근 얼굴을 가진 약 10세 여자아이 한 명으로, 앰버의 기본 외형 및 참고 얼굴과 대체로 부합한다. 혼혈 정체성 자체는 외모만으로 확정할 수 없다. 눈은 크게 뜨고 입술은 살짝 벌려 놀람을 표현한다. 참고의 머리 위 보호구는 없고 옷은 거의 가려져 비교하기 어렵다. 지도와 신발, 뒷좌석 찰리는 프레임 밖이므로 결함으로 보지 않는다. 황량한 들판, 죽은 나무, 폐가와 검문 시설은 보이나, 등장 허용 대상이 아닌 경찰 두 명도 보인다. 판독 가능한 글자는 없다.",
    "hard_violations": [
     "앰버만 등장하도록 제한한 지시와 달리 검문소에 경찰 두 명을 추가했다."
    ],
    "physics": "앰버의 하체는 등받이에 가려져 있지만 조수석에 앉아 상체와 목을 뒤로 돌린 자세로 해석 가능하다. 공중에 떠 있다는 증거는 없다. 좌석과 대시보드는 차체에 고정되어 있고, 경찰·바리케이드·초소·차량은 도로에 지지된다."
   },
   {
    "label": "A",
    "direction": "앰버의 얼굴과 두 눈은 후방 카메라 쪽을 향한다. 앞유리 너머 검문소는 앰버 뒤에 있으므로 검문소를 바라보는 순간이 아니다. 경찰 두 명은 캠핑카 쪽을 향한다. 손에 든 지도는 내려다보지 않는다.",
    "built_space": "차량 후방에서 앞유리를 보는 구도로, 왼쪽 운전석 등받이 하나와 운전대 하나, 오른쪽 가장자리의 조수석 일부, 대시보드, 중앙 거울 하나가 보인다. 앰버는 왼쪽 등받이 바로 앞이자 운전대 쪽에 있어 지정된 오른쪽 조수석 배치와 맞지 않는다. 얼굴뿐 아니라 상체와 지도까지 포함하고 실내도 넓게 보여 지정된 클로즈업에서 벗어난다. 밖에는 바리케이드 세 구간, 초소 하나, 경찰차 두 대가 보이며 검문소가 먼 도로 끝이라기보다 차량 가까이에 있다. 거울에는 판별 가능한 인물 반사가 없다.",
    "entities": "금발의 약 10세 여자아이 한 명이며 둥근 얼굴과 큰 눈은 앰버의 기본 특징에 부합한다. 혼혈 여부는 이미지로 확정할 수 없다. 참고의 남색 티셔츠와 갈색 멜빵 작업복 대신 갈색 깃 달린 상의를 입고, 머리 위 보호구도 없다. 종이 지도는 손에 들려 있다. 교체 신발과 찰리는 프레임 밖이다. 폐가와 죽은 나무, 검문 시설 외에 금지된 경찰 두 명이 등장한다. 바리케이드의 ‘POLICE’는 명확하게 읽힌다.",
    "hard_violations": [
     "앰버가 지정된 오른쪽 조수석이 아니라 운전석 쪽에 배치되어 있다.",
     "앰버만 등장하도록 제한한 지시와 달리 경찰 두 명을 추가했다.",
     "바리케이드에 판독 가능한 ‘POLICE’ 글자를 노출해 읽을 수 있는 글자 금지 지시를 위반했다."
    ],
    "physics": "앰버는 좌석 위에서 상체를 뒤로 돌린 자세로 보이며, 자세 자체가 물리적으로 불가능하지는 않다. 지도 가장자리를 손으로 잡고 있어 지지점이 있다. 경찰은 도로에 발을 딛고 있고 바리케이드·초소·차량도 지면에 놓여 있다. 지지 없이 떠 있는 물체는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "조수석 위치는 맞지만 검문소가 아닌 차내 뒤쪽을 바라보고, 얼굴 클로즈업 대신 실내를 넓게 보여주며 허용되지 않은 경찰 두 명을 추가했다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "얼굴은 더 크게 보이지만 검문소를 등진 시선, 운전석 쪽 인물 배치, 추가 경찰과 읽히는 ‘POLICE’ 글자가 핵심 지시를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 몸과 얼굴을 뒤로 돌려 화면 왼쪽의 차내 후방을 바라본다. 앞유리 밖 도로의 바리케이드와 검문소는 얼굴 뒤편에 있으므로 시선의 목표가 아니다. 멀리 선 경찰 두 명은 캠핑카 쪽을 향한다. 무기나 이동 중인 물체는 보이지 않는다.",
        "built_space": "후방 실내에서 앞을 보는 구도다. 왼쪽 운전석 등받이 하나와 그 앞 운전대 하나, 오른쪽 조수석 등받이 하나, 대시보드와 앞유리가 보인다. 앰버는 오른쪽 조수석에 있어 좌석 관계는 맞는다. 그러나 어깨 옆의 좁은 좌석 조각만 보여야 하는 지시와 달리 양쪽 좌석 및 실내 대부분이 노출된다. 도로에는 바리케이드 세 구간, 초소 하나, 경찰차 한 대가 보인다. 문제될 반사는 없다.",
        "entities": "금발과 둥근 얼굴을 가진 약 10세 여자아이 한 명으로, 앰버의 기본 외형 및 참고 얼굴과 대체로 부합한다. 혼혈 정체성 자체는 외모만으로 확정할 수 없다. 눈은 크게 뜨고 입술은 살짝 벌려 놀람을 표현한다. 참고의 머리 위 보호구는 없고 옷은 거의 가려져 비교하기 어렵다. 지도와 신발, 뒷좌석 찰리는 프레임 밖이므로 결함으로 보지 않는다. 황량한 들판, 죽은 나무, 폐가와 검문 시설은 보이나, 등장 허용 대상이 아닌 경찰 두 명도 보인다. 판독 가능한 글자는 없다.",
        "hard_violations": [
         "앰버만 등장하도록 제한한 지시와 달리 검문소에 경찰 두 명을 추가했다."
        ],
        "physics": "앰버의 하체는 등받이에 가려져 있지만 조수석에 앉아 상체와 목을 뒤로 돌린 자세로 해석 가능하다. 공중에 떠 있다는 증거는 없다. 좌석과 대시보드는 차체에 고정되어 있고, 경찰·바리케이드·초소·차량은 도로에 지지된다."
       },
       {
        "label": "B",
        "direction": "앰버의 얼굴과 두 눈은 후방 카메라 쪽을 향한다. 앞유리 너머 검문소는 앰버 뒤에 있으므로 검문소를 바라보는 순간이 아니다. 경찰 두 명은 캠핑카 쪽을 향한다. 손에 든 지도는 내려다보지 않는다.",
        "built_space": "차량 후방에서 앞유리를 보는 구도로, 왼쪽 운전석 등받이 하나와 운전대 하나, 오른쪽 가장자리의 조수석 일부, 대시보드, 중앙 거울 하나가 보인다. 앰버는 왼쪽 등받이 바로 앞이자 운전대 쪽에 있어 지정된 오른쪽 조수석 배치와 맞지 않는다. 얼굴뿐 아니라 상체와 지도까지 포함하고 실내도 넓게 보여 지정된 클로즈업에서 벗어난다. 밖에는 바리케이드 세 구간, 초소 하나, 경찰차 두 대가 보이며 검문소가 먼 도로 끝이라기보다 차량 가까이에 있다. 거울에는 판별 가능한 인물 반사가 없다.",
        "entities": "금발의 약 10세 여자아이 한 명이며 둥근 얼굴과 큰 눈은 앰버의 기본 특징에 부합한다. 혼혈 여부는 이미지로 확정할 수 없다. 참고의 남색 티셔츠와 갈색 멜빵 작업복 대신 갈색 깃 달린 상의를 입고, 머리 위 보호구도 없다. 종이 지도는 손에 들려 있다. 교체 신발과 찰리는 프레임 밖이다. 폐가와 죽은 나무, 검문 시설 외에 금지된 경찰 두 명이 등장한다. 바리케이드의 ‘POLICE’는 명확하게 읽힌다.",
        "hard_violations": [
         "앰버가 지정된 오른쪽 조수석이 아니라 운전석 쪽에 배치되어 있다.",
         "앰버만 등장하도록 제한한 지시와 달리 경찰 두 명을 추가했다.",
         "바리케이드에 판독 가능한 ‘POLICE’ 글자를 노출해 읽을 수 있는 글자 금지 지시를 위반했다."
        ],
        "physics": "앰버는 좌석 위에서 상체를 뒤로 돌린 자세로 보이며, 자세 자체가 물리적으로 불가능하지는 않다. 지도 가장자리를 손으로 잡고 있어 지지점이 있다. 경찰은 도로에 발을 딛고 있고 바리케이드·초소·차량도 지면에 놓여 있다. 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "조수석 위치는 맞지만 검문소가 아닌 차내 뒤쪽을 바라보고, 얼굴 클로즈업 대신 실내를 넓게 보여주며 허용되지 않은 경찰 두 명을 추가했다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "얼굴은 더 크게 보이지만 검문소를 등진 시선, 운전석 쪽 인물 배치, 추가 경찰과 읽히는 ‘POLICE’ 글자가 핵심 지시를 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 몸과 얼굴을 뒤로 돌려 화면 왼쪽의 차내 후방을 바라본다. 앞유리 밖 도로의 바리케이드와 검문소는 얼굴 뒤편에 있으므로 시선의 목표가 아니다. 멀리 선 경찰 두 명은 캠핑카 쪽을 향한다. 무기나 이동 중인 물체는 보이지 않는다.",
        "built_space": "후방 실내에서 앞을 보는 구도다. 왼쪽 운전석 등받이 하나와 그 앞 운전대 하나, 오른쪽 조수석 등받이 하나, 대시보드와 앞유리가 보인다. 앰버는 오른쪽 조수석에 있어 좌석 관계는 맞는다. 그러나 어깨 옆의 좁은 좌석 조각만 보여야 하는 지시와 달리 양쪽 좌석 및 실내 대부분이 노출된다. 도로에는 바리케이드 세 구간, 초소 하나, 경찰차 한 대가 보인다. 문제될 반사는 없다.",
        "entities": "금발과 둥근 얼굴을 가진 약 10세 여자아이 한 명으로, 앰버의 기본 외형 및 참고 얼굴과 대체로 부합한다. 혼혈 정체성 자체는 외모만으로 확정할 수 없다. 눈은 크게 뜨고 입술은 살짝 벌려 놀람을 표현한다. 참고의 머리 위 보호구는 없고 옷은 거의 가려져 비교하기 어렵다. 지도와 신발, 뒷좌석 찰리는 프레임 밖이므로 결함으로 보지 않는다. 황량한 들판, 죽은 나무, 폐가와 검문 시설은 보이나, 등장 허용 대상이 아닌 경찰 두 명도 보인다. 판독 가능한 글자는 없다.",
        "hard_violations": [
         "앰버만 등장하도록 제한한 지시와 달리 검문소에 경찰 두 명을 추가했다."
        ],
        "physics": "앰버의 하체는 등받이에 가려져 있지만 조수석에 앉아 상체와 목을 뒤로 돌린 자세로 해석 가능하다. 공중에 떠 있다는 증거는 없다. 좌석과 대시보드는 차체에 고정되어 있고, 경찰·바리케이드·초소·차량은 도로에 지지된다."
       },
       {
        "label": "A",
        "direction": "앰버의 얼굴과 두 눈은 후방 카메라 쪽을 향한다. 앞유리 너머 검문소는 앰버 뒤에 있으므로 검문소를 바라보는 순간이 아니다. 경찰 두 명은 캠핑카 쪽을 향한다. 손에 든 지도는 내려다보지 않는다.",
        "built_space": "차량 후방에서 앞유리를 보는 구도로, 왼쪽 운전석 등받이 하나와 운전대 하나, 오른쪽 가장자리의 조수석 일부, 대시보드, 중앙 거울 하나가 보인다. 앰버는 왼쪽 등받이 바로 앞이자 운전대 쪽에 있어 지정된 오른쪽 조수석 배치와 맞지 않는다. 얼굴뿐 아니라 상체와 지도까지 포함하고 실내도 넓게 보여 지정된 클로즈업에서 벗어난다. 밖에는 바리케이드 세 구간, 초소 하나, 경찰차 두 대가 보이며 검문소가 먼 도로 끝이라기보다 차량 가까이에 있다. 거울에는 판별 가능한 인물 반사가 없다.",
        "entities": "금발의 약 10세 여자아이 한 명이며 둥근 얼굴과 큰 눈은 앰버의 기본 특징에 부합한다. 혼혈 여부는 이미지로 확정할 수 없다. 참고의 남색 티셔츠와 갈색 멜빵 작업복 대신 갈색 깃 달린 상의를 입고, 머리 위 보호구도 없다. 종이 지도는 손에 들려 있다. 교체 신발과 찰리는 프레임 밖이다. 폐가와 죽은 나무, 검문 시설 외에 금지된 경찰 두 명이 등장한다. 바리케이드의 ‘POLICE’는 명확하게 읽힌다.",
        "hard_violations": [
         "앰버가 지정된 오른쪽 조수석이 아니라 운전석 쪽에 배치되어 있다.",
         "앰버만 등장하도록 제한한 지시와 달리 경찰 두 명을 추가했다.",
         "바리케이드에 판독 가능한 ‘POLICE’ 글자를 노출해 읽을 수 있는 글자 금지 지시를 위반했다."
        ],
        "physics": "앰버는 좌석 위에서 상체를 뒤로 돌린 자세로 보이며, 자세 자체가 물리적으로 불가능하지는 않다. 지도 가장자리를 손으로 잡고 있어 지지점이 있다. 경찰은 도로에 발을 딛고 있고 바리케이드·초소·차량도 지면에 놓여 있다. 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 2,
   "A": 1
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2,
    "verdict_ko": "조수석 위치는 맞지만 검문소가 아닌 차내 뒤쪽을 바라보고, 얼굴 클로즈업 대신 실내를 넓게 보여주며 허용되지 않은 경찰 두 명을 추가했다."
   },
   {
    "label": "A",
    "score": 1,
    "verdict_ko": "얼굴은 더 크게 보이지만 검문소를 등진 시선, 운전석 쪽 인물 배치, 추가 경찰과 읽히는 ‘POLICE’ 글자가 핵심 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh5_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "앰버",
    "path": "<bytes:934888>",
    "asset_id": "25957d1c-1e57-42c5-bf5c-3d3fe67590c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aaf40d-db51-7f0c-b9b7-f0ac40b1d755",
  "confined_fp": {
   "base_key": "confinedfp::0261ae55cee7",
   "apt_reason": "이 샷은 캠핑카 내부 조수석을 배경으로 하며, 앰버가 조수석에 앉아 앞유리를 통해 전방의 경찰 검문소를 바라보는 상황입니다. 차량 내부의 정확한 좌석 배치와 인물의 시선 방향이 어긋나면 화면의 일관성과 몰입을 해칠 수 있으므로 평면도 형태의 레이아웃 가이드가 필요합니다.",
   "fixed": true,
   "mismatches": [
    "The LOCATION specifies looking through the windshield toward the checkpoint, but the drawn camera points sideways inward at Amber from beside the passenger seat. The windshield and checkpoint are consequently outside this indicated close-up, toward screen-right."
   ]
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S48sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T18:11:55.569120+00:00",
  "fingerprint": "0fa6dc2eeb2a3d80cf7b2a7d9acca7100e3c3c368ca7b3162cd1cd423ce415ed",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S48sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S48sh5_sel.png",
  "source_sha256": "64e26758e1ad7672ffe53ffcb0163c992bb4cf30c34703e75baae531a4464db1",
  "file": "S48sh5_cine.png",
  "staged_sha256": "cc301bfe2cd7d63114dfce8e7ce6c6f10544aea0037cc701a26d4ded13a4e545",
  "latency_ms": 10614
 },
 "era_assess::01eb3d1e2773b936": {
  "subjects": [],
  "subject_text": "쓰레기 수거선 갑판\n대형 수거선의 넓은 금속 갑판. 수거용 크레인과 그물, 쌓인 해양 폐기물과 고철이 낮빛 아래 드러난다.",
  "identity": "canonical",
  "scope_id": "L147",
  "scope_role": "location_exterior",
  "scope_sha": "007731c4bba546f4"
 },
 "groupbg::collection_deck": {
  "input_fingerprint": "ddda1847e8c8fde8",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "collection_deck",
    "tags": [
     "S2sh3"
    ]
   },
   "context_sig": "681bc2b5ba32f5eb"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n쓰레기 수거선 갑판: 바다에서 건져 올린 쓰레기가 쏟아지는 크고 거친 금속 갑판. (특징: 철제 선박 갑판; 위에서 우수수 떨어지는 쓰레기 무더기; 그물에 엉켜 있는 낡은 고철 로봇)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 수거선의 갑판으로 우수수 떨어지는 쓰레기들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n쓰레기 수거선 갑판: 바다에서 건져 올린 쓰레기가 쏟아지는 크고 거친 금속 갑판. (특징: 철제 선박 갑판; 위에서 우수수 떨어지는 쓰레기 무더기; 그물에 엉켜 있는 낡은 고철 로봇)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 수거선의 갑판으로 우수수 떨어지는 쓰레기들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_collection_deck_b9e7b9.png",
  "asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150",
  "input_asset_ids": [
   "b077dc93-9217-44ee-9e06-e7c3ca0dcd21"
  ],
  "origin_tag": "S2sh3",
  "place_text": "On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.",
  "origin_inputs": {
   "place_text": "On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.",
   "time_of_day_en": "day",
   "conti_asset_id": "b077dc93-9217-44ee-9e06-e7c3ca0dcd21"
  }
 },
 "S2sh3::bgfirst_bg": {
  "input_fingerprint": "d391387eb9ee00b4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S2sh3__bgfirst_bg.png",
  "asset_id": "22001d0d-2139-4c65-801f-8fe7b7978313",
  "input_asset_ids": [
   "b077dc93-9217-44ee-9e06-e7c3ca0dcd21",
   "ad655f79-61dd-4341-a8cb-31e490c3f150"
  ]
 },
 "groupbg::dump_emergence": {
  "input_fingerprint": "a058d0a8ccb0e7ce",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "dump_emergence",
    "tags": [
     "S6sh15",
     "S6sh17"
    ]
   },
   "context_sig": "f05b52e5623c370f"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 쓰레기장: 산처럼 거대하게 쌓인 폐기물과 고철 더미로 이루어진 구역. (특징: 하늘을 가릴 듯 높은 거대한 쓰레기 산; 우그러진 로봇 팔, 기어 등 각종 고철 부품; 부서진 낡은 주크박스 몸체와 나팔형 스피커; 바닥에 깔린 모포와 널브러진 공구, LP판)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 더미에서 몸을 서서히 드러내는 건.... 고릴라 모양의 로봇이다.\n- 앰버를 내려다보다 방긋 웃더니, 이내 와락 껴안는다!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 쓰레기장: 산처럼 거대하게 쌓인 폐기물과 고철 더미로 이루어진 구역. (특징: 하늘을 가릴 듯 높은 거대한 쓰레기 산; 우그러진 로봇 팔, 기어 등 각종 고철 부품; 부서진 낡은 주크박스 몸체와 나팔형 스피커; 바닥에 깔린 모포와 널브러진 공구, LP판)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 더미에서 몸을 서서히 드러내는 건.... 고릴라 모양의 로봇이다.\n- 앰버를 내려다보다 방긋 웃더니, 이내 와락 껴안는다!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_dump_emergence_3bffab.png",
  "asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903",
  "input_asset_ids": [
   "c554fa18-115e-4e6b-9462-3d5f65b3a88e"
  ],
  "origin_tag": "S6sh15",
  "place_text": "Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.",
  "origin_inputs": {
   "place_text": "Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.",
   "time_of_day_en": "day",
   "conti_asset_id": "c554fa18-115e-4e6b-9462-3d5f65b3a88e"
  }
 },
 "S6sh15::bgfirst_bg": {
  "input_fingerprint": "f2e37e8b99b592df",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15__bgfirst_bg.png",
  "asset_id": "8d75d911-3d18-4135-b6ab-fc26e495c901",
  "input_asset_ids": [
   "c554fa18-115e-4e6b-9462-3d5f65b3a88e",
   "27c70e61-a2a2-4e06-b8da-d8a7b9325903"
  ]
 },
 "era_assess::f2fe65929099e9dd": {
  "subjects": [],
  "subject_text": "라울의 컨테이너 앞 공터\n낡은 철제 컨테이너 출입문 앞에 마련된 작은 공터. 주변에 컨테이너 하우스가 늘어서고 바닥에는 물 호스가 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L169",
  "scope_role": "location_exterior",
  "scope_sha": "d5e7f7fd70a3c7c2"
 },
 "S16sh3::bgfirst_bg": {
  "input_fingerprint": "c93785fe45234431",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3__bgfirst_bg.png",
  "asset_id": "58f097e3-cb89-4739-96b1-20ba4de0ccc6",
  "input_asset_ids": [
   "083f7cfb-ee18-4327-afb0-899eddfe6874",
   "a2e9a0db-ecb3-47e7-96eb-291c9f873ffb"
  ]
 },
 "groupbg::banana_market_stall": {
  "input_fingerprint": "d286953e23f5d0a1",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "banana_market_stall",
    "tags": [
     "S22sh3",
     "S22sh7",
     "S22sh9"
    ]
   },
   "context_sig": "6d6114ecb3b8f4ca"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 그러다 바나나를 파는 가게 앞에 멈추는 찰리.\n- 인파 틈으로 사라지는 찰리. 그 자리로 뛰어오는 앰버와 라울.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 그러다 바나나를 파는 가게 앞에 멈추는 찰리.\n- 인파 틈으로 사라지는 찰리. 그 자리로 뛰어오는 앰버와 라울.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_banana_market_stall_8677ef.png",
  "asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958",
  "input_asset_ids": [
   "774388e6-e167-4f76-a648-7a8dce68ade1"
  ],
  "origin_tag": "S22sh3",
  "place_text": "At the street-facing banana display of an outdoor market stall in the refugee settlement.",
  "origin_inputs": {
   "place_text": "At the street-facing banana display of an outdoor market stall in the refugee settlement.",
   "time_of_day_en": "day",
   "conti_asset_id": "774388e6-e167-4f76-a648-7a8dce68ade1"
  }
 },
 "S22sh3::bgfirst_bg": {
  "input_fingerprint": "16b23c64cd826776",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3__bgfirst_bg.png",
  "asset_id": "e6ccf295-a01a-4188-a4ca-5ce36e1609e9",
  "input_asset_ids": [
   "774388e6-e167-4f76-a648-7a8dce68ade1",
   "0333daa2-9ee6-4388-bf59-c510fb37f958"
  ]
 },
 "groupbg::reception_clearing": {
  "input_fingerprint": "4459c25d668f8883",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "reception_clearing",
    "tags": [
     "S25sh12",
     "S25sh19",
     "S26sh7"
    ]
   },
   "context_sig": "4d84b3e4221710d9"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 지하 하수도: 지상에서 맨홀과 사다리를 통해 내려오는 습하고 좁은 지하 콘크리트 통로. (특징: 수직으로 이어진 낡은 철제 사다리; 갈라지는 두 갈래의 콘크리트 길; 바닥의 물기와 어두운 조명) / 인천 난민촌 피로연장: 빈 공터에 천막을 치고 음악을 틀어놓은 조악한 파티장. (특징: 공터 위로 드리워진 허름한 천막; 빛나는 주크박스; 순간적으로 꺼졌다가 과부하로 밝게 터져나가는 가로등과 전구들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- (*피로연장이라고 해봤자 그냥 빈 공터에 천막 치는 정도)\n- /피로연장 -N\n- 하객들은 한가운데 무리 지어있고, 그들을 총과 몽둥이로 위협하고 있는 상황\n\nTIME OF DAY (lock): sunset.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 지하 하수도: 지상에서 맨홀과 사다리를 통해 내려오는 습하고 좁은 지하 콘크리트 통로. (특징: 수직으로 이어진 낡은 철제 사다리; 갈라지는 두 갈래의 콘크리트 길; 바닥의 물기와 어두운 조명) / 인천 난민촌 피로연장: 빈 공터에 천막을 치고 음악을 틀어놓은 조악한 파티장. (특징: 공터 위로 드리워진 허름한 천막; 빛나는 주크박스; 순간적으로 꺼졌다가 과부하로 밝게 터져나가는 가로등과 전구들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- (*피로연장이라고 해봤자 그냥 빈 공터에 천막 치는 정도)\n- /피로연장 -N\n- 하객들은 한가운데 무리 지어있고, 그들을 총과 몽둥이로 위협하고 있는 상황\n\nTIME OF DAY (lock): sunset.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_reception_clearing_3488b1.png",
  "asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1",
  "input_asset_ids": [
   "b30b0e2b-b930-4260-9eaf-f19f33633ca4"
  ],
  "origin_tag": "S25sh12",
  "place_text": "In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.",
  "origin_inputs": {
   "place_text": "In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.",
   "time_of_day_en": "sunset",
   "conti_asset_id": "b30b0e2b-b930-4260-9eaf-f19f33633ca4"
  }
 },
 "S25sh12::bgfirst_bg": {
  "input_fingerprint": "d9dcc640cf231108",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12__bgfirst_bg.png",
  "asset_id": "44dbc893-09f5-4451-985f-95d270aac79f",
  "input_asset_ids": [
   "b30b0e2b-b930-4260-9eaf-f19f33633ca4",
   "7f28b2fd-054b-41f4-b119-0c0d0c391ef1"
  ]
 },
 "era_assess::f3c165a73d599006": {
  "subjects": [],
  "subject_text": "인천 난민촌 도로와 맨홀 주변\n컨테이너 주거지 사이를 지나는 도로. 노면에 둥근 맨홀 뚜껑이 있고 길을 따라 가로등이 이어진다.",
  "identity": "canonical",
  "scope_id": "L177",
  "scope_role": "location_exterior",
  "scope_sha": "50531cbe908dd43e"
 },
 "S27sh12::bgfirst_bg": {
  "input_fingerprint": "9bfccaadca8a231e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12__bgfirst_bg.png",
  "asset_id": "164bd776-d0dc-43e4-bf32-6ef9f6356b78",
  "input_asset_ids": [
   "29e8ef08-4d4f-4022-ac9f-847c4599b99d",
   "785d4e92-6642-43a7-a3f4-70e3c4f7a9b3",
   "2b4b407b-73bc-405e-8265-11c6b8d30ef7"
  ]
 },
 "era_assess::678a9ff44c021387": {
  "subjects": [],
  "subject_text": "무인점포 내부\n신발과 가방, 식품, 생활용품이 진열된 무인 매장. 여러 진열대 사이로 통로가 나 있고 의약품 코너와 현금인출기가 설치돼 있다.",
  "identity": "canonical",
  "scope_id": "L195",
  "scope_role": "location_interior",
  "scope_sha": "ef0c39aef8ac337b"
 },
 "S41sh14::bgfirst_bg": {
  "input_fingerprint": "6eba475ba1d90762",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14__bgfirst_bg.png",
  "asset_id": "6398fb6a-6f79-4ad1-90ee-719e7320869a",
  "input_asset_ids": [
   "2b2818ac-406a-4c9d-805c-3eb8df283def",
   "830dae63-3f71-4c94-9045-5c77a027e48e"
  ]
 },
 "groupbg::camp_gate": {
  "input_fingerprint": "9acb2b35bb0ae94e",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "camp_gate",
    "tags": [
     "S11sh1",
     "S42sh11",
     "S42sh19",
     "S42sh2"
    ]
   },
   "context_sig": "d84d14ae7a17fa66"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 입구와 철창문 앞: 통금을 알리는 방송 장비가 설치된 거대한 수용소 출입구 구역. (특징: 난민 거주 구역을 분리하는 거대한 철문; 곳곳에 설치된 낡은 확성기 스피커; 가로등이 소등되어 컴컴해지는 거리)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 출입문 닫기 전에 헐레벌떡 뛰어오는 난민들, 그 틈에 현우도 끼여서 간신히 들어온다.\n- 42. 난민촌 입구 – N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 입구와 철창문 앞: 통금을 알리는 방송 장비가 설치된 거대한 수용소 출입구 구역. (특징: 난민 거주 구역을 분리하는 거대한 철문; 곳곳에 설치된 낡은 확성기 스피커; 가로등이 소등되어 컴컴해지는 거리)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 출입문 닫기 전에 헐레벌떡 뛰어오는 난민들, 그 틈에 현우도 끼여서 간신히 들어온다.\n- 42. 난민촌 입구 – N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_gate_b39179.png",
  "asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2",
  "input_asset_ids": [
   "ed925a67-f2a3-41f5-9413-65abf6516887"
  ],
  "origin_tag": "S42sh2",
  "place_text": "At ground level near the refugee settlement entrance at night, looking into an opened body bag.",
  "origin_inputs": {
   "place_text": "At ground level near the refugee settlement entrance at night, looking into an opened body bag.",
   "time_of_day_en": "night",
   "conti_asset_id": "ed925a67-f2a3-41f5-9413-65abf6516887"
  }
 },
 "S42sh2::bgfirst_bg": {
  "input_fingerprint": "d9da0cfaff19a919",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2__bgfirst_bg.png",
  "asset_id": "0428d63b-de1f-42c5-ad4e-e5abb4f95632",
  "input_asset_ids": [
   "ed925a67-f2a3-41f5-9413-65abf6516887",
   "83651a7a-615d-46c7-ba87-2459a6bc8bf2"
  ]
 },
 "era_assess::912540b6f9cc78bf": {
  "subjects": [],
  "subject_text": "찰리가 홀로 걷는 지방도로\n도시 밖으로 길게 이어지는 지방도로. 양옆으로 길가 지면이 드러나고 인근 숲과 산길로 이어지는 갈림길이 있다.",
  "identity": "canonical",
  "scope_id": "L212",
  "scope_role": "location_exterior",
  "scope_sha": "53aa184621808354"
 },
 "groupbg::rural_walking_road": {
  "input_fingerprint": "3d470ba3998b1a53",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "rural_walking_road",
    "tags": [
     "S52sh1"
    ]
   },
   "context_sig": "f7a1a96b56ac9a0d"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n찰리가 홀로 걷는 지방도로: 비 온 뒤 젖어 있는 외곽 차도를 독특한 복장을 한 인물이 걷는 풍경. (특징: 넓은 챙의 밀짚모자; 커다란 고무장화; 알록달록한 비닐 우비; 적막한 시골 도로)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 52. 지방도로 – D\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n찰리가 홀로 걷는 지방도로: 비 온 뒤 젖어 있는 외곽 차도를 독특한 복장을 한 인물이 걷는 풍경. (특징: 넓은 챙의 밀짚모자; 커다란 고무장화; 알록달록한 비닐 우비; 적막한 시골 도로)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 52. 지방도로 – D\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_rural_walking_road_5c3d2f.png",
  "asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566",
  "input_asset_ids": [
   "364391b6-7633-4178-a784-c9f3499899a6"
  ],
  "origin_tag": "S52sh1",
  "place_text": "On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.",
  "origin_inputs": {
   "place_text": "On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.",
   "time_of_day_en": "day",
   "conti_asset_id": "364391b6-7633-4178-a784-c9f3499899a6"
  }
 },
 "S52sh1::bgfirst_bg": {
  "input_fingerprint": "84b0daefeff7cd87",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S52sh1__bgfirst_bg.png",
  "asset_id": "4bc8f5c0-751d-4e2e-8589-aaf2e188a7ae",
  "input_asset_ids": [
   "364391b6-7633-4178-a784-c9f3499899a6",
   "2c211078-5c11-4ab2-abaa-988d6405f566"
  ]
 },
 "groupbg::forest_trap_site": {
  "input_fingerprint": "490cf693a7eb6de1",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "forest_trap_site",
    "tags": [
     "S56sh14",
     "S56sh8"
    ]
   },
   "context_sig": "fe15f055e3293da6"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /숲 길 -D\n- 현우가 도착한 곳은 찰리가 있던 장소, 그러나 아무도 없고...\n- 순간 트랩에 빠지는 현우와 앰버\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /숲 길 -D\n- 현우가 도착한 곳은 찰리가 있던 장소, 그러나 아무도 없고...\n- 순간 트랩에 빠지는 현우와 앰버\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_forest_trap_site_08ebbc.png",
  "asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847",
  "input_asset_ids": [
   "b891d386-6990-4b27-8c39-13f03b4127da"
  ],
  "origin_tag": "S56sh8",
  "place_text": "On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.",
  "origin_inputs": {
   "place_text": "On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.",
   "time_of_day_en": "day",
   "conti_asset_id": "b891d386-6990-4b27-8c39-13f03b4127da"
  }
 },
 "S56sh8::bgfirst_bg": {
  "input_fingerprint": "a24eece6f5f15e38",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8__bgfirst_bg.png",
  "asset_id": "7e75322c-4c67-4f9e-95d0-69b421b60a39",
  "input_asset_ids": [
   "b891d386-6990-4b27-8c39-13f03b4127da",
   "0948ec1a-4562-45ba-965f-9a74745e6847"
  ]
 },
 "S59sh36::bgfirst_bg": {
  "input_fingerprint": "5f87be3da579985f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh36__bgfirst_bg.png",
  "asset_id": "c373b774-4997-4590-9fd6-e9ba287393af",
  "input_asset_ids": [
   "f5ca6eee-4fd8-46c4-909f-e7354ed6208f",
   "dfd97579-913b-4622-86a5-0d6cd32f2ea6"
  ]
 },
 "S60sh4::bgfirst_bg": {
  "input_fingerprint": "0749be9b35ce6836",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh4__bgfirst_bg.png",
  "asset_id": "ab433a49-41f6-42f3-9c2c-d68dd6bf3b32",
  "input_asset_ids": [
   "6a6601d1-910a-4eba-a636-eb4b443cdfb9",
   "239110df-f90f-4ca2-b3ef-4a8f15c13fe9",
   "7cc6d446-071f-4801-886a-a504b66fe88c"
  ]
 },
 "era_assess::61cd8d32cafd0b75": {
  "subjects": [],
  "subject_text": "익산 마을의 고장 난 로봇 보관 폐창고\n기계 부품과 고장 난 로봇 몸체가 산더미처럼 쌓인 대형 폐창고. 반쯤 열린 문으로 빛이 들어오고 구석에는 상자들이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L238",
  "scope_role": "location_interior",
  "scope_sha": "bbf27bbaa44eebab"
 },
 "S64sh7::bgfirst_bg": {
  "input_fingerprint": "297b9287270f504b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7__bgfirst_bg.png",
  "asset_id": "2c542339-d519-40a3-aee9-40b352e75792",
  "input_asset_ids": [
   "cbafc64b-e6e6-4acf-8c80-6fb23ad1a5d9",
   "e1a9f48b-d9e5-4c90-99a8-fadcfb66b2d8"
  ]
 },
 "groupbg::village_truck_ambush": {
  "input_fingerprint": "063be8a200e52513",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "village_truck_ambush",
    "tags": [
     "S67sh49",
     "S67sh70"
    ]
   },
   "context_sig": "89cb53bfef34571e"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 빠르게 멀어지는 민병대 트럭!\n- 트럭 앞에서 대기하고 있는...B-200.\n- 찰리, 다친 몸을 이끌고 B-200의 폭파지점으로 향한다.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 빠르게 멀어지는 민병대 트럭!\n- 트럭 앞에서 대기하고 있는...B-200.\n- 찰리, 다친 몸을 이끌고 B-200의 폭파지점으로 향한다.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
  "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
  "input_asset_ids": [
   "f64f2986-6e3d-4497-9926-988ba8db7762"
  ],
  "origin_tag": "S67sh49",
  "place_text": "On the road leading out of the village, directly ahead of the departing militia truck convoy at night.",
  "origin_inputs": {
   "place_text": "On the road leading out of the village, directly ahead of the departing militia truck convoy at night.",
   "time_of_day_en": "night, bright moonlight",
   "conti_asset_id": "f64f2986-6e3d-4497-9926-988ba8db7762"
  }
 },
 "S67sh49::bgfirst_bg": {
  "input_fingerprint": "fe84cf31c82330ef",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh49__bgfirst_bg.png",
  "asset_id": "5b89bd25-951b-4371-9526-4d76774415ae",
  "input_asset_ids": [
   "f64f2986-6e3d-4497-9926-988ba8db7762",
   "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  ]
 },
 "S67sh70::bgfirst_bg": {
  "input_fingerprint": "fe7986440477c893",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh70__bgfirst_bg.png",
  "asset_id": "24c611fd-00da-42ad-afe0-c4324cb61919",
  "input_asset_ids": [
   "0909ba75-540e-4c7a-b173-2342ee20e201",
   "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  ]
 },
 "era_assess::78e004d7320278aa": {
  "subjects": [],
  "subject_text": "익산과 서남부 내륙의 지방도로\n완만한 능선과 들판 사이로 이어지는 지방도로. 마을 밖 갈림길에서 도로가 두 방향으로 나뉘며 멀리 낮은 산자락이 보인다.",
  "identity": "canonical",
  "scope_id": "L236",
  "scope_role": "location_exterior",
  "scope_sha": "ed8bfc6afb2c3029"
 },
 "groupbg::open_truck_bed": {
  "input_fingerprint": "f5dd248f6c584f8e",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "open_truck_bed",
    "tags": [
     "S68sh7"
    ]
   },
   "context_sig": "b534b76a0a5277eb"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산과 서남부 내륙의 지방도로: 차량 여러 대가 열을 지어 달리다가 길이 양옆으로 갈라지는 확 트인 농로 및 국도. (특징: 나란히 달리는 지붕 없는 구형 트럭들; Y자 형태로 갈라지는 교차로/갈림길 노면; 흙먼지를 일으키며 멀어지는 타이어 궤적)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 짐칸에 앉은 찰리 B-200의 가슴 부품을 꺼내본다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산과 서남부 내륙의 지방도로: 차량 여러 대가 열을 지어 달리다가 길이 양옆으로 갈라지는 확 트인 농로 및 국도. (특징: 나란히 달리는 지붕 없는 구형 트럭들; Y자 형태로 갈라지는 교차로/갈림길 노면; 흙먼지를 일으키며 멀어지는 타이어 궤적)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 짐칸에 앉은 찰리 B-200의 가슴 부품을 꺼내본다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_bed_9f9871.png",
  "asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49",
  "input_asset_ids": [
   "aa8d3a60-5597-4fad-886f-90bc9f48a604"
  ],
  "origin_tag": "S68sh7",
  "place_text": "In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.",
  "origin_inputs": {
   "place_text": "In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.",
   "time_of_day_en": "day",
   "conti_asset_id": "aa8d3a60-5597-4fad-886f-90bc9f48a604"
  }
 },
 "S68sh7::bgfirst_bg": {
  "input_fingerprint": "d508cb28332e1ada",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S68sh7__bgfirst_bg.png",
  "asset_id": "16f6020b-ef17-4edc-b91e-0f09780d6056",
  "input_asset_ids": [
   "aa8d3a60-5597-4fad-886f-90bc9f48a604",
   "f8675ac4-ad27-42ab-a598-758f4a775a49"
  ]
 },
 "era_assess::8b9cd93cb333094d": {
  "subjects": [],
  "subject_text": "크리스의 선박 선원실\n두 명 정도가 누울 수 있는 작은 선원실. 벽 가까이 소파가 놓여 있고 미닫이 출입문으로 바깥 통로와 연결된다.",
  "identity": "canonical",
  "scope_id": "L263",
  "scope_role": "location_interior",
  "scope_sha": "963e62a141b2a059"
 },
 "S79sh10::bgfirst_bg": {
  "input_fingerprint": "2d74ae5216f638dc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10__bgfirst_bg.png",
  "asset_id": "9ca43f94-b62c-4185-bbd5-d874b7069c57",
  "input_asset_ids": [
   "471ec738-cab9-44f7-b95c-4082e4ab2709",
   "a3a7948a-4ed6-4034-bdbc-66720a85c15b"
  ]
 },
 "S82sh14::bgfirst_bg": {
  "input_fingerprint": "629fe2bea12b2c63",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14__bgfirst_bg.png",
  "asset_id": "09c1e32d-ad58-4bd4-a51a-d8b7432ecb22",
  "input_asset_ids": [
   "90735080-f327-4c19-9101-6d98f10c3031",
   "a491e625-4c37-4571-b3ba-85a3daf996f9"
  ]
 },
 "S83sh5::bgfirst_bg": {
  "input_fingerprint": "8df7f491a6de2e97",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh5__bgfirst_bg.png",
  "asset_id": "3bc12182-b8a5-4b6b-bbfb-1fc47da3fb74",
  "input_asset_ids": [
   "481275cf-ea10-437e-8452-befc99e03a7f",
   "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a"
  ]
 },
 "S5sh16::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-19: 「양팔의 구도이상해서 경찰 팔이 3개야」 — 변환본. 원본은 경찰 둘",
  "at": "2026-09-19T15:25:15.679163+00:00",
  "rejected_cine": {
   "file": "S5sh16_cine.png",
   "staged_sha256": "e680ef142c37eed59f0ed698211d69891305615c969e091465597163c76872a0",
   "source_sha256": "c1fc7f5f50a357c35bb36404585a3b4a60d5e26f85887338b7ffac2cae6cdf7b",
   "fingerprint": "ab1e61464228cf5de73ebc12cff266750efb2d8d9bfd11a33cd748e9f544a12f",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S5sh9::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-19: 「눈안에 다시 눈이 있는 형태가 되버렸어 로봇이」 — 변환본",
  "at": "2026-09-19T15:25:15.770202+00:00",
  "rejected_cine": {
   "file": "S5sh9_cine.png",
   "staged_sha256": "617e55290eab768dae735553b450cce3b28f0259205f1deffc5f83482adbe1fc",
   "source_sha256": "a15a7d8290428f184ab9893dd66a0410ffb4b5afce07b0fb7c8e1225b88eb810",
   "fingerprint": "7a3cf5a1f9a848b5e2fb45c0142859c1c142760a880a54126e12fb8e2d317b9f",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S88sh7::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-19: 「이번에는 얼굴이 완전히 사람이야」 — 변환본. 원본은 로봇 얼굴",
  "at": "2026-09-19T15:25:15.856810+00:00",
  "rejected_cine": {
   "file": "S88sh7_cine.png",
   "staged_sha256": "aff566e2d8800a79dcb7e95e696aa05001b17f0ea30c95b7d6b0e90f0e842f5d",
   "source_sha256": "7d080a0d4f6e8fe08484f88103c0d363f9d7478cbcf9a2fd6a2e0a5c236a9a73",
   "fingerprint": "c1209631d4c224358be2e59d030fad1d187034714947141e50e21bd9eac01756",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S88sh24::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-19: 「찰리와 함께 현우만 나오는거 아냐? 현우가 두명」 — 변환본. 원본은 현우 하나",
  "at": "2026-09-19T15:25:15.942646+00:00",
  "rejected_cine": {
   "file": "S88sh24_cine.png",
   "staged_sha256": "a804b50e8fb1b9dc66be7551c71779f1cea98a9d6b753ec015cbb708c67c65e7",
   "source_sha256": "da6692e6ae88397a49e23c8c10011710b17e5befe65242fb6b15134f1879cae2",
   "fingerprint": "08612094e655a5862499c0d8d8edfce3121fb2475cf52b09aaa1411015cb7f8a",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S90sh9::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-19: 「챨리 홀로그램 색상이 완전히 실사야 푸른색이 아니고」 — 변환본. 원본은 푸른 홀로그램",
  "at": "2026-09-19T15:25:16.030079+00:00",
  "rejected_cine": {
   "file": "S90sh9_cine.png",
   "staged_sha256": "773cc223a59cf0ccddb4a6e7d785b57f4863629ec91ac7cd28f934645d748a23",
   "source_sha256": "bf65b76a57907d1a378f4bdf8d1f2d8db7b266365e7873de8c900edfca28671f",
   "fingerprint": "8cd74973d0dffb728df2f857d0d55f5b140a36ba2f1836c4404c9a25904468cf",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S59sh18::cine::fb_grok": {
  "applied": false,
  "attempted_at": "2026-09-19T18:45:35.214631+00:00",
  "fingerprint": "591254e8ad682ce0f22b290bdedc827773180182c4c740401d97f6e681322712",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-quality",
  "multiroll_tag": "still_S59sh18_cine_fb_grok",
  "slot": "fb_grok",
  "pack": "24.202608252115",
  "source_file": "S59sh18_sel.png",
  "source_sha256": "3440298fbed40e27061050e5fe60f9b68d3d7f14255b26d8187cf97afbff013b",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"code\\\":\\\"imagine:content-moderated\\\",\\\"error\\\":\\\"Generated image rejected by content moderation.\\\",\\\"usage\\\":{\\\"cost_in_usd_ticks\\\":600000000}}\",\"provider_name\":\"xAI\",\"is_byok\":false}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S59sh18::cine::fb_mai": {
  "applied": false,
  "attempted_at": "2026-09-19T18:45:46.994859+00:00",
  "fingerprint": "b7ad7a169b94b3bdd5eaf29d74a52ac9014e1ce2f712c15974fc8155367cf85b",
  "fingerprint_version": 2,
  "provider": "mai",
  "endpoint": "openrouter/chat-completions",
  "model": "microsoft/mai-image-2.6",
  "multiroll_tag": "still_S59sh18_cine_fb_mai",
  "slot": "fb_mai",
  "pack": "24.202608252115",
  "source_file": "S59sh18_sel.png",
  "source_sha256": "3440298fbed40e27061050e5fe60f9b68d3d7f14255b26d8187cf97afbff013b",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"error\\\":{\\\"code\\\":\\\"content_safety_violation\\\",\\\"message\\\":\\\"Input content violated imagegen safety policies.\\\",\\\"details\\\":\\\"Input content violated imagegen safety policies.\\\"}}\",\"provider_name\":\"Azure\",\"is_byok\":false,\"provider_error_code\":\"content_safety_violation\"}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S59sh18::cine::fb_seedream": {
  "applied": true,
  "attempted_at": "2026-09-19T18:46:20.914322+00:00",
  "fingerprint": "cfc795848814aff3b5582511f4fcffcfa51bf18e8dec36ed25d758135f7fe847",
  "fingerprint_version": 2,
  "provider": "seedream",
  "endpoint": "openrouter/images+aspect_ratio",
  "model": "bytedance-seed/seedream-5-0-pro",
  "multiroll_tag": "still_S59sh18_cine_fb_seedream",
  "slot": "fb_seedream",
  "pack": "24.202608252115",
  "source_file": "S59sh18_sel.png",
  "source_sha256": "3440298fbed40e27061050e5fe60f9b68d3d7f14255b26d8187cf97afbff013b",
  "file": "S59sh18_cine_fb_seedream.png",
  "staged_sha256": "bbe1565a2741b41620a8ce5726f4f2b8db48523927e43324bf5093eb1ada8a5b",
  "latency_ms": 105424
 },
 "S60sh52::cine::fb_grok": {
  "applied": false,
  "attempted_at": "2026-09-19T18:50:28.314870+00:00",
  "fingerprint": "df76b877d3627d2bcf0bcf8564c1b5e1796e414666fb0432c706068d72ecdbee",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-quality",
  "multiroll_tag": "still_S60sh52_cine_fb_grok",
  "slot": "fb_grok",
  "pack": "24.202608252115",
  "source_file": "S60sh52_sel.png",
  "source_sha256": "379fe86ed947aad1b47cd62cd811540512e0dd2bb758f344743f49dded9f4e1e",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"code\\\":\\\"imagine:content-moderated\\\",\\\"error\\\":\\\"Generated image rejected by content moderation.\\\",\\\"usage\\\":{\\\"cost_in_usd_ticks\\\":600000000}}\",\"provider_name\":\"xAI\",\"is_byok\":false}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S60sh52::cine::fb_mai": {
  "applied": false,
  "attempted_at": "2026-09-19T18:50:34.742540+00:00",
  "fingerprint": "767ad6855f4b4ca88fba3a0fa9e77923de6e95d6304e7b3674befc9d2d2fb1b0",
  "fingerprint_version": 2,
  "provider": "mai",
  "endpoint": "openrouter/chat-completions",
  "model": "microsoft/mai-image-2.6",
  "multiroll_tag": "still_S60sh52_cine_fb_mai",
  "slot": "fb_mai",
  "pack": "24.202608252115",
  "source_file": "S60sh52_sel.png",
  "source_sha256": "379fe86ed947aad1b47cd62cd811540512e0dd2bb758f344743f49dded9f4e1e",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"error\\\":{\\\"code\\\":\\\"content_safety_violation\\\",\\\"message\\\":\\\"Input content violated imagegen safety policies.\\\",\\\"details\\\":\\\"Input content violated imagegen safety policies.\\\"}}\",\"provider_name\":\"Azure\",\"is_byok\":false,\"provider_error_code\":\"content_safety_violation\"}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S60sh52::cine::fb_seedream": {
  "applied": true,
  "attempted_at": "2026-09-19T18:51:00.184763+00:00",
  "fingerprint": "cbd2c149670c02ccfe399572aee7a457cca1735376be5ceb017939ebbad300f0",
  "fingerprint_version": 2,
  "provider": "seedream",
  "endpoint": "openrouter/images+aspect_ratio",
  "model": "bytedance-seed/seedream-5-0-pro",
  "multiroll_tag": "still_S60sh52_cine_fb_seedream",
  "slot": "fb_seedream",
  "pack": "24.202608252115",
  "source_file": "S60sh52_sel.png",
  "source_sha256": "379fe86ed947aad1b47cd62cd811540512e0dd2bb758f344743f49dded9f4e1e",
  "file": "S60sh52_cine_fb_seedream.png",
  "staged_sha256": "5176391cbe1e93564c01ab0b64d6fe9c0bac83dcb15fce99368be47aa5a00895",
  "latency_ms": 116563
 },
 "S85sh11::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-19: 「가슴 부분 모양도 다르고, 눈과 눈주변이 모두 사람얼굴과 합성」 — 09-20 재생성 뒤 변환본이 로봇 머리 위에 사람 턱·입을 그림(원본은 로봇 얼굴)",
  "at": "2026-09-19T20:28:55.902108+00:00",
  "rejected_cine": {
   "file": "S85sh11_cine.png",
   "staged_sha256": "b68fafa65059258b4ed83a9672563fd4840c20d071ec602407ebe7fbbf0e1810",
   "source_sha256": "183b99f30fc03d60a70aa4573a32e26baeb7d53b8c22c0325a635170cef2e87e",
   "fingerprint": "70d432c0c1c8bd63f7ae88060c91327cf30cbae12fe6e14eaf870973bb2bc17b",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S20sh6::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-20: 「갑자기 페드로가 경찰 복 입고 있어」 — 변환본만. 원본은 경찰이 경찰복",
  "at": "2026-09-20T02:10:10.139219+00:00",
  "rejected_cine": {
   "file": "S20sh6_cine.png",
   "staged_sha256": "c211685d503f41449d59a4d79699dbea4d736931d2ef09394640895ebfc0904c",
   "source_sha256": "c172a85989425f3bd212cfd185f0b24a473807c5d0dbb325809cad539609ff31",
   "fingerprint": "5ff872d8bac6d6661d980bdf3abbca159535d52b3ddba1be585b86d3ef29484b",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S79sh10::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-20: 「또 사람 얼굴이 나와」 — 변환본만. 원본은 붉은 눈 로봇 가면",
  "at": "2026-09-20T02:10:10.222628+00:00",
  "rejected_cine": {
   "file": "S79sh10_cine.png",
   "staged_sha256": "d0a85a7892b608f7f297d423cd0ba7d63febf4e181f2513a15cb5f7bebbc2f5a",
   "source_sha256": "979462f3ea3883f87b379229412a8ca135c23d0c1deaf08a4ff5688010ebceee",
   "fingerprint": "2270e89a0b53e423744e31fd2adc798b11793a068ed9947fdb4f80a72e8b6e88",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 },
 "S91sh5::cine_human": {
  "decision": "keep_original",
  "reason": "사용자 육안 2026-09-20: 「또 사람 얼굴이 나와」 — 변환본이 왼쪽 위에 사람 얼굴을 더했다",
  "at": "2026-09-20T02:10:10.304934+00:00",
  "rejected_cine": {
   "file": "S91sh5_cine.png",
   "staged_sha256": "e42655c9e70d7ff908d2d37476ad59e28d6c60c524b738045b3057f2e688c030",
   "source_sha256": "5bf0ba2da94a9e2dd2415c29edbcf50451868254516eca0223745e056ed87424",
   "fingerprint": "e4b7079f4586f9bb29941039a7635c1fddf44d9cd203f346e51bd82921084e4a",
   "provider": "xai",
   "model": "grok-imagine-image"
  }
 }
}